OCR — read the text off a scan
Read the text in a scan, on your own machine.
Drop your scan here
or click to choose it — pick a language on the next step
The heavy job is the one to keep at home.
OCR is what people usually hand to a server, because it is expensive. Expensive is the reason not to: a scanned passport, medical letter or contract is the last thing you want sitting in someone's processing queue.
The scan is kept, always.
The words are written invisibly behind the original image. We do not offer to replace the scan with typeset text: recognition is never perfect, and keeping the picture means a misread can never damage the document.
One download, then nothing.
The language model comes from this site the first time you use it and is then cached. That is the only request this page ever makes, and the footer counts it. After that, OCR works with no connection like everything else here.
You can switch to another tab while this runs — it keeps going in the background. Closing the tab stops it for good; nothing is queued anywhere else.
What to expect
- Clean printed text: near-perfect.
- Faxes and photocopies: good, with occasional wrong characters.
- Handwriting: not attempted. We will say which pages we skipped.
- Tables keep their words, not their grid.
Here vs. the app
On this page OCR is unlimited and needs no account, one file at a time. The Android app does not have OCR yet.
Free here means all of it, one file at a time, with no account: reading a document of any length, as many documents as you like. What Pro adds is a different capability rather than a bigger allowance — doing many files in one go, and writing the words back into the PDF so the file itself is searchable.
Reading on your device
0%
- Loading the language model waiting
- Reading each page waiting
- Writing the text behind the scan waiting
One block per page. Filled blocks are pages already read.
This is your own processor working through the pages, several at a time. There is no queue in front of you and no faster tier — the only thing between you and the result is arithmetic.
Done — on this device
- words found
- —
- characters
- —
- your file
- untouched — nothing was written back into it
- time taken
- —
Wanting a searchable PDF instead?
That is a different thing from the text above: the same words written back into the document as an invisible layer, so the file itself becomes searchable in any reader while looking exactly as it does now. It is what Pro will do with this page, and Pro is not on sale yet, on either surface. There is no button here for it because there is nothing behind the button — when it exists, it will appear.
Reading documents at work? We are asking firms what they need.
Questions
Is it safe to read the text off a PDF online?
Nothing is uploaded, so there is no copy on a server to trust. The counter at the foot of this page shows bytes sent, and it stays at zero.
Does this work offline?
After the first run, yes. The language model is downloaded once, and then it works with the network off like everything else here.
Is there a file size limit?
60 MB. Past that a browser tab stops being able to write the finished file out and sits unresponsive for minutes — it does not crash, it stops being usable — so we refuse rather than hang.
Do I need an account?
No. There is nothing to sign in to, and no limit on how many files you run through it.
Is there a watermark?
No. Nothing is added to the file.
What happens to my file?
It stays on your device. The counter at the foot of every page shows bytes sent and third-party requests, so you can check that rather than take our word for it.