OCR — read the text off a scan

Read the text in a scan, on your own machine.

Drop your scan here

or click to choose it — pick a language on the next step

Up to 60 MB · no account, no watermark · stays on this device

The heavy job is the one to keep at home.

OCR is what people usually hand to a server, because it is expensive. Expensive is the reason not to: a scanned passport, medical letter or contract is the last thing you want sitting in someone's processing queue.

The scan is kept, always.

The words are written invisibly behind the original image. We do not offer to replace the scan with typeset text: recognition is never perfect, and keeping the picture means a misread can never damage the document.

One download, then nothing.

The language model comes from this site the first time you use it and is then cached. That is the only request this page ever makes, and the footer counts it. After that, OCR works with no connection like everything else here.

Questions

Is it safe to read the text off a PDF online?

Nothing is uploaded, so there is no copy on a server to trust. The counter at the foot of this page shows bytes sent, and it stays at zero.

Does this work offline?

After the first run, yes. The language model is downloaded once, and then it works with the network off like everything else here.

Is there a file size limit?

60 MB. Past that a browser tab stops being able to write the finished file out and sits unresponsive for minutes — it does not crash, it stops being usable — so we refuse rather than hang.

Do I need an account?

No. There is nothing to sign in to, and no limit on how many files you run through it.

Is there a watermark?

No. Nothing is added to the file.

What happens to my file?

It stays on your device. The counter at the foot of every page shows bytes sent and third-party requests, so you can check that rather than take our word for it.