OCR PDF
Make scanned PDFs searchable with OCR.
How to use OCR PDF
Upload the scanned or image-only PDF.
Choose the language of the document text.
Click Run OCR to recognize the text on every page.
Download the searchable PDF.
What OCR adds to a scan
A scanned PDF is a stack of photographs. Your computer can display it and can do nothing else with it: search finds nothing, copy returns nothing, a screen reader has nothing to read. OCR examines the pixels, recognises the characters, and writes a text layer invisibly behind the image.
The page looks exactly the same afterwards — that is deliberate. What changes is that the document becomes searchable, quotable, indexable by your operating system and accessible to assistive technology. A folder of a hundred scanned invoices goes from unusable to searchable in one pass.
What determines whether OCR works well
Accuracy on a clean 300 DPI scan of printed text is typically 98-99%. It falls off sharply with input quality, and the fixes are all on the scanning side rather than the recognition side.
- Resolution: 300 DPI is the sweet spot. Below 200 DPI accuracy drops steeply.
- Straightness: a page scanned at a slight angle loses accuracy line by line.
- Contrast: crisp black on white beats a grey photocopy of a photocopy.
- Typeface: ordinary serif and sans-serif text is easy; decorative and script faces are not.
- Handwriting: expect poor results. Printed text is the design target.
Proofreading the result
Even at 99% accuracy, a 2,000-word page contains around twenty errors. They cluster in predictable places: 1 versus l versus I, 0 versus O, rn read as m, and punctuation that vanishes. A spellchecker will not catch most of them, because the results are usually still valid words.
Check numbers first and hardest. A misread digit in an invoice total, an account number or a dosage is the error that actually costs something, and it is the one no automated check will flag for you.
OCR without sending your documents anywhere
Every step runs inside your browser, using the WebAssembly and Canvas APIs your device already ships with. No file is uploaded, no copy is kept, and nothing sits in a queue on a server waiting to be deleted — which is also why the tool keeps working with the network disconnected.
Cloud OCR services process the entire content of every document you send, and the documents worth OCR-ing are the archives: medical records, legal files, historical correspondence, financial paperwork. Recognising them locally means the text is extracted on your machine and stays there.
Frequently Asked Questions
What does OCR do to my PDF?
OCR (optical character recognition) reads the images in a scanned PDF and adds an invisible text layer on top, so you can search, select and copy the text while the pages still look the same.
Which languages are supported?
English, Spanish, French, German, Italian and Portuguese. Pick the language the document is written in: the model is downloaded once (a few MB) and cached for next time.
Is OCR done in the browser?
Yes. Recognition runs locally on your device, so your scanned documents are never uploaded to a server.
Will the page look different after OCR?
No. The original scan is kept exactly as it is, and the recognised text is placed invisibly behind it. You see the same page; your computer now also sees words.
How accurate is it?
Around 98-99% on a clean, straight 300 DPI scan of printed text. Lower resolution, skew, poor contrast or unusual fonts each cost accuracy. Handwriting is not reliably recognised.
Which languages are supported?
The major Latin-script languages, including English, Spanish, French, German, Italian and Portuguese. Selecting the right language matters: it loads the correct dictionary and accented characters, which measurably improves accuracy.
Can I OCR a photo of a document rather than a scan?
Yes, though results are worse. Photographs bring perspective distortion and uneven lighting. Shoot flat from directly above in even light, or use your phone's document scanner mode, which corrects both.
Does OCR make the file bigger?
Slightly. The text layer is a few kilobytes per page against images measured in megabytes, so the increase is usually under 1%.
Related tools
Keep reading
- How to Edit a Scanned PDF
Your PDF has no text in it — the pages are photographs. Here are the three ways round that, and how to tell which one your document needs.
Looking for something else? Browse all pdf tools.