ToolGenie
PDFPublished 3 min read

How to Edit a Scanned PDF

You open a PDF, try to change a date, and nothing happens. The cursor will not select anything. The text tool does not find any text. It looks like a document and behaves like a picture.

That is because it is a picture. If the file came from a scanner, a photocopier or a phone camera, each page is a photograph of a page, and there is no text in the file to edit. Everything that feels broken about editing it follows from that one fact.

Confirm it in two seconds

Open the PDF and try to drag-select a sentence. If the text highlights, you have a real text PDF and a different problem. If nothing highlights — or a whole page highlights as one block — it is a scan.

A second check: use the reader's search for a word you can plainly see on the page. If it finds nothing, there is nothing to find.

Option one: overlay (fastest, works always)

Do not try to change what is on the page — put something new on top of it. Cover the wrong figure with a white box and type the right one over it. Add your signature. Fill in the blank lines. Tick the boxes.

This works on any PDF, scanned or not, needs no conversion, and cannot damage the layout because it never touches it. For the overwhelming majority of edits — filling a form, correcting a date, signing — this is not a workaround. It is the right answer.

  • Match the white box to the page background: a scan is rarely pure white, so sample the colour rather than assuming.
  • Match the font size to the surrounding text, and pick a similar typeface.
  • Zoom to 100% to check the alignment. Corrections that sit a millimetre off the baseline are obvious in print.

Option two: OCR, then edit properly

OCR reads the pixels, recognises the characters, and adds a real text layer to the document. The page looks identical afterwards; what changes is that the file now contains text — searchable, copyable, and readable by a screen reader.

This is what you want when the document is going into an archive, when you need to quote from it, or when you will need to find it again by its contents. It is also the necessary first step if you intend to convert it to Word.

Option three: convert to Word and rewrite

If you need to genuinely rewrite paragraphs rather than patch a line, run OCR and then convert to Word. You get an editable document with real text you can restructure freely.

Expect to do some tidying. OCR is 98-99% accurate on a clean 300 DPI scan and worse on anything less, and conversion has to reconstruct the layout by inference. Proofread the numbers especially hard — a misread digit in an amount or an account number is the error that costs something, and no spellchecker will flag it because it is still a valid number.

Getting a better scan in the first place

Everything downstream is easier if the scan is good, and most bad scans come from photographing a page badly rather than from bad equipment.

  • 300 DPI. Below 200 DPI, OCR accuracy falls off steeply.
  • Straight. A page scanned at a slight angle loses accuracy line by line.
  • Even lighting. Photograph in daylight from the side, never under a lamp that casts your own shadow.
  • Directly overhead. Shooting at an angle produces perspective distortion that OCR handles poorly.
  • Use your phone's document scanner mode if it has one — it corrects perspective and flattens lighting automatically.

One thing to be careful with

If you are covering something rather than correcting it — hiding an account number before sending a statement, for instance — a white or black box is not enough on a *text* PDF, because the text underneath survives. On a scan there is no text underneath, so a box genuinely does hide it.

The catch is knowing which you have. If you ran OCR first, your scan now contains text, and a drawn box no longer hides anything. Redact properly, or cover before you OCR.

Common questions

Why can't I select the text in my PDF?

Because there is no text in it. The pages are images from a scanner or camera. Run OCR to add a text layer, or overlay your changes on top, which works either way.

Is OCR accurate enough to rely on?

On a clean, straight 300 DPI scan of printed text, 98-99%. That still means around twenty errors in a 2,000-word page, clustered around 1/l/I and 0/O. Always proofread, and check numbers first.

Can OCR read handwriting?

Not reliably. OCR is built for printed text, and handwriting recognition is a much harder problem with far worse results. Expect to transcribe handwritten sections yourself.

Tools mentioned in this guide

Everything below runs in your browser. No file is uploaded.

Keep reading

All guides