Converting
How to Convert a Scanned Document to a Searchable PDF
Take scans or images of paper and produce a PDF you can search, copy from and file — with the right settings at every step.
A scan is a photograph of a page. It looks like a document and behaves like a picture: you cannot search it, you cannot copy a line out of it, and a screen reader finds nothing on it. OCR is what closes that gap.
From paper to searchable PDF
- 1Scan or photograph the pages. 300 DPI greyscale is right for ordinary printed text.
- 2Convert images to PDF if your scanner produced JPGs rather than a PDF.
- 3Merge the pages into one document if they came out separately.
- 4Remove blank pages and rotate anything sideways.
- 5Run OCR to add a text layer.
- 6Compress last, once everything else is settled.
Scanner settings that matter
| Setting | Recommended | Why |
|---|---|---|
| Resolution | 300 DPI | Below 200 the OCR accuracy falls away; above 400 adds size, not accuracy |
| Colour mode | Greyscale | Roughly a third the size of colour with no loss on black-and-white originals |
| Format | PDF or JPEG | TIFF is enormous and rarely needed |
| Straightening | On | Skewed text is the single biggest cause of OCR errors |
What OCR actually does
OCR reads the picture, recognises the shapes as characters, and writes an invisible text layer positioned exactly over the visible image. The page still looks like the scan — the original picture is untouched — but a search now finds words, and selecting text copies real characters.
It is not perfect. Clean printed text at 300 DPI is recognised extremely well; a faint fax, a handwritten note or a dense table with faint rules will produce errors. Always spot-check by searching for a word you know appears in the document.
Why the order matters
OCR reads the image as it is at the time you run it. Compress first and you have OCR reading a degraded picture, which produces more mistakes than necessary. Do every lossless step first, run OCR, then compress — the text layer is unaffected by compression because it is text, not image.
Do it now
OCR PDF
Extract text from scanned documents. It runs in this browser tab — your file is not uploaded anywhere.
Open OCR PDFFrequently asked questions
What DPI should I scan at?+
300 DPI for text. Higher adds file size without helping recognition; lower starts to cost you accuracy.
Can OCR read handwriting?+
Not reliably. It is built for printed characters. Neat block capitals sometimes work; ordinary cursive does not.
Does OCR change how my scan looks?+
No. The text layer is invisible and sits behind the image. The page looks exactly the same.
Should I OCR before or after compressing?+
Before. OCR on a compressed image makes more mistakes, and compressing afterwards does not damage the text layer.