4-part guide

OCR and scanned documents: turning pictures of words back into words

Short answer

OCR converts the pixels of a scanned page into machine-readable characters and stores them as an invisible text layer aligned with the image. The page still looks identical, but it becomes searchable, selectable and copyable. Accuracy runs above 98 percent on clean 300 DPI print and far lower on handwriting.

A scanned page is a photograph. Your computer sees a grid of grey pixels where you see a sentence, which is why Ctrl+F finds nothing and copy-paste returns an empty clipboard. Optical character recognition reads those pixels and works out which letters they represent, then writes the result back into the file as an invisible text layer sitting exactly on top of the image. These guides cover running OCR on a PDF, pulling text out of photos and screenshots, testing whether a file is searchable, and rescuing bad scans.

Every guide in this topic

Tools for this job

Other topics