Scanned PDF or searchable PDF? How to tell, and how to convert
, 5 min read
Two PDFs can look identical and behave completely differently. One is a document. The other is a photograph of a document.
The two second test
Open the file and try to select a line of text with your mouse. If you get a neat blue highlight over the words, the PDF contains real text. If you get a rectangle over the whole page, or nothing at all, it is a scan.
You can also press Ctrl+F and search for a word you can see. No results means no text layer.
What OCR adds
Optical character recognition looks at the picture, recognises the shapes as letters, and writes the result into the file as an invisible layer sitting exactly behind the visible image. The page looks unchanged, but now it can be searched, selected and copied.
Make PDF Searchable does this and keeps the original appearance. If you only want the words, Scanned PDF to Text gives you a plain text file.
Getting better accuracy
Recognition quality depends almost entirely on the input:
- Aim for text that is sharp when you zoom to 100 percent. Blur is the main cause of errors.
- Photograph pages flat and square, not at an angle.
- Use even light. A shadow across the page confuses the engine more than dim light does.
- Pick the right language, and choose English plus Hindi for mixed documents.
If your pages are phone photos rather than scanner output, run Document Scanner first. It flattens the lighting, whitens the paper and darkens the ink, which usually lifts the confidence score noticeably.
Why bother
A searchable archive changes how you work. Years of bills and statements become findable in seconds. Your operating system indexes them, so a search for a policy number surfaces the right file without you remembering where you put it.
The OCR engine runs in your browser, so medical reports and bank statements never travel anywhere. The first run downloads the recognition engine, then it is cached for later.