Make a scanned PDF searchable

A scanned document is a picture of text, which is why searching it finds nothing. This reads the words with optical character recognition and writes them back as an invisible text layer, so the page looks the same but Ctrl+F starts working.

How to make a scanned PDF searchable

  1. 1Open the scanned PDFThe tool checks whether the document already has a text layer, so you do not run OCR on a file that does not need it.
  2. 2Pick the language and qualityTwelve languages are available, including English, Spanish, French, German, Chinese, Japanese, Korean, Arabic, Hindi and Russian. Higher DPI reads small print better and takes longer.
  3. 3Run the recognitionPages are processed one at a time with progress shown. The result keeps the original page images with the recognised text placed invisibly behind them.

Questions

Are my files uploaded to a server?

No. The file is read by your browser and processed in the page. It never reaches a server, so there is no upload to wait for, no queue, and no copy of your document sitting somewhere with a retention policy.

Which languages are supported?

English, Spanish, French, German, Portuguese, Italian, Simplified Chinese, Japanese, Korean, Arabic, Hindi and Russian. The recognition data for the language you pick is downloaded the first time you use it.

Why is it slow?

Every page is rendered to an image and then read character by character, in your browser rather than on a server farm. A long document can take several minutes. Higher DPI settings are more accurate and slower.

How accurate is it?

Clean, straight, high-resolution scans of ordinary printed text read very well. Handwriting, faint photocopies, unusual fonts and skewed pages are where accuracy drops.

Does the page still look the same?

Yes. The original page image is kept and the recognised text is placed invisibly behind it, so nothing shifts visually but the words become selectable.

Related tools