Make scanned PDFs searchable with OCR
A scanned document is really a stack of photographs. It looks like text, but your computer only sees pixels, so you cannot search for a name, copy a paragraph or let a screen reader read it aloud. OCR changes that. It recognizes the characters on each page and places an invisible text layer on top of the image, turning a flat scan into a PDF you can actually work with.
This is invaluable for digitized archives, old contracts, receipts photographed on a phone, printed reports scanned at the office, and research material from books or library copies. Once recognized, a file becomes findable in your desktop search and much easier to quote from.
Thirteen languages, one per document
Choosing the right language matters because the recognizer uses language-specific knowledge of letters and words. The Document language list covers English, German, French, Spanish, Italian, Portuguese, Polish, Dutch, Turkish, Russian, Ukrainian, Japanese and Chinese (Simplified). Each run uses a single language, so a document that mixes two languages will be recognized best in whichever one dominates.
The Scan quality setting controls the resolution each page is rendered at before recognition. More accurate uses a higher resolution and catches small or faint characters more reliably. Faster trades a little accuracy for speed, which can be worthwhile on long, clean documents.
Recognition runs on your own hardware
Nothing is shipped off to a cloud engine here. The recognition software runs inside your browser, reading each rendered page on your device and building the text layer locally. The only thing fetched from the site is the language data needed for recognition, never your document. Scanned passports, contracts and medical forms therefore stay where they started.
Because your own processor does the work, speed depends on your device. A recent laptop moves through pages briskly, while an older phone can take a good while on a long document. Keep the tab open until the progress finishes.
Honest limitations
OCR is powerful but not perfect. Expect occasional errors with handwriting, heavy stamps, unusual fonts, tight tables or pages scanned at an angle. The output pages are rebuilt from rendered images of the originals with the recognized text layered underneath, so any text that was already selectable is replaced by the new recognized layer, and file size may differ from the original. Rotate crooked pages first with Rotate PDF for better results.
What to do after OCR
Pull the recognized words into a plain file with PDF to Text, or rebuild them as an editable document with PDF to Word. Scans often produce large files, so Compress PDF is a natural next step. If your scans are still loose photos, combine them into one document with JPG to PDF before running recognition.
Frequently asked questions
What does OCR do to a PDF?
OCR, or optical character recognition, reads the letters in a scanned image and turns them into real text. The result looks the same as your scan but gains an invisible text layer, so you can search, select and copy the words.
Which languages does OCR PDF support?
Thirteen: English, German, French, Spanish, Italian, Portuguese, Polish, Dutch, Turkish, Russian, Ukrainian, Japanese and Chinese (Simplified). Pick one per run from the Document language list.
Is OCR PDF free and private?
Yes. There is no charge or account, and recognition runs in your browser rather than on a server. Your scans stay on your device the entire time.
How long does OCR take?
It depends on the number of pages, the scan quality setting and the power of your device. A few pages usually finish quickly, while long documents on an older phone take noticeably longer. The first run for a language also loads its recognition data.
Should I choose Faster or More accurate?
More accurate renders pages at a higher resolution and is the better default, especially for small print. Faster can be enough for large, clear type and saves time on long documents.
Why are there mistakes in the recognized text?
OCR accuracy depends on the source. Blurry photos, skewed pages, handwriting, decorative fonts and low-contrast scans produce more errors, as does choosing the wrong language.
Can I get the recognized text as a Word or text file?
Yes. Run OCR first, then open the searchable result in PDF to Text for a plain .txt file or in PDF to Word for an editable document.
