Skip to content
freeonlinefileconvert

OCR PDF

Add a hidden layer of text to a scanned PDF so you can search it and copy from it, while every page looks exactly the same.

or drop it here · up to 50 MB, free

Bigger files? Pro converts files up to 500 MB

This is bigger than the free version

You can still do it: convert it now and pay €1.99 only if you want to download the result.

Free covers files up to 50 MB and up to 20 files at a time.

Free · No sign-up · No watermark · Your files never leave your device

OCR PDF in 3 steps

  1. Select your scanned PDF, or drop it on the box.
  2. Choose the language of the text (and a second one if the document mixes two), then click Run OCR.
  3. Download the searchable PDF, or all the recognised text as a TXT file.

How OCR turns a scan into a searchable PDF

A scanned page is only a picture of paper: you cannot search it, select a sentence or copy a figure from it. OCR (optical character recognition) reads the letters in that picture and adds them to the page as an invisible layer of text, placed exactly over each word. The page looks the same as before, but Ctrl+F now finds words, you can highlight and copy text in Acrobat, Preview, Chrome or Edge, and Windows, macOS, Google Drive or your document system can index the file.

The text is recognised on your device by Tesseract, an open-source OCR engine. Each page is read at up to 300 DPI; pages that already contain text are kept as they are, and pages scanned sideways or upside down are turned the right way up. It works best on clear printed text such as letters, invoices, contracts, books and forms. Handwriting, very small print, faded pages and photos taken at an angle give more mistakes, and tables come out as text, not as cells. The first time you use a language, its reading model (1 to 3 MB) is downloaded from this site and kept in your browser.

Frequently asked questions

Will my PDF look different after OCR?

No. The original pages, pictures and drawings are not changed or compressed again. OCR only adds an invisible text layer on top of each page, so the file looks the same and grows by a few kilobytes per page. Pages scanned sideways are turned upright so they are easier to read.

Which languages can it read?

Twenty-four: English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Turkish, Indonesian, Vietnamese, Japanese, simplified and traditional Chinese, Korean, Russian, Ukrainian, Arabic, Hindi, Catalan, Czech, Greek, Romanian and Swedish. Pick two at once for documents that mix languages, such as Spanish and English.

How accurate is the text recognition?

On a clean scan of printed text, nearly every word is right: in our tests on office scans at 200 DPI, more than 98 percent of the characters were read correctly. Handwriting, tiny or decorative fonts, blurry scans and coloured backgrounds lower the accuracy, so check names and numbers that matter.

How long does OCR take?

A few seconds per page on a recent computer and roughly twice as long on a phone, plus a short download of the language model the first time. A 20-page scan usually takes one to three minutes. Keep the tab open while it works; pages that already have text are skipped.

Can I get the text as a TXT file?

Yes. Next to the searchable PDF you can download all the recognised text as a TXT file, page by page, ready to paste into Word, Google Docs or an email. For photos and screenshots, use our Image to Text tool, which works the same way.

Is my PDF uploaded to a server?

No. The pages are read by an OCR engine that runs inside your browser, on your own device, so contracts, medical letters and ID documents are never sent to us or stored anywhere. Only the OCR engine and the language model are downloaded from this site, once, and kept in your browser.

Related tools

You are one of the first

FreeOnlineFileConvert has just launched. Create a free account now and join our Family & Friends list: the full Pro plan, free for life.

Pro0 €6,99 €
Join for free