Cover illustration

You've scanned a document, saved it as a PDF, and now you want to search for a word or copy a paragraph — but you can't. The text isn't selectable. Why? Because a scanned PDF is just an image of text, not actual text. It's like taking a photo of a book page: you can see the words, but your computer can't read them.

The solution is OCR — Optical Character Recognition. OCR software analyzes the image, recognizes the characters, and adds a hidden text layer behind the image. Suddenly, your scanned PDF becomes searchable, selectable, and editable.

This guide explains everything you need to know about PDF OCR: how it works, which tools to use, how to get the best accuracy, and how to troubleshoot common problems.

What Is OCR and How Does It Work?

OCR (Optical Character Recognition) is technology that converts images of text into actual text characters. Here's a simplified version of what happens:

  1. Image preprocessing: The software deskews (straightens) the image, adjusts contrast, removes noise, and binarizes (converts to black and white).
  2. Layout analysis: It identifies text blocks, lines, words, and individual characters, distinguishing text from images and tables.
  3. Character recognition: For each character, it compares the shape against known character patterns (using pattern matching or machine learning).
  4. Post-processing: It uses dictionaries and grammar rules to correct errors (e.g., "he1lo" → "hello").
  5. Output: The recognized text is placed as an invisible layer behind the original image, making the PDF searchable.
Modern OCR = AI: Today's best OCR tools use deep learning (neural networks) trained on millions of documents. This gives them much better accuracy than older pattern-matching systems, especially for poor-quality scans, handwritten text, and complex layouts.

When Do You Need OCR?

You need OCR when your PDF is:

You don't need OCR when your PDF was created digitally (from Word, Excel, a website, etc.) — these already have a text layer.

How to Tell If a PDF Needs OCR

Try these simple tests:

  1. Select test: Open the PDF and try to click and drag to select text. If you can select individual words, it has a text layer (no OCR needed). If you select the whole page as an image, it needs OCR.
  2. Search test: Press Ctrl+F (Cmd+F on Mac) and search for a word you can see on the page. If "No matches found," it probably needs OCR.
  3. Copy test: Copy some text and paste it into Notepad. If you get gibberish or nothing, it needs OCR.

Best Free OCR Tools

Desktop Tools

1. Tesseract OCR (Free, Open Source)

Tesseract is the most popular open-source OCR engine, originally developed by HP and now maintained by Google. It supports 100+ languages and is used by many other tools under the hood.

2. Microsoft OneNote (Free)

OneNote has built-in OCR that's surprisingly good. You can insert an image or PDF printout, right-click, and copy text.

3. Microsoft Lens (Free, Mobile)

Microsoft's scanning app includes OCR and can save directly as searchable PDF or Word.

4. Mac Preview + Shortcuts (Free, Built-in)

macOS has built-in OCR (Live Text) since macOS Monterey. You can select text directly from images in Preview, and use Shortcuts to OCR PDFs.

Online Tools

1. Google Drive (Free)

Upload a PDF or image to Google Drive, right-click → Open with → Google Docs. Google automatically runs OCR and converts it to an editable document.

2. Online OCR (onlineocr.net)

Dedicated free OCR service supporting 40+ languages.

3. i2OCR

Another free online OCR tool with no registration.

4. PDF24 OCR

Part of the PDF24 toolkit, free and unlimited.

Paid Tools (Worth Considering)

ToolPriceBest For
Adobe Acrobat Pro$12.99/monthBest overall accuracy, integrates with PDF workflow
ABBYY FineReader$199 one-timeHighest accuracy, especially for complex layouts
Kofax OmniPage$149 one-timeEnterprise-grade OCR, batch processing
Readiris$129 one-timeGood accuracy, user-friendly interface

Step-by-Step OCR Tutorials

Method 1: Google Drive (Easiest, Free)

  1. Go to drive.google.com and sign in
  2. Click "New" → "File upload" → select your scanned PDF
  3. Once uploaded, right-click the file → "Open with" → "Google Docs"
  4. Wait 10-30 seconds — Google runs OCR automatically
  5. The document opens as editable text in Google Docs
  6. Review and correct any OCR errors
  7. File → Download → PDF Document (.pdf) to save as a new searchable PDF
Pro Tip: Google Docs OCR works best with files under 2MB and 10 pages. For larger files, split them first, OCR each part, then merge.

Method 2: Adobe Acrobat Reader (Free, Searchable PDF)

Note: The free Acrobat Reader doesn't have OCR, but you can use the free online trial or one of the methods above. If you have Acrobat Pro:

  1. Open the scanned PDF in Acrobat Pro
  2. Tools → Scan & OCR
  3. Click "Recognize Text" → "In This File"
  4. Select language and output style
  5. Click "Recognize Text"
  6. Wait for processing
  7. File → Save As to save the searchable PDF

Method 3: PDF24 Creator (Free, Windows Desktop)

  1. Download and install PDF24 Creator (free from pdf24.org)
  2. Open PDF24 Creator → select "OCR" tool
  3. Add your scanned PDF (drag & drop or click to add)
  4. Select language (e.g., English)
  5. Choose output: "Searchable PDF" (keeps original image + text layer)
  6. Click "Start" and wait
  7. Download the searchable PDF

Method 4: Mac Preview / Live Text (Free, Built-in)

  1. Open the scanned PDF in Preview
  2. Hover over text — your cursor should turn into a text selection tool (if Live Text is supported)
  3. Select and copy text as needed
  4. For full PDF OCR, use the Shortcuts app:
    • Open Shortcuts → create new shortcut
    • Add "Extract Text from Image" action
    • Set input to PDF pages
    • Run the shortcut on your PDF

Method 5: Microsoft OneNote (Free)

  1. Open OneNote → create a new note
  2. Insert → File Printout → select your PDF
  3. Wait for the pages to insert as images
  4. Right-click any page image → "Copy Text from Picture"
  5. Paste the text somewhere to edit
  6. (For full PDF, repeat for each page or use a different method)

How to Get the Best OCR Accuracy

OCR accuracy depends heavily on the quality of your input. Follow these tips for 99%+ accuracy:

Scan Quality Tips

  1. Use 300 DPI minimum: 300 DPI is the sweet spot for OCR. Below 200 DPI, accuracy drops sharply. Above 600 DPI, files get huge with little accuracy gain.
  2. Scan in color or grayscale: Color scans give OCR more information to work with. Pure black-and-white (1-bit) scans can lose detail.
  3. Ensure even lighting: Shadows and glare confuse OCR. Use a flatbed scanner or good lighting if photographing.
  4. Straighten the page: Skewed text reduces accuracy. Most scanners have auto-deskew, or use software to straighten before OCR.
  5. Clean the scanner glass: Dust and smudges create artifacts that OCR may misread as characters.
  6. Remove staples and flatten pages: Curved pages cause distortion at the edges.
  7. Use high contrast: Black text on white background is ideal. Faded or colored text is harder.

Software Settings

  1. Select the correct language: Most OCR tools let you choose the document language. Choosing the right one dramatically improves accuracy. For multi-language docs, select all applicable languages.
  2. Use dictionary correction: Enable dictionary/grammar checking in post-processing.
  3. Choose the right output: "Searchable PDF" keeps the original image with an invisible text layer — best for most uses. "Text only" extracts just the text but loses formatting.
  4. Process in batches: For multi-page documents, OCR all pages at once for consistent results.

Pre-Processing

Many scanning apps (Microsoft Lens, Adobe Scan, CamScanner) do this automatically.

OCR for Different Content Types

Printed Text

Accuracy: 98-99.9% with good scans

Printed text is what OCR does best. Clean, high-contrast, 300 DPI scans of standard fonts yield near-perfect results.

Handwritten Text

Accuracy: 60-95% depending on handwriting

Handwriting is much harder. Modern AI-based OCR (Google, Azure, AWS) can handle neat handwriting, but messy cursive is still challenging. For important handwritten documents, expect to manually correct errors.

Tables

Accuracy: 70-95%

Tables are tricky because OCR needs to understand the grid structure. Tools like ABBYY FineReader and Adobe Acrobat Pro have dedicated table recognition. Google Docs can convert tables to HTML tables reasonably well.

Multi-Column Layouts

Accuracy: 80-95%

Newspapers, magazines, and academic papers with multiple columns can confuse OCR's reading order. Advanced tools (ABBYY, Adobe) handle this better than basic ones.

Non-Latin Scripts

Accuracy: Varies by language

OCR supports 100+ languages, but accuracy varies. Latin scripts (English, Spanish, French) are best supported. Chinese, Japanese, Arabic, and Devanagari work well with modern tools but may require specific language packs.

Low-Quality Documents

Accuracy: 50-85%

Faxed documents, old newspapers, photocopies of photocopies, and faded text are challenging. Pre-processing (enhance contrast, despeckle) helps, but manual correction will be needed.

Common OCR Problems and Solutions

"The OCR output is full of errors"

"The text is searchable but formatting is messed up"

"OCR doesn't recognize my language"

"The file is too large for online OCR"

"I can search but can't copy text properly"

OCR and Accessibility

OCR is essential for document accessibility. A scanned PDF without OCR is completely inaccessible to screen readers — blind users can't read it at all. Adding OCR makes the text available to assistive technology.

For fully accessible PDFs, after OCR you should also:

Adobe Acrobat Pro has an "Accessibility" tool that can auto-tag PDFs after OCR. For free tools, LibreOffice can export tagged PDFs.

Batch OCR: Processing Many Files

If you have hundreds of PDFs to OCR, manual processing isn't practical. Here are batch solutions:

Free Batch Options

Paid Batch Options

Cloud APIs (For Developers)

If you need to OCR thousands of files programmatically:

OCR vs. PDF Text Extraction: What's the Difference?

People often confuse these two:

OCRText Extraction
InputImage-based PDF (scanned)Text-based PDF (digital)
ProcessRecognizes characters from pixelsReads existing text layer
Accuracy95-99% (depends on quality)100% (exact copy)
SpeedSlow (1-5 seconds/page)Instant
ToolsTesseract, Adobe, Google DriveAny PDF reader, copy-paste

If your PDF already has a text layer (created from Word, a website, etc.), you don't need OCR — just copy the text directly. OCR is only for scanned/image PDFs.

Frequently Asked Questions

Q: Is OCR 100% accurate?

No. Even the best OCR makes occasional mistakes, especially with poor-quality scans, unusual fonts, or complex layouts. For clean, 300 DPI scans of standard printed text, accuracy is typically 98-99.9%. Always proofread important documents after OCR.

Q: Can OCR read handwriting?

Modern AI-based OCR (Google Cloud Vision, Azure, AWS) can read neat printed-style handwriting with reasonable accuracy (70-90%). Cursive handwriting is still very difficult and may require specialized handwriting recognition tools or manual transcription.

Q: Does OCR work on phone photos of documents?

Yes, if the photo is clear, well-lit, and the document is flat. Apps like Microsoft Lens and Adobe Scan automatically enhance the image (deskew, crop, enhance contrast) before OCR, giving good results. Blurry or angled photos will have poor accuracy.

Q: What's the best free OCR tool?

For most users, Google Drive is the easiest and most accurate free option — just upload and open with Google Docs. For desktop use on Windows, PDF24 Creator is excellent. For Mac, the built-in Live Text works well for individual text selection. For batch processing, OCRmyPDF (based on Tesseract) is the best free option.

Q: Can OCR recognize tables and preserve them in Excel?

Some tools can. Adobe Acrobat Pro and ABBYY FineReader have dedicated table recognition and can export to Excel. Google Docs can convert tables to HTML tables (which you can copy into Excel). Basic OCR tools may output table text as plain text, losing the grid structure.

Q: Is it safe to upload sensitive documents to online OCR tools?

For confidential documents (tax, medical, legal), prefer desktop OCR tools that process files locally on your computer. Online tools upload your file to a server — while reputable ones delete files quickly, there's always some risk. See our online PDF tools safety guide for more details.

Q: How long does OCR take?

Typically 1-5 seconds per page, depending on page complexity and your computer/server speed. A 10-page document takes 10-50 seconds. Very large or complex documents (high DPI, many images) may take longer.

Summary

OCR transforms scanned PDFs from static images into searchable, selectable, editable documents. The key points:

With OCR, your entire paper archive can become searchable and accessible. For more PDF guides, explore our complete article library.