← All guides

How to OCR a PDF Before Translating (and When to Skip It)

OCR is only necessary when your PDF has no text layer — here's the ten-second check, the free tools, and what OCR-then-translate still won't do for you.

·6 min read

You only need to OCR a PDF before translating it if the PDF has no text layer — that is, if it's a scan or a photo rather than a digitally created file. The check takes ten seconds: open the PDF and try to select a few words with your cursor. If individual words highlight, the text layer exists and most translators will accept the file as-is — skip OCR entirely. If the cursor just drags an empty box across the page, the PDF is an image, and you either run OCR first or use a translator that does its own OCR internally.

This guide covers what OCR actually does, the full set of checks for whether your file needs it, free ways to run it, and the catch nobody mentions up front: OCR-then-translate as a two-step pipeline still loses the document's layout, which is why one-pass tools exist.

What OCR actually does

OCR — optical character recognition — turns a picture of text into machine-readable text. Given a scanned page, an OCR tool detects where the writing is, recognizes each character, and outputs either a plain text file or, more usefully, a 'searchable PDF': the same page image with an invisible text layer laid underneath it, so the document looks unchanged but can now be selected, searched, and copied.

That's the entire job. OCR doesn't translate anything, and it doesn't understand the document — it recognizes shapes as characters. It's a prerequisite for translation, not a substitute for it: translation software operates on text, and OCR is what produces text from an image.

How to tell whether your PDF needs OCR

The selection test above settles it most of the time, but a couple of extra checks catch the edge cases:

  • Select text with the cursor. Words highlight individually → text layer exists. An empty selection box → image-only, needs OCR.
  • Search the document (Ctrl+F / Cmd+F) for a word you can see on the page. Found → text layer. Not found → no text layer, or a bad one.
  • Copy a paragraph and paste it into a plain text editor. Clean text → you're fine. Nothing, or garbled characters → the file needs OCR (or was OCRed badly — see below).
  • Check for mixed documents. PDFs assembled from multiple sources can have digital pages and scanned pages in the same file, so test a page from each part, not just page one.

One trap worth knowing: some scanners and copiers OCR automatically, so your PDF may already have a text layer — of unknown quality. If the copy-paste test produces text that's recognizably wrong (jumbled words, character soup), a translator will faithfully translate that garbage. In that case re-running OCR from the page images gives the translation a clean starting point.

Free ways to OCR a PDF

If you want or need OCR as its own step, you don't have to pay for it:

  • Your scanner's own software. Most scanner and multifunction-printer utilities have a 'searchable PDF' output option — turning it on at scan time is the cheapest OCR you'll ever run.
  • Free OCR web tools. Plenty of sites OCR an uploaded PDF and return a searchable one. Quality varies, and you're uploading your document to a third party — think twice before sending anything sensitive to a site you picked from a search results page.
  • Google Drive. Uploading an image PDF and opening it as a Google Doc commonly extracts the text. It's a text extraction, though — the layout is discarded in the process.
  • Desktop PDF apps you already own. If you have a paid PDF suite, its text-recognition feature produces a searchable PDF with generally solid quality.

For any OCR route, the same input rule applies: clear, flat scans of printed text recognize well, while low-resolution captures and heavy handwriting are unreliable no matter which tool runs the recognition.

The catch: OCR + translate still loses the layout

Here's what surprises people after they've dutifully OCRed their file. OCR gives you text; it doesn't give the next tool in the chain any obligation to respect the page. If you copy the recognized text into a translator, you get translated plain text — every table, column, and form field flattened. If you upload the searchable PDF to a document translator, the result is commonly re-flowed rather than reproduced: translated text is rarely the same length as the original, so lines wrap differently, tables and multi-column pages suffer, and the output needs manual cleanup to look like the document you started with.

There's a subtler failure mode too. OCR outputs text in the order it decides to read the page, and on complex layouts — two columns, a table, a form with scattered fields — that order can interleave content that a human would never read together. Translate that stream and the confusion compounds. None of this means OCR-then-translate is wrong; it means the two-step pipeline hands you the words while quietly dropping the document. If the document part matters, budget real time to rebuild it — or skip the pipeline.

The one-pass alternative: skip the separate OCR step

Tools built specifically for translating scanned documents fold OCR, translation, and re-typesetting into a single pass, which dissolves the whole question of OCRing first. Reglyph works this way: you upload the scanned PDF (or a photo of a page) exactly as it is — no preparation, no separate OCR tool. OCR reads the page internally, the original text is erased from the page image, and the translation is typeset back in its place, so tables, stamps, figures, and numbers stay where they were on the original. The output is a translated PDF that still looks like your document, with an optional bilingual side-by-side export for verifying names and figures against the source. It handles 12+ languages, runs in the browser (phones included), and the first 5 pages are free with no credit card; paid use starts at $5.

So: OCR first, or not?

A quick decision guide:

  • The PDF already passes the selection test → no OCR needed. Send it to a document translator directly.
  • It's a scan and you only need the meaning → free OCR plus any translator is fine. Accept that the layout is gone.
  • It's a scan and you need a usable document back → use a one-pass tool that OCRs, translates, and re-typesets together, and skip the manual OCR step entirely.
  • It has a text layer, but a bad one → treat it as a scan: re-OCR from the images, or hand it to a tool that does.
  • It's destined for an authority that requires certification → do the machine pass for the draft and layout, then have a certified human translator review and sign it.

Translate your scanned document now

Upload a scanned PDF or a photo — Reglyph OCRs it, translates it, and rebuilds the page so tables, stamps, and figures stay exactly where they were.

Translate 5 pages free
No credit card · Files auto-delete in 24h · Never used for training

Frequently asked

Do translation tools OCR a PDF automatically?

Some do, many don't. Tools built for digital PDFs commonly expect an existing text layer and fail or return nothing on scans, while tools built for scanned documents run OCR internally as part of the job. If a translator rejects your file or outputs a blank, a missing text layer is the usual reason.

How do I know if my PDF has already been OCRed?

Try to select and copy text, or search for a word you can see on the page. If selection and search work and the copied text reads cleanly, it has a usable text layer. If the text copies out garbled, it was OCRed poorly and is worth re-processing before translation.

Is OCR free?

It can be. Scanner software often has a searchable-PDF option, free OCR websites exist (be careful with sensitive documents), and Google Drive can extract text from an image PDF. Free routes give you the text; what they don't preserve is the document's layout.

Does OCR translate the document?

No. OCR only converts the picture of the text into machine-readable text in the original language. Translation is a separate step — which is why many people use a single tool that does both, plus the re-typesetting, in one pass.

Translate a specific language

More guides

Choose your plan

Simple, scan-friendly pricing. Pages are pages — no multiplier for scanned files.

Pay-as-you-go
Translate a few pages. No commitment.
$5/ one-time
What's included:
  • $0.50 / page
  • 10 scanned pages, never expire
  • Up to 100 MB per file
Buy pages
Most popular
Lite
For individuals with scans to translate every month.
$15/ per month
What's included:
  • $0.13 / page
  • 120 pages every month
  • Up to 100 MB per file
  • Pages reset each month
Subscribe
Pro
For high-volume users and whole-book scans.
$39/ per month
What's included:
  • $0.06 / page
  • 700 pages every month
  • Unused pages roll over (up to 2,100)
  • Big files up to 300 MB
Subscribe