Editing Content

How to tell whether a PDF is scanned

Two identical-looking documents behave completely differently. Here is how to tell which one you have, in about three seconds.

Almost every confusing thing about a PDF traces back to one question: does this document contain text, or a picture of text? The two look identical and behave nothing alike.

Three tests

Three quick tests
TestReal textA scan
Try to select a line by dragging across itA selection highlight appears, following the wordsNothing selects, or the whole page selects as one block
Search for a word you can seeIt is foundNo matches
Zoom right in on a letterThe edges stay crisp at any magnificationThe letter goes soft and pixelated

The selection test is the fastest and the most reliable. Any one of the three is enough.

Verso will just tell you

Choose Make Searchable and Verso checks first. If the document already carries real text, it says so rather than doing unnecessary work. That is a fourth test, and the least effort of all.

The in-between case

Documents are not always all one or the other. A born-digital report can contain a scanned appendix. A page can be real text with a photographed table dropped into the middle of it. If search finds a word on page three and misses the same word on page eleven, that is what is happening.

Running recognition over the whole document is harmless in that case: the pages that already have text keep it, and the pages that do not gain it.

Why it matters before you start work

Knowing which kind of document you have saves you from a series of small confusions:

  • Text editing appears not to work, because there is nothing to click into
  • Highlighting and underlining do nothing, because there is nothing to select
  • Search finds nothing, so you conclude the document does not mention what it clearly mentions
  • Read Aloud stays silent
  • Converting to Word produces a document containing one large picture
  • Redaction has no text to remove — and this one matters, because you may believe a document is safe when it is not

All six are the same problem, and one run of recognition fixes all of them at once.

Frequently asked questions

How do I know if my PDF is scanned?

Try to select a line. If a highlight follows the words, it has real text. If nothing selects, or the whole page selects as one block, it is a scan.

Why does searching my PDF find nothing?

Because there is nothing to search: the page is a picture. Run text recognition over it and search will work.

Can a PDF be partly scanned?

Yes, and it is common — a born-digital report with a scanned appendix, or a page with a photographed table in it. Running recognition over the whole document is harmless: pages that already have text keep it.

Does a scanned PDF look different?

Usually not on screen. Zooming right in is the giveaway: scanned letters go soft, while real text stays crisp at any magnification.

Why does this matter for redaction?

Because redaction removes text, and a scan has none. Knowing which kind of document you are working with is part of knowing whether a redacted copy is actually safe to send.