What a PDF is, and why it behaves the way it does
A plain explanation of the format, where it came from, and the one distinction that explains most of the confusion people have with it.
PDF stands for Portable Document Format. The name is unusually honest: the whole point of the format is that a document travels from one machine to another and still looks like itself when it arrives.
The problem it was invented to solve
In the late 1980s, sending someone a document was a gamble. The fonts you used might not exist on their machine. Their word processor might reflow your careful layout into something unrecognisable. Printing it on a different printer could move every line. A page that looked right on your screen had no obligation to look right on anyone else's.
The answer, first released in 1993, was to stop describing a document as content to be laid out and start describing it as a page that has already been laid out. A PDF says where every mark goes, carries the fonts it needs along with it, and leaves nothing for the receiving machine to decide. That is why a PDF looks the same on a phone, a laptop, a plotter and a print shop's press.
The format was published as an open standard in 2008 and is now maintained as ISO 32000, which is why it outlived the company politics of the 1990s and became the way the world exchanges documents that must not move.
What a page actually contains
A PDF page holds a set of instructions for drawing: put this glyph here at this size, fill this rectangle, place this image in this box. Because the instructions are absolute, the page is fixed. Because they describe marks rather than meaning, a PDF has no idea that a run of glyphs is a heading, or a paragraph, or a name.
That single fact explains most of what people find surprising about PDFs:
| What people expect | What the format actually does |
|---|---|
| Text reflows when the window narrows | Nothing reflows. The page is a fixed canvas, so a narrow window means smaller text or scrolling. |
| Editing a line is like editing a document | There are no paragraphs to push around, so an editor has to re-lay the line and fit it back into the space it had. |
| Everything on a page can be searched | Only if the page carries text. A page can equally carry a picture of text, which looks identical and contains no words at all. |
| A black box hides what is under it | A rectangle is just another mark. The text underneath is still there unless something removes it. |
| The file is one flat picture | It is usually a structured document: real text, real vector shapes, real images, plus information about itself. |
The distinction that explains almost everything
There are two kinds of PDF, and they look identical on screen.
A born-digital PDF was produced by a program — a word processor, a design tool, an accounting package. Its pages carry real text. You can select a sentence, search for a word, copy a paragraph and edit a line.
A scanned PDF was produced by a camera or a scanner. Each page is a photograph, wrapped in the PDF format. It looks like a document and behaves like a picture: nothing to select, nothing to search, nothing to click into.
Text recognition is the bridge. Running Make Searchable in Verso reads the words in the picture and adds them to the page as real text, so a scan becomes selectable, searchable and editable. It happens on your Mac rather than on someone's server, which matters when the scan is a medical record or a contract.
What else a PDF carries
Besides the marks on the page, a PDF can carry a surprising amount:
- Its own details — a title, an author, keywords, and the dates it was created and changed. This travels with the file, which is why clearing it before sharing is worth doing.
- An outline — the table of contents a long report was published with, which good readers turn into a navigable list.
- Annotations — highlights, notes, drawings and stamps, stored as objects on top of the page rather than burned into it. That is why they can be moved and removed, and why flattening exists for when you need them not to be.
- Form fields — real, fillable boxes rather than lines drawn to look like boxes.
- Links — regions of a page that go somewhere when clicked.
- Encryption — a password requirement, and restrictions on printing or copying.
What a PDF is good at, and what it is not
A PDF is the right format when the layout is part of the meaning: a contract, an invoice, a plan, a scientific paper, a form, anything that will be printed, anything that must look identical to everyone who opens it, anything you need to sign.
It is the wrong format when the content will keep changing and the layout does not matter. A document still being drafted belongs in a word processor; a table still being calculated belongs in a spreadsheet. Converting out of PDF is possible and often useful, but a conversion is a translation, and some of the original layout will always be interpreted rather than reproduced.
Why the format has lasted
Thirty years is a long time in software. PDF has lasted because it solved its problem completely rather than partly: a page that cannot move cannot be broken by the next machine, the next operating system or the next decade. Formats that promised more flexibility have come and gone; the one that promised a page would stay put is still how the world sends documents that matter.
Frequently asked questions
What does PDF stand for?
Portable Document Format. The name describes the goal: a document that travels to another machine and still looks exactly like itself.
Who invented the PDF?
It was created at Adobe, with the first version released in 1993. The format was published as an open standard in 2008 and is now maintained as ISO 32000, which is why anyone can build software that reads and writes it.
Why does a PDF look the same on every device?
Because the page is already laid out. A PDF describes where every mark goes and carries the fonts it needs, so the machine opening it has nothing left to decide.
Why can I not select the text in some PDFs?
Because those pages are pictures of text rather than text. A scanned document looks identical on screen and contains no words at all until text recognition adds them.
Can a PDF be edited?
Yes. Verso edits the actual text on a page rather than pasting a box over it, and can reorder, rotate, merge and split pages. A scanned page has to be recognised first, because there is nothing to click into until it is.
Is a PDF a picture?
Usually not. Most PDFs are structured documents holding real text, vector shapes and images. Scanned PDFs are the exception: each of their pages is a single photograph.
Does a PDF contain information about me?
It can. A PDF carries its own title, author and keywords, plus creation and modification dates, and those travel with the file. Clearing them before you share a document is a sensible habit.
Why does a black box over a name not actually hide it?
Because a rectangle is just another mark on the page, and the text underneath is untouched — anyone can select it and copy it straight back out. Real redaction removes the content rather than covering it.