← All guides

Guide

PDF to Word: What Survives and What Doesn't

Converting a PDF to Word works well or not at all, and which one you get is decided before you start — by whether your PDF contains text or a photograph of text. This explains how to tell, what a conversion can honestly recover, and why the tools that promise perfect layout are the ones to be most careful with.

There are two kinds of PDF, and only one converts

A PDF made from a Word file, a web page or an export has a text layer: the characters are stored as characters. Select a sentence in a reader and it highlights word by word — that is a text PDF, and it converts.

A PDF made by a scanner or a phone camera is a photograph of a page. There are no characters in it at all, just an image that happens to look like writing. Try to select text and you get nothing, or you get a rectangle over the whole page.

That is the whole diagnosis and it takes two seconds. If you cannot select individual words, no converter can extract them, and a converter that returns an empty document is telling you the truth about your file rather than failing.

Why layout cannot come across

A PDF does not describe a document the way Word does. It describes a page: this glyph at these coordinates, in this font, at this size. There is no paragraph, no table, no column — only marks in positions that happen to look like those things to a human eye.

Word is the opposite. It stores structure and works out positions afterwards, which is what lets text reflow when you change a margin.

So converting is not translation, it is reconstruction. The text and its reading order come across because that information is genuinely there. Everything else — which marks form a table, where a column begins, which line is a heading — has to be inferred from spacing, and inference is guesswork with a confident face.

The failure that is hardest to catch

This is the real argument for expecting less from a converter.

When a tool guesses wrong about a table, it does not produce something obviously broken. It produces a table — neatly formatted, plausible, with a number in the wrong column. When it guesses wrong about a two-column layout, it interleaves the columns into paragraphs that read almost sensibly.

You can proofread a document that came out as plain paragraphs, because the damage is visible and you know what you lost. You cannot easily proofread a document that came out looking right. For anything with figures in it — an invoice, a statement, a bill of quantities — the confident reconstruction is the more dangerous output.

That is why extracting text and paragraphs, and saying plainly that the rest did not survive, is the safer behaviour rather than a limitation.

Getting the best result

Try to select a word in the PDF first. If you can, extract the text and accept that you are rebuilding the formatting yourself — which is usually faster than repairing a bad reconstruction anyway.

If you cannot select, you have a scan. Decide whether you need the text badly enough to OCR it and proofread it, or whether retyping the part you actually need is quicker. For a single table, retyping usually wins.

And think about where the file goes. The documents people convert are contracts, CVs, payslips and bank statements — the exact category you would not email to a stranger. A converter that runs in your browser never has the file to lose.

Common questions

Why is my converted Word file empty?

Your PDF is almost certainly a scan — a photograph of text with no text layer underneath. There is nothing to extract. Try selecting a single word in the PDF: if you cannot, no converter can either without OCR.

PDF to Word
Will my tables and images come across?

No. A PDF stores glyph positions, not structure, so tables and columns have to be inferred from spacing. Text and paragraph breaks survive because that information genuinely exists; the rest is reconstruction.

Some converters promise to keep the layout. Are they lying?

Not exactly — they attempt the reconstruction. The problem is how it fails: a wrong guess produces a neat, plausible table with a value in the wrong column, which is far harder to catch than plain paragraphs. For documents with figures, be sceptical of confident output.

How do I convert a scanned PDF?

You need OCR, which recognises characters from the image. Expect systematic errors between 0 and O, 1 and l, 5 and S, and check any figures against the original. For one table, retyping is often faster and safer.

Is it safe to convert a contract or a payslip online?

It depends whether the file leaves your device. Most converters upload, which puts your document on someone's server under their retention policy. One that runs in the browser never receives it, and you can confirm that in the Network tab.

Compress a PDF without uploading it

Tools in this guide

Share
WhatsAppXin