You have a PDF with a table in it and you need those numbers in a spreadsheet. The conversion runs, you open the result, and instead of tidy columns you get one long smear of text in column A — or rows that look right until you notice a figure has landed in the wrong place.
Whether that happens is decided almost entirely by one thing, and it is not the tool.
A PDF does not contain a table
This is the part that surprises people, and everything else follows from it.
A PDF stores where marks go on a page. “Draw the text Platform subscription at these coordinates. Draw 360.00 at those coordinates.” There is no row, no column, no cell, and no statement anywhere in the file that those two things belong to the same record. A table in a PDF is an appearance of a table, assembled in your head by your eyes.
So converting one to a spreadsheet is not reading data out of a structure. It is looking at the positions of several hundred pieces of text and inferring a grid that was never written down. That inference is usually right, and when it is wrong it is wrong in ways that look careless rather than impossible.
The measurement: the same data, twice
We took nine rows of charges and laid them out two ways. One with ruled lines around the cells, the way a word processor or an accounting package produces. One aligned with spaces in a monospaced font, the way a great many reports and system exports produce. Same numbers, same order, same page.
Then we put both through PDF to Excel and opened the results:
- With ruled lines: 10 rows across 5 columns. The header and all nine charges, every value in its own cell.
- With spacing only: 2 rows across 2 columns. Extraction found no grid, fell back to treating the page as text, and the entire table arrived as a single block.
That is the figure above, and it is the whole lesson. Nothing about the data changed. The only difference was whether the boundaries between columns existed as objects in the file, or only in the appearance of the page.
Why ruled lines matter so much
When a table has drawn lines, those lines are real objects in the PDF with real coordinates. Extraction can find them, build the grid from them, and assign each piece of text to the cell it falls inside. This approach is usually called lattice extraction, and it is close to reliable because it is reading something that was actually recorded.
Without lines there is nothing to read, so extraction falls back to inferring columns from the gaps between text — stream extraction. It looks for vertical channels of whitespace running down the page and calls those the column boundaries. On a clean, generously spaced table that works well. It gets into trouble when:
A value is long enough to close the gap. One supplier name that runs wider than the others can bridge the whitespace channel between two columns, and from that row down the alignment may be read differently.
Columns are tight. A narrow gutter between two number columns can be the same width as the ordinary space between words, and nothing in the file distinguishes the two.
A cell is empty. Blank cells widen the whitespace channel in some rows and not others, which can look like a column boundary that moves.
The table spans a page break. Each page is inferred independently, so the second page may be read with a different column count from the first.
What to check before you trust the result
Three checks, in the order that catches the most for the least effort.
Count the columns. If the spreadsheet has fewer columns than the table had, extraction fell back to text mode and nothing further is worth inspecting — go and get a better source file.
Re-total one column. Sum a column in the spreadsheet and compare it with the total printed in the PDF. This single check catches misaligned values, dropped rows and numbers that arrived as text, and it takes ten seconds. If the PDF prints no total, sum the first five rows by hand.
Look at the longest row. Whatever goes wrong usually goes wrong first on the row with the widest entry, because that is the row most likely to have closed a whitespace gap.
Getting a better source beats getting a better tool
If the table began life somewhere other than a PDF, the shortest route is almost always backwards rather than forwards. A report that was exported to PDF from an accounting system can usually be exported again as CSV from that same system, and no amount of extraction will beat the original.
If the PDF is all you have and it has no ruled lines, you have three options and they are all honest:
Convert it anyway and reconcile the totals, accepting that you are checking rather than trusting. Convert to Word instead if what you need is the text rather than the arithmetic — text extraction does not care about columns. Or retype it, which for a nine-row table is genuinely faster than fixing a bad conversion, and which nobody wants to hear.
And if the page is a scan
A scanned table has no text at all, so there is nothing to extract and nothing to infer from. It needs OCR first to put a recognised text layer behind the image, and only then can extraction attempt a grid.
Be realistic about that chain: OCR is a best guess at the characters, and extraction is then a best guess at the structure of those guesses. Two inferences stacked. For a table of figures that matters, a scan is the case where checking every total is not optional.
The short version
Before you convert, look at the table and ask one question: are there lines, or does it just look like there are? If there are lines, the conversion will probably be clean. If the columns are held together by spacing alone, expect to check it — and reconcile a total either way.
You can try your own file on PDF to Excel; the column count in the result will tell you within seconds which of the two cases you are in.