Convert a Scanned PDF to Excel

A scan holds a picture of a table, not a table. Tick OCR and the characters are recognised on your own device — the document is never uploaded.

Quick answer:

Open the PDF table extractor, tick Read scanned pages (OCR) and press Extract tables. A scanned PDF holds a photograph of each page rather than any characters, which is why it otherwise converts to an empty sheet. The columns are the robust part: all five columns of a five-column statement came back at every quality we tested, down to 56 dpi — while cells exactly right fell from 100% at 150 dpi to 85%. The recognition runs in your browser, so the document never leaves your device.

Why did my scanned PDF convert to an empty sheet?
Because it holds a photograph of each page rather than any characters. Try selecting a word in a reader: if nothing highlights, tick Read scanned pages (OCR) and run it again.
How do I know if my PDF is scanned?
Select a word with the mouse, or search the document for a word you can see. A scan highlights nothing and finds nothing.
Will the columns survive a poor scan?
Yes. All five columns of a five-column statement came back at 150, 100, 75 and 56 dpi, and at 1.5° of skew. It is the individual characters that suffer, not the layout.
What scanner settings should I use?
300 dpi, greyscale or black and white, de-skew on, and not the smallest-file-size mode. Flatten the page so the crease does not throw a shadow through a column.
Is my scanned document uploaded?
No. Rendering and recognition both run in your browser. The only network request is the ~6 MB recognition engine, fetched once and then cached.
📄 Open the PDF table extractor

OCR runs on your device. No account, no upload, no page limit.

Why your scanned PDF converted to an empty spreadsheet

Because there was nothing in it to put in cells. This is the single most common surprise with these documents and it is worth understanding before you try another tool.

A PDF made by software — exported from Word, generated by a billing system, printed to PDF — stores characters, each with a position on the page. A converter reads those characters and works out which ones share a row and which share a column. A PDF made by a scanner or a copier stores one photograph per page. There are no characters in it at all. Open it in a reader and try to select a word: if nothing highlights, that is what you have.

So a converter with nothing to read has two honest options: produce an empty sheet, or say so. This one says so, and points you at the tick that fixes it — Read scanned pages (OCR). With that on, each page is rendered to an image, the characters are recognised, and the recognised words are then treated exactly like the characters a normal PDF would have supplied.

The structure survives what the characters do not

This is the finding that matters for a spreadsheet, and it runs the opposite way to what people expect. Column layout is robust. Individual digits are fragile.

We rendered the same five-column statement at falling quality and measured what came back:

Scan qualityColumns recoveredCells exactly right
150 dpi, light grain5 of 5100%
100 dpi5 of 591%
75 dpi5 of 591%
56 dpi, heavy grain5 of 585%
150 dpi, heavy grain and hard JPEG compression5 of 597%
150 dpi, rotated 0.5°5 of 5100%
150 dpi, rotated 1.5°5 of 591%

All five columns came back every single time, including from a 56 dpi scan so grainy it is unpleasant to read. The cells are where quality shows: 150 dpi gives you a table you can trust, 56 dpi gives you a table you have to check.

On a denser page the same story from the other side. A 30-row statement at 150 dpi gave 60 of 60 money figures exactly right; a rough 82 dpi version of the same page gave 53 of 60. Seven wrong figures in a statement is not a rounding error, it is a wrong answer — so on a poor scan, check the numbers.

Getting a better scan is cheaper than fixing a bad one

Everything above is decided before the file reaches any converter. If you are the one doing the scanning:

If the scan already exists and cannot be redone, run it anyway — the columns will be right, and correcting a handful of figures in the preview is quicker than it sounds.

Generations: why a photocopy of a fax is a different problem

Documents in circulation are rarely first-generation. A supplier prints an invoice, faxes it, the recipient photocopies it for a file, and somebody eventually scans that copy to send onward. Each of those steps thickens the strokes slightly and adds speckle, and the damage compounds — by the third generation the counter inside an e has closed up and a recogniser reads it as a c or an o.

Fax is the worst of the chain by a wide margin. Standard fax resolution is roughly 200 by 100 dots per inch — deliberately squashed vertically, because the standard was designed for telephone lines rather than legibility. Characters arrive taller than they are detailed, and a decimal point is a single dot that may not have survived the journey. If a faxed document is all you have, expect the columns and verify every figure.

Two practical consequences. First, always chase the earliest generation you can reach — asking the sender for the original PDF takes a minute and removes the problem entirely, because a software-generated PDF needs no recognition at all. Second, if you must re-copy a document before scanning it, raise the contrast on the copier rather than lowering it; pale, broken characters are far harder to recover than slightly heavy ones.

Two things a scan loses that a normal PDF keeps

The ruling lines

When a PDF is generated by software, a bordered table carries its own geometry: the rules are vector paths with exact coordinates, and reading them gives a grid that matches the printed one exactly. A scan has no paths — the borders are dark pixels in a picture. So a ruled table that has been scanned is read the same way an unruled one is, by looking at the gaps between words. That works well, and it is worth knowing that a scan gives up the more precise of the two methods.

The page's own structure across pages

Grouping several pages into one continuous table relies on comparing headings and column positions between pages, and both come back slightly noisier from a scan. The grouping still works — over ten multi-page fixtures the rule got 10 of 10 right with nothing wrongly merged — but a scan is where you should glance at the preview before trusting the join.

Mixed documents: a scanned insert among typed pages

Very common in practice — a contract exported from software with a signed page photographed and dropped back in, or a report with one page of appendix that was faxed. Whether a page needs recognising is decided per page, not per document. Pages that carry real text are read directly, which is both faster and exact; only the pages that need it are rendered and recognised. The preview and the page list mark which pages went through OCR, so you know which ones to check.

That also means you never have to split the file first, and you never pay the recognition cost on pages that did not need it.

Checking the result, and what gets marked for you

A cell that cannot be what its column is gets a ⚠. The classic case is a zero read as a capital O, giving 12O0.00 in a column of money — not a hesitant reading, an impossible one. On deliberately corrupted figures that check flagged 6 of 6 with no false alarms, and it raised nothing at all on 6 correctly extracted typed pages.

Its limit is stated plainly because it matters: it catches roughly one OCR mistake in five. Most of the rest are in description text, where there is nothing to check a word against — SAINSBURVS for SAINSBURYS is wrong and looks like a perfectly ordinary word. A page with no marks is not a page with no errors.

Double-click any cell in the preview to correct it. The correction goes into the file you download, and the extraction underneath is left alone, so changing the format or the sheet settings afterwards keeps your edits.

Where to go next

Frequently Asked Questions

Open the PDF table extractor, tick Read scanned pages (OCR) and press Extract tables. The page is rendered, the characters are recognised in your browser, and the recognised words are laid out into columns the same way a normal PDF's text would be. Your document is never uploaded — only the recognition engine is downloaded, about 6 MB the first time.
Because a scanned PDF contains a photograph of each page rather than any characters, so there was nothing to put in cells. Open it in a reader and try to select a word: if nothing highlights, that is what you have. Tick Read scanned pages (OCR) and run it again. A converter that hands back a blank sheet without saying this is the reason is not helping you.
The columns are far more reliable than the characters. All five columns of a five-column statement came back at every quality we tested — 150, 100, 75 and 56 dpi, and at 1.5 degrees of skew. Cells exactly right fell from 100% at 150 dpi to 85% at 56 dpi. On a dense 30-row page, 60 of 60 money figures were exact at 150 dpi and 53 of 60 on a rough 82 dpi version.
300 dpi, greyscale or black and white, de-skew on, and not the smallest-file-size compression mode. 150 dpi is the comfortable floor. Flatten the page so the crease does not throw a shadow through a column, because a shadow is read as ink.
That works without splitting the file. Whether a page needs recognising is decided per page: pages with real text are read directly, which is faster and exact, and only the pages that need it are rendered and recognised. The preview marks which pages went through OCR so you know which to check.
Not as geometry. In a software-generated PDF the borders are vector paths with exact coordinates and the grid can be read straight from them. In a scan the borders are just dark pixels, so the columns are worked out from the gaps between words instead. That still works well — it is simply the less precise of the two methods, and a scan gives up the better one.
No. The rendering and the recognition both run in your browser on your own device. The only network request is for the recognition engine itself, about 6 MB the first time, which your browser then caches. Adobe, Smallpdf and the other OCR converters for scanned PDFs upload your file to their servers to do this.
Open it in any reader and try to select a word with the mouse. If text highlights, the PDF holds real characters and needs no recognition. If nothing highlights, or the whole page selects as one block, it is an image and you need the OCR tick. Another tell: search the document for a word you can see on the page — a scan finds nothing.
Yes, and expect the worst results of anything here. Standard fax resolution is roughly 200 by 100 dots per inch — deliberately squashed vertically, because the format was designed for telephone lines rather than legibility — so characters arrive tall and short of detail, and a decimal point is a single dot that may not have survived. Take the columns, then verify every figure against the original.
Yes, and the number of generations matters more than the quality of any single copy. Each print, fax and photocopy thickens the strokes and adds speckle, and the damage compounds until the counter inside an e closes up and reads as a c. Chase the earliest generation you can reach — asking the sender for the original PDF removes the problem entirely, since a software-generated PDF needs no recognition at all.
Because it is running on your device rather than on a rack of servers, and it is seconds per page. A long document is a genuine wait. The Extract button becomes a Cancel while it runs and the pages already read are kept, so you can stop and take what you have. Extracting a page range instead of the whole document is the quickest way to shorten it.
Yes — recognition and output format are independent choices, so tick Read scanned pages and set Format to CSV. One catch that is not about scanning: a CSV has no concept of sheets, so if the document holds several tables they all land in one file in one grid. Choose .xlsx when there is more than one.
Coloured paper is usually fine. A highlighter over figures is not — it lowers the contrast between ink and background in exactly the region you most want read, and a marker that has bled through from the other side of the sheet is worse again. Scan in greyscale rather than colour and raise the contrast if your scanner offers it.

Related Tools