Extract the text from a PDF

Read the words out of a PDF, including scanned documents where the pages are really just images. Add the file and the text comes back ready to copy.

⬆️

A photo of a page, a screenshot, or a scan

Works on scans and photos where the text is part of the picture and cannot be selected.

The two kinds of PDF

Some PDFs contain real text you can already select in any reader. Others, particularly anything that came from a scanner or a photocopier, contain nothing but a picture of each page. The second kind is why searching a scanned contract finds nothing and why you cannot copy a sentence out of it. This tool handles that case by rendering each page and reading the characters back out of the image.

Why page count matters here

Every page is rendered and then recognised separately, so the work grows in direct proportion to length. A two-page letter is quick; a two-hundred-page report is a different proposition entirely. The tool processes up to twenty pages in one run, which covers the great majority of documents people actually need to read text out of. For anything longer, split the PDF first and do the section you need.

Checking the result

Scanned text is recognised, not read, and recognition makes mistakes. The usual ones are a lowercase l becoming a 1, rn merging into m, and hyphenated words at line ends being joined oddly. Skim the output against the original before relying on it, and be especially careful with reference numbers, dates and amounts, where an error is both easy to make and easy to miss.

Frequently asked questions

My PDF already has selectable text. Do I need this?

Probably not. If you can select and copy the text in a normal PDF reader, copying it directly will be faster and perfectly accurate. This tool earns its keep on scans, photographs of documents, and PDFs exported as images.

How many pages can it handle?

Up to twenty in a single run. Beyond that the job takes long enough that splitting the document and doing the part you need is genuinely the better approach.

Why did it return nothing at all?

Usually because the pages have no readable text on them, for example a PDF of photographs. It can also happen with a very low-resolution scan where the letters have too few pixels to identify. A cleaner scan at 300 dpi will normally fix it.

More ways to use this tool