PDF to text guide
How to convert a scanned PDF to text with OCR
A scanned PDF contains page images, while PDF to Text reads text objects already stored in the file. Converting a scan therefore requires a recognition step first.

Stage 1: recognize the scanned pages
Open OCR PDF with the scanned document. The current recognizer is intended for printed English and accepts up to 25 scanned pages in one OCR job. Download the searchable result, then test several words and manually compare names, dates, and numbers with the original images.
If the source is blurred, skewed, dark, or too small, improve the page images before treating the recognized text as usable.
Stage 2: extract the recognized text
Open the reviewed OCR result in PDF to Text. Choose whether approximate spacing and page headings should be included, inspect the preview, and download the UTF-8 TXT file.
PDF to Text reads the recognized text layer created in the first stage. It does not return the page images or preserve the original visual layout.
Clean the TXT without losing meaning
Line wrapping, columns, tables, headers, and footers can create awkward reading order. Compare each corrected passage with the scan before removing line breaks or repeated text. Keep the searchable PDF beside the TXT so page context remains available.
FAQ
Why does PDF to Text return little or nothing from a scan?
The pages may contain only images. Run OCR PDF first, review the recognized layer, and then extract the text.
Can I send a scanned PDF directly to PDF to Text?
You can open it, but PDF to Text does not recognize image-only words. The OCR stage is required when no selectable text exists.
Will the TXT preserve the scanned page layout?
No. TXT stores plain text, so columns, tables, spacing, and page design may need manual cleanup.