PDF text

PDFTXT

Extract text from PDF

A PDF can hold a real text layer, or it can be a photograph of a page. This conversion pulls the text that is already there into a .txt file you can search, diff or load into another tool. It does not invent characters from pixels.

No account on the public uploader. Guest jobs stay within five conversions and 250 MB per day. Business rates start at €0.80/GB for OST → PST.

Talk to the AI — it replies
IACOPO

Describe what you want, then continue to convert.

Extraction is not OCR

If you can select text in a PDF reader, there is something to extract. If the file is a scan, PDF to TXT will not magically type it. That honesty is the point of this page — consumer tools often blur the two.

Archival PDF/A is a different conversion: it changes the container profile, it does not emit a .txt.

Layout versus stream

PDF text extraction is not a perfect page layout. Columns, headers and footnotes can reorder. You get a text dump, not InDesign.

Validation

We check that the output is text from a PDF, not an empty file from a corrupt or encrypted source. Encrypted PDFs fail with a diagnostic.

Questions

Scanned invoices?

Without a text layer, extraction stays empty or useless. Use a dedicated OCR workflow elsewhere; this page will not pretend.

Keep the PDF as well?

This job returns text. If you still need an archival PDF, run PDF to PDF/A as a separate conversion.

Password-protected PDFs?

We do not bypass encryption. Unlock the file with the password you own, then upload.