Plain text is the most reusable form a document can take: it pastes cleanly into emails and forms, feeds into search tools and analysis scripts, and opens on anything. PDF to TXT extracts the text layer of every page into one UTF-8 text file, with pages separated by blank lines, leaving formatting and images behind.
How to use PDF to TXT
- Add the PDF. If it's a scan, run OCR PDF first.
- Select Extract Text.
- Download the TXT file.
- Open it in any text editor.
What's extracted
- The text layer, in stored order: VelloDoc reads the text layer of each page in the order it's stored in the PDF and writes it to a text file encoded as UTF-8, so accented letters and non-Latin scripts are preserved.
- Pages separated by a blank line: Pages follow one another, separated by a blank line.
- Line breaks mirror the page: Line breaks inside paragraphs usually mirror the line breaks on the page, so a long paragraph may be split across several lines that you can rejoin if needed.
Where extraction struggles
Text order depends on how the PDF was built.
- Columns, headers and footers: Multi-column layouts may interleave columns, headers and footers repeat on every page, and text inside tables appears line by line.
- Characters without a proper mapping: Some PDFs made by unusual software store characters without a proper mapping to letters, which produces garbled output even though the page looks fine. Running OCR PDF and extracting again usually solves that.
- Scans produce nothing until OCR adds a text layer.
Choosing the output format
- For editing with paragraphs and page breaks, PDF to Word is more convenient.
- For documentation: For a starting point for documentation, PDF to Markdown adds page headings.
- For tables: If you need tables in a spreadsheet, PDF to Excel splits rows into cells.
- Where to read more: The searchable PDF guide explains why some PDFs have no text to extract.
When people use PDF to TXT
Quoting from reports
Copy exact wording from a long report without fighting awkward PDF text selection.
Feeding text into analysis tools
Prepare documents for search indexes, word counts or text analysis scripts that expect plain text.
Reading with assistive technology
Open document text in a reader app or large-print editor that handles plain text better than PDF.
Limitations to know
PDF to TXT extracts existing text only, so scans need OCR first. Multi-column layouts may be reordered, repeated headers and footers are included, and all formatting, tables and images are dropped.
PDF to TXT: common questions
Why is my text file empty?
The PDF has no text layer, which is typical of scans and photos. Run OCR PDF, then extract again.
Why is the text garbled?
Some PDFs store characters without information about which letters they represent. OCR PDF recreates the text from the page image, which usually fixes this.
Is formatting kept?
No. A TXT file holds plain characters only, so bold, fonts, tables and images are left behind.
What encoding is used?
The file is saved as UTF-8, which every modern editor reads and which supports all languages.