VelloDoc
Convert from PDF

Getting content out of a PDF: text, data, images and archives

A PDF is built to be read, not reused. It records where each character and image sits on a page, but not the paragraphs, tables and styles of the document it came from. Getting content back out is therefore a question of what you actually need: the words, the numbers, the pictures, a web version or a long-term archive. This guide explains what each conversion can recover and what it can't. You can choose the route that gives the most useful result with the least cleanup.

Convert from PDF tools

Tools in this area: 12. Each one links to a detailed page with its settings, limits and common questions.

PDF to JPG

Render every page of a PDF as a JPG image at 120, 150 or 200 DPI, downloaded together in a ZIP.

Convert from PDF

PDF to PNG

Render PDF pages as lossless PNG images with crisp text and lines, ideal for documentation and diagrams.

Convert from PDF

PDF to WebP

Turn PDF pages into compact WebP images for website previews, online menus and in-app catalogs.

Convert from PDF

PDF to Word

Recover the text of a PDF as an editable DOCX with paragraphs and page breaks, ready to restyle.

Convert from PDF

PDF to Excel

Pull lines of text from PDF pages into an XLSX workbook, splitting columns where the text has wide gaps.

Convert from PDF

PDF to PowerPoint

Place each PDF page on its own slide as a high-resolution image, sized to match the shape of the page.

Convert from PDF

PDF to TXT

Extract all readable text from a PDF into a plain UTF-8 text file for quoting, search or analysis.

Convert from PDF

PDF to Markdown

Extract PDF text into a Markdown file with a heading for each page, a starting point for wikis and notes.

Convert from PDF

PDF to HTML

Convert a PDF into one self-contained HTML page with real, reflowing text and embedded images.

Convert from PDF

Extract Images

Save the photos and graphics embedded in a PDF in their original format and resolution, without screenshots.

Convert from PDF

PDF to PDF/A

Rewrite a PDF as a self-contained PDF/A-2 archival file with embedded fonts and an sRGB color profile.

Convert from PDF

AI PDF Workspace

Search and ask questions across several PDFs at once. Compare dates, amounts and clauses, spot conflicts and get one summary.

Convert from PDF

Why PDFs are hard to convert back

Inside most PDFs, text is stored as characters placed at exact positions on a page. There's often no record that a group of lines forms a paragraph, that certain words are a heading, or that a grid of numbers is a table with rows and columns. A converter has to rebuild that structure from positions, which works well for simple documents and poorly for complex layouts. Scanned PDFs are harder still, because they contain no characters at all, only pictures of pages. That's why the most reliable path is always the original source file when it exists, and why VelloDoc's converters favor clean, honest output over imitations of the original layout. The format guide explains the underlying trade-offs.

Recovering the words

Three converters recover text for different destinations. PDF to Word produces an editable DOCX with the text of each page as paragraphs and page breaks preserved, ready to restyle, but it doesn't rebuild fonts, tables, images or columns. PDF to TXT writes all text to a UTF-8 file, the most portable form for quoting, searching and analysis. PDF to Markdown adds a heading for each page, a convenient starting point for wikis and note apps. If the output looks garbled even though the page reads normally, the PDF stores characters without a proper mapping to letters. Running OCR and converting again usually fixes it.

Recovering numbers and tables

PDF to Excel creates one worksheet per page and turns each line of text into a row. It splits cells wherever the text has a tab or a wide gap, since that's how many PDFs lay out columns. Simple, evenly spaced tables come across well. Merged cells, wrapped text and empty cells need tidying, because spacing is the only clue to structure. Values arrive as text, so convert them to numbers before calculating. If the data first came from a system export, asking for the CSV file is always more accurate than pulling it out of a PDF. And when you need a printable table, CSV to PDF does the reverse.

Pictures of pages or the images inside

Two different jobs are easy to confuse. Turning pages into images gives you a picture of each whole page. PDF to JPG, PDF to PNG and PDF to WebP convert every page at 120, 150 or 200 DPI and give you a ZIP. JPG suits photos, PNG suits text and diagrams, and WebP suits websites. Extract Images instead pulls out the photos and graphics stored inside the PDF, in their original format and resolution. It can be far better quality than a page render. For presenting pages, PDF to PowerPoint places each page on a slide sized to the page shape, as an image you can annotate but not edit.

A PDF as a web page

Long PDFs are awkward on phones and opaque to search engines. PDF to HTML turns a document into a single web page file. Pages with a text layer become real text that fits the screen and can be selected, searched and read aloud, with their images included. Scanned pages are added as page images, so nothing is lost. Exact layout, columns and fonts are simplified in favor of readability. For documentation sites that use Markdown, the Markdown converter is the better fit, and publishing the original PDF alongside the HTML keeps a printable version available for people who want one.

Scans need OCR first

Every text-based conversion depends on a text layer. A scanned or photographed PDF has none, so PDF to Word returns empty pages, PDF to Excel returns empty sheets and PDF to HTML falls back to page images. OCR PDF adds recognized text in any of 39 languages, after which every text converter works normally. Recognition isn't perfect, especially with small print, handwriting and complex layouts, so check names, numbers and dates in the converted result. Cleaning scans before recognition, with the optimization tools for straightening and blank-page removal, measurably improves accuracy, and what makes a PDF searchable explains why.

Answers from many PDFs at once

Sometimes you don't need to convert anything. You need an answer that's spread across several files. The AI PDF Workspace reads the text of up to 20 PDFs so you can search all of them at once, right in your browser. It can also ask Claude, an AI model made by Anthropic, to compare the documents, list where they disagree or write one summary, with a file name and page number for each point. Search is free and never leaves your browser. AI answers send the text to Anthropic only when you ask and agree, so check each page reference before you rely on it.

Archiving instead of converting

Sometimes the goal isn't to reuse content but to keep the document readable for decades. PDF to PDF/A rewrites a file to the PDF/A-2 archival standard with Ghostscript, embedding fonts and an sRGB color profile so the file doesn't depend on anything outside itself. No converter can guarantee that every input passes strict validation, so check the result with a validator like veraPDF when compliance is mandatory. Repair damaged files first with Repair PDF, give the document a meaningful title with View / Edit Metadata, and make scans searchable before archiving so they stay findable.

Common workflows

Reuse text from a scanned report

Add a text layer to a scanned report, then recover its paragraphs as an editable Word document that you can quote and restyle.

Move a price list into a spreadsheet

Make a scanned or printed price list searchable, then pull its rows into a workbook where you can clean and compare prices.

Check a contract against its amendments

Make scanned amendments searchable, then ask which dates, amounts and terms apply now, with a page reference for each answer.

Publish a report as web content

Turn a report into a reflowing web page, save its original photos for your site, and create page previews for the download link.

Archive records properly

Rebuild a damaged file, give it a clear title and author, and convert it to PDF/A for long-term retention.

Questions about converting PDFs to other formats

Can a PDF be converted back into an identical Word file?

Rarely. PDFs record positions, not document structure, so converters recover text reliably but not the exact layout. Use the original source file whenever you have it.

Why do my conversions come out empty?

The PDF is almost certainly a scan with no text layer. Run OCR PDF first, then convert the result.

What's the difference between PDF to PNG and Extract Images?

PDF to PNG renders whole pages as pictures. Extract Images saves the individual photos and graphics stored inside the PDF, in their original format and resolution.

Which conversion is best for reusing tables?

PDF to Excel gives a head start on simple tables. For accurate data, ask for the original CSV or spreadsheet the PDF was made from.

Does converting change my original PDF?

No. Each conversion creates a new file, and the temporary copy on the server is deleted after your download finishes.

Related guides