What PDF metadata is
PDFs store metadata in two places.
- The document information dictionary holds a handful of standard fields: title, author, subject, keywords, the application that created the original document, the software that produced the PDF, and the creation and modification dates.
- The XMP stream: An XMP stream (a block of XML data inside the file) often repeats those details and can add more. That can include document IDs, the names of editing tools and history notes written by the software that touched the file.
- Almost none of it is typed by a person: Almost none of this is typed by a person. Word processors fill in the author from the account name, scanners and printers add their model, and templates carry their original titles into every document made from them.
What metadata can reveal
- The author field often has a full name or a computer username, which matters when a document should be anonymous, for example in blind peer review or a whistleblowing report.
- Company and software names: A company name, a department or a client name can appear in the title or subject of a template. Software versions show which tools an organization uses.
- Dates that contradict a story: Dates can also contradict a story. A report described as new may carry a creation date from years ago, and a modified date can show that a "final" document was changed later.
- Keywords sometimes preserve internal project names.
- Invisible on the page: None of this is visible when the recipient simply reads the pages, but anyone who opens the document properties can see it.
Check a PDF's metadata first
- In your viewer: Most PDF viewers show the main fields under a menu item like Document Properties.
- With View / Edit Metadata: You can also add the file to View / Edit Metadata, which reads the title, author, subject and keywords as soon as the PDF is loaded and lets you change them.
- XMP is usually hidden: Viewers rarely show the XMP stream, so a document that looks clean in its properties can still carry details there.
- When in doubt, remove everything: If you aren't sure what a file contains, removing all metadata is the safer choice.
Removing metadata with VelloDoc
- What the tool clears: Remove Metadata clears the document information dictionary and deletes the XMP stream, then saves a clean copy. The pages themselves, including all of their text and images, stay exactly as they were.
- How to run it: Add the PDF, enter its password if it's protected, select Remove Metadata and download the result.
- Adding a title back: If the published file should still have a helpful title, open the clean copy in View / Edit Metadata afterwards. Fill in only the fields you want to share. A good title helps search engines and screen readers present the file properly.
- Check before sending: Check the result in your viewer's document properties before sending it.
What metadata removal doesn't remove
Metadata is only one layer of a PDF.
- Comments, forms, attachments and bookmarks: Comments can include their authors' names, and form fields keep whatever was typed into them. Attached files travel inside the PDF. Bookmarks can have revealing labels, and hidden layers can hold content you don't see.
- Camera data in embedded photos: Photos embedded in the pages may still contain camera details, including location, in their own image data.
- The page content itself: And the page content itself, from names in the text to text hidden in white or under an image, is untouched.
- Redact what is on the page: For content on the page, use Redact PDF, which deletes what's inside the marked areas.
- Flatten for a visible-only copy: If you need a copy with only what's visible, Flatten PDF rebuilds every page as a new image. That leaves out comments, attachments, hidden text and the original photos. The catch is that you can't select the text anymore.
- Rename the file too: The file name is separate from all of this, so rename the file too.
A pre-sharing checklist
Work from content to properties.
- Send only what's needed: First decide what the recipient actually needs and remove everything else with Remove Pages or Extract Pages.
- Redact, then flatten: Redact sensitive details inside the remaining pages, and flatten the document if comments, form data or hidden content could be a problem.
- Remove metadata last: Remove metadata as the last editing step, so no earlier step can leave details behind.
- Add a password after that: If the file needs a password, add it with Protect PDF after the metadata is gone.
- Rename and check: Finally, give the file a neutral name and open it once more to confirm the pages and properties look the way you expect.
- Where to read more: The guide to what a PDF can reveal covers each of these risks in more depth.
Questions
Does removing metadata change the pages?
No. Remove Metadata clears the document properties and the XMP stream. The text and images on the pages stay exactly as they were.
Will the file name still show my name?
Yes, if your name is in it. The file name isn't part of the PDF's metadata, so rename the file before sending it.
Are comments and annotations removed?
No. Annotations can include their authors' names. Delete them in a PDF editor, or use Flatten PDF to create a copy that has only page images.
Does it remove location data from photos inside the PDF?
No. Remove Metadata doesn't change embedded images. Flatten PDF rebuilds each page as a new image, which leaves out the original photos and their camera data.
Should I remove metadata before or after adding a password?
Before. Remove the metadata first, then protect the clean copy with Protect PDF.