Accounting systems print a whole month of invoices into one PDF. So do payroll runs, bank statements and dispatch notes. The documents inside are not the same length, so splitting every three pages goes wrong on the first four-page invoice, and typing page ranges for two hundred documents is not work anyone should do. But the file already knows where the divisions are: the invoice number changes, or the word Invoice appears again. This tool reads that and cuts there.
How to use Split by Text
- Add the PDF holding the batch.
- Choose how a new document is recognised: by the text in an area of the page changing, or by a page containing certain words.
- For the area, give its position as percentages of the page. The default is the top third, which covers most reference numbers and letterheads.
- For words, type the phrase that opens each document, such as Invoice number or Statement of account.
- Select Split by Text and download the ZIP. Each document is a separate PDF, named after the text that started it.
The two ways of finding a boundary
Both do the same job from opposite directions, and which fits depends on whether your documents have a reference number or a heading.
- When the text in an area changes: The chosen part of every page is read, and a new document starts whenever what it says is different from the page before. This is the one to use when each invoice carries its own number in the same place.
- When a page contains certain words: Every page is searched for a phrase, and each page that contains it starts a new document. This suits batches where the first page of each document says something the others do not.
- Continuation pages stay put: In the area mode, a page where the area is empty is treated as a continuation of the document before it. That is exactly how a three-page invoice behaves, since only page one carries the number.
- Anything before the first boundary: Pages ahead of the first match are not dropped. They are kept together as the first part, so a covering letter at the front of a batch is not lost.
Setting the area
The area is given in percentages rather than millimetres, so the same four numbers keep working on A4, Letter and next month's batch alike.
- Start with the default: Left 0, top 0, width 100, height 33 reads the top third of every page. Reference numbers and letterheads live there, and it is the right answer more often than not.
- Narrowing to a corner: If the top third also holds a date that changes on every page, narrow the area. Left 60, top 0, width 40, height 15 reads only the top-right corner, where invoice numbers usually sit.
- Too wide and every page splits: An area that includes body text or a page number changes on every page, so every page becomes its own document. If that happens, make the area smaller.
- Too narrow and nothing splits: An area with no text in it cannot find a boundary, and the tool says so rather than returning the file unchanged. Widen it and try again.
What comes back
- One file per document: Each part is a separate PDF holding the pages from one boundary to the next, so a two-page invoice is two pages and a five-page invoice is five.
- Named after what split it: In the area mode, each file is named after the text that was found, so invoices-02-INV-1002.pdf can be filed without opening it. In the words mode, files are named by the page they start on.
- Delivered as a ZIP: A batch split produces many files, and one download is far more reliable than a burst of separate ones. A batch that turns out to hold only one document comes back as a plain PDF instead.
- Pages are copied whole: Nothing is re-rendered. Text, images and vector graphics are carried across as they were, so the parts are as sharp and as searchable as the batch they came from.
When people use Split by Text
A month of invoices in one export
Split on the invoice number and every customer's invoice becomes its own file, named after the number, ready to attach to the right email.
Payslips for a whole payroll run
Each payslip starts with the employee's name or reference in the same corner. Splitting there produces one file per person in a single pass.
Bank statements by account
A combined export holding several accounts splits on the account number, giving one statement per account without reading a single page number.
Limitations to know
The PDF must have real, selectable text. A scanned batch is a picture of text and nothing will be found; run OCR PDF on it first, then split. Text is matched exactly as it appears, so a reference that is drawn as a barcode or an image is invisible to this tool. Bookmarks, form fields and attachments are not carried into the parts.
Split by Text: common questions
Nothing was found. What went wrong?
Either the area you chose has no text in it, or the PDF has no selectable text at all. Check the second possibility first: try selecting a word in the document with your mouse. If you cannot, it is a scan, and OCR PDF has to run before any text-based split will work.
Every single page came out as its own file.
The area you set includes something that changes on every page, most often a page number, a date or body text. Make the area smaller and aim it at the part that only changes between documents, such as the corner holding the reference number.
How do I know which area to use?
Draw it on the page preview: drag the orange box over the spot where each document's number is printed. The figures in the settings follow the box, as percentages of the page, and can be typed instead.
What happens to a document that runs over several pages?
It stays in one piece. In the area mode, continuation pages have nothing in the chosen area, and a page with an empty area is treated as belonging to the document before it. That is why an invoice of any length comes out whole.
Can I split on something other than a number?
Yes. The words mode matches any phrase, so Statement of account, Dear or Page 1 of all work if they appear once per document. Choose a phrase that appears on the first page of every document and nowhere else.