Pull structured data out of PDFs and into a spreadsheet
Watch a Drive folder for new invoices, read them with OCR, use AI to pull out the fields you actually need, and land them in a spreadsheet ready for bookkeeping.
What this looks like
One step on screen at a time. Read it, do it, tap Next.
Pull structured data out of PDFs and into a spreadsheet
Step 1 of 12
Watch the folder for new PDFs
Create a workflow and add a Google Drive Trigger watching your invoices folder for new files.
Checkpoint
Dropping a test PDF into the folder fires the trigger.
Before you begin
Tick these off as you sort them. Stopping halfway to create an account is how a fifteen-minute build becomes an hour.
0 of 10 ready
Ticks are saved on this device, so you can come back to them.
- Have ready
Where invoices land — by upload, forwarding rule, or manual drag-and-drop. Keep it exclusive to documents you want extracted.
- Account
This walkthrough uses an OCR API reached over HTTP to turn scanned PDFs into readable text before anything else happens.
- API key
Found in your OCR provider's dashboard, used to authenticate the extraction request.
- API key
Used to read the OCR text and pull out the specific fields you want, since OCR alone only gives you raw text, not structure.
- Have ready
Columns matching the fields you want: InvoiceNumber, VendorName, InvoiceDate, TotalAmount, FileLink.
- Account
Used to watch the folder, download files, and write the extracted results.
- Time
The two AI steps take patience to prompt correctly — budget extra time for that.
The basics — ticked once, remembered everywhere
- Account
Either the hosted version or a self-hosted install. Everything here works the same on both — the only visible difference is the shape of your webhook URLs.
Create a workspace - Know-how
Adding a node, connecting two of them, and pressing Execute. If any of that is new, the free primer covers it in about ten minutes.
- Credential
You will paste at least one secret during setup. A password manager beats a notes app, and it stops you pasting a live key into a chat window by accident.
The shape of the workflow
The nodes you will end up with, left to right, in the order you add them.
4 things to set here- 1Google Drive Trigger — the trigger — everything starts here
- 2Google Drive — Download the file
- 3HTTP Request — Run OCR on the document
- 4Gemini — Extract the fields you need with AI
Skip the node hunting
Import a starter file with every node for this build already on the canvas, named after its step and wired in order, each one carrying a sticky note with the values to enter. You still connect your own accounts and fill the fields — that is what the walkthrough takes you through — but you never start at a blank canvas.
In n8n: Workflows → Import from file.
What you will build
Jump straight to any step, or print the list and tick it off beside your n8n tab.
- 1Watch the folder for new PDFs
- 2Download the file
- 3Run OCR on the document
- 4Extract the fields you need with AI
- 5Parse the AI's response into real fields
- 6Validate the extracted data before saving it
- 7Write good extractions to the sheet
- 8Flag anything that failed validation
- 9Test with real invoices, then activate
- 10Drop in a clean, typed invoice
- 11Drop in a poor-quality scan
- 12Drop in a non-invoice PDF
If something goes wrong
Everything ready?
Start walkthrough