Invoice intake with matching
Ten documents the way they really arrive: PDF, stamped scan, skewed phone photo, crumpled receipt, e-invoice. Four errors and a tight discount deadline are planted. All companies are made up.
Why bother? Because retyping costs money.
A PDF in the inbox, a stamped letter, a phone photo, a receipt from a jacket pocket, plus e-invoices as XML. Someone types in the values, digs out the purchase order and delivery note and compares line by line.
Almost one invoice in five becomes an exception that has to be sorted out by hand: wrong quantity, different price, missing purchase order. On average it takes over 9 days for an invoice to get through. US figures; there are no current German ones.
No more retyping. Overbilled quantities and duplicate invoices surface before payment, not after. Discount deadlines stop slipping: 2% for paying in 10 instead of 30 days works out at 36.7% a year. And every decision can be traced.








Recorded on 01/10/2026. Draft: Gemini's free tier allows 20 requests per day and model, so the readings come from three Gemini Flash versions. The final run will be Gemini 3.8 Flash throughout.
Not a pilot. Five points make the difference.
79% of companies report rolling out AI agents. Only 11% run them in production (Gartner 2026). The model is rarely the problem. These five points are.
Approval
Nothing gets paid automatically. A human reviews amber cases, red ones are blocked, green ones are posted with one click.
Traceable
Every value shows where it came from. Every finding names its rule and numbers.
Permissions
The model only reads. No write access to the ERP, no access to payments.
Cost
0.48 US¢ per document, measured. 1,000 documents a month: about €4.16.
Test
188 of 189 fields, 8 of 9 documents fully correct. All 4 planted errors and the discount deadline caught.
The model reads. Code does the maths.
Model only for reading
Matching, tax and mandatory fields are rules. Rules belong in code: same input, same output, auditable by accounting.
E-invoices without a model
Since 2025 German companies must accept e-invoices, and from 2027 or 2028, depending on turnover, issue them. I parse the XML directly: zero cents, zero errors. AI only where there's paper.
Source location over speed
DeepSeek V4.1 Flash reads almost as well (186 vs 188 of 189 fields) and is 9.6× faster. But its boxes miss, see below. Without a source location nobody can verify, so Gemini.


| Model | Docs | Fields correct | Time avg | Cost avg | Locations |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | 5 | 101/102 | 5.6 s | 0.50 US¢ | precise |
| Gemini 3.7 Flash | 3 | 61/61 | 5.2 s | 0.54 US¢ | precise |
| Gemini 3.6 Flash | 6 | 127/128 | 5.2 s | 0.44 US¢ | precise |
| DeepSeek V4.1 Flash | 9 | 186/189 | 0.5 s | 0.22 US¢ | imprecise |
All 14 Gemini runs combined: 289/291 fields. The 188/189 above counts the 9 readings in the demo. Times measured on the free tier, which is throttled under load. DeepSeek is a Chinese provider: a minus for real customer data.
Three bugs. All mine.
Missed the duplicate at first
On the photo the model reads “Weber & Söhne”, in the PDF the full company name. My matcher keyed on the name and saw two suppliers. Now: invoice number plus amount.
Shipping flagged as “not delivered”
Shipping never appears on a delivery note. The matcher reported €8.21 overbilled. Lines without a delivery-note line are no longer quantity-checked.
My test set was wrong, not the model
The receipt prints no quantity. The model correctly returned none, my test set expected 1. Ten “errors” vanished once I fixed the test set.
* Vendor benchmark incl. staff time. A human still reviews flagged cases, so real savings are smaller than the difference. Mid-sized companies typically see 800–5,000 incoming invoices a month.
Cloud-API
Gemini 3.8 FlashThe demo runs on the free tier with made-up documents. In production: paid tier, no training on inputs, DPA with Google.
EU cloud
Vertex AI · europe-west3Same model, processed in Frankfurt. Switch: endpoint and credentials.
On-premise
PaddleOCR-VL / Qwen3-VL-8BShould run on a 16 GB graphics card (Qwen3-VL-8B quantised). Benchmark on my RTX 5070 Ti is next, results will land here.