Extract fields from PDF invoices. Vendor, invoice number, date, amount, VAT. Deterministic rules. No ML. No data leaves your machine.
Every month, you get invoices from random SaaS tools that don't sync with your accounting software.
You spend an hour downloading PDFs, finding the VAT number, and typing it into a spreadsheet. It's not hard work — it's just tedious. And it repeats every single month.
Existing tools are expensive and complex.
Full accounting suites cost $50+/month and require setup. OCR tools are inaccurate. You need something that just extracts the fields and gets out of the way.
Feed it a PDF invoice. Get back structured fields:
Vendor name extracted from the invoice.
Invoice or receipt number.
Invoice date (YYYY-MM-DD).
Total amount and currency (USD, EUR, GBP).
VAT or tax ID number.
$ echo '{"pdf_path":"/path/to/invoice.pdf"}' | python extract.py
{"vendor": "Acme Corp", "invoice_number": "INV-2026-001", "invoice_date": "2026-01-15", "total": 1200.0, "currency": "USD", "vat_number": null, "line_items": []}
PyPDF2 for text extraction, deterministic regex for field parsing. No model, no API call, no data leaves your machine.
An LLM extracting invoice fields is a black box you cannot audit.
When an extractor says "the total is $1200", you need to know why. A rule you can read is a rule you can trust, debug, and improve.
EUR49 one-time. Runs locally. No subscription, no cloud, no data collection.
Buy now - EUR49