invoice-extractor.

Extract fields from PDF invoices. Vendor, invoice number, date, amount, VAT. Deterministic rules. No ML. No data leaves your machine.

Live - EUR49 one-time
Buy now - EUR49

The problem

Every month, you get invoices from random SaaS tools that don't sync with your accounting software.

You spend an hour downloading PDFs, finding the VAT number, and typing it into a spreadsheet. It's not hard work — it's just tedious. And it repeats every single month.

Existing tools are expensive and complex.

Full accounting suites cost $50+/month and require setup. OCR tools are inaccurate. You need something that just extracts the fields and gets out of the way.

What it does

Feed it a PDF invoice. Get back structured fields:

vendor

Vendor name extracted from the invoice.

invoice_number

Invoice or receipt number.

invoice_date

Invoice date (YYYY-MM-DD).

total + currency

Total amount and currency (USD, EUR, GBP).

vat_number

VAT or tax ID number.

How it works

$ echo '{"pdf_path":"/path/to/invoice.pdf"}' | python extract.py
{"vendor": "Acme Corp", "invoice_number": "INV-2026-001", "invoice_date": "2026-01-15", "total": 1200.0, "currency": "USD", "vat_number": null, "line_items": []}

PyPDF2 for text extraction, deterministic regex for field parsing. No model, no API call, no data leaves your machine.

Why deterministic

An LLM extracting invoice fields is a black box you cannot audit.

When an extractor says "the total is $1200", you need to know why. A rule you can read is a rule you can trust, debug, and improve.

Buy

EUR49 one-time. Runs locally. No subscription, no cloud, no data collection.

Buy now - EUR49