2026 — ongoing

tax

The same ingest-and-reconcile treatment applied to the documents that only matter once a year, and only if you can find them.

Private repository — the writing here describes the approach rather than linking to code.

  • Python
  • SQLite
  • structured extraction

Donation receipts, dividend statements, PAYG summaries, work-expense invoices, trust distributions, managed-fund statements. Documents that arrive across twelve months, matter for one week, and are individually too small to justify filing properly — which is exactly why they are never where you need them in July.

Structurally this is the same tool as bills: ingest, classify, extract into a schema, validate, file, make it queryable. The interesting part was not building another CLI but working out which parts of bills were genuinely about bills and which were about documents. The answer turned out to be most of it, and the split produced a shared core rather than a fork.

That is the test I now apply to each of these tools. If a second one in the same family cannot be built mostly from the first, the first was not factored properly.

The repository is private.