/ Case study · Startup (own product)
LLM as adjudicator, not decision-maker: an auditable AI architecture for financial reconciliation
How I architect AI for workflows where a wrong answer costs money: deterministic engines generate candidates, the model only chooses among them under a strict schema, and every decision is traceable.
- 0Model calls in the deterministic-only baseline
- ≥85%Auto-match target by volume before MVP (exit gate, not a result)
- 100%Decisions with source provenance — by design
The problem
Mid-market groups in the Gulf with two to six legal entities close their books by hand. Finance teams match bank lines, invoices and intercompany balances in spreadsheets, and every month-end turns into a scramble. The obvious “AI” answer — let an LLM read everything and decide — fails exactly where it matters: it can’t be audited, it isn’t repeatable, and one confident wrong match costs more trust than a hundred correct ones earn.
I’m building Tasweya (تسوية), a reconciliation and close-automation layer for these groups. This case study is about the architecture pattern, not the product pitch.
Constraints
- Sits beside existing accounting systems with read-only access. It is not a general ledger and doesn’t file tax.
- Multi-tenant from day one: every row is scoped to a tenant and group.
- Data residency in the UAE.
- Bilingual Arabic/English with right-to-left support as a first-class requirement, not a translation pass.
- Every automated decision must be explainable to an auditor months later.
Architecture
- Ingestion normalises every source (bank statements, ERP exports, documents) into one canonical
FinancialEventwith mandatory provenance back to an immutable, content-addressed source file. - A pure, I/O-free matching engine (a plain Python library with no network or database access) generates ranked match candidates deterministically.
- The LLM only adjudicates among engine-generated candidates and must answer in a strict JSON schema. It cannot invent a match the engine didn’t propose.
- Every outcome — automatic, model-assisted or human — is written to an append-only decision log.
- Deterministic-only mode is first-class: the system works with 0 model calls, so AI is an accelerator, never a dependency.
- Postgres row-level security, rolled out from permissive to enforcing, isolates tenants at the database layer.
- Deployed as a modular monolith (API and worker from one image) on Azure UAE North.
Key decisions
- Model as judge, not author. Constraining the model to choose from candidates removes the largest class of hallucination risk and makes its output testable.
- Provenance is not optional. A decision without a traceable source is rejected at the schema level, so by design 100% of decisions carry their source provenance.
- The engine is the moat, not the model. Models will commoditise; the matching and mapping logic, and the data it learns from, won’t.
- Security mapped up front. The architecture is reviewed against the OWASP Top 10 for LLM Applications (2025) and the OWASP web Top 10, rather than retrofitted.
How it’s evaluated
- Phase 0 runs the engine with no model calls on real anonymised data from at least five groups.
- The exit gate before building the MVP is ≥85% auto-match by transaction volume. If deterministic matching can’t get close, adding an LLM won’t save it.
- Model-assisted decisions are measured against the deterministic baseline on the same data, so the model has to earn its place.
What I’d do next
- Publish the eval harness pattern (golden sets per entity type, regression on every engine change).
- Add a reviewer-facing explanation view that renders the provenance chain for any decision.
- Bring the same pattern to client workflows: document intake, KYC review queues and ops triage all fit the “deterministic candidates, model adjudicates, human approves” shape.
Stack
- Python 3.12
- FastAPI
- Pydantic v2
- PostgreSQL (RLS)
- SQLAlchemy 2 / Alembic
- React + TypeScript
- Azure UAE North
- Azure AI Foundry
- Azure Document Intelligence