In build Case studyIn diagnosisDeterministic by design

Invoice Drafting Where the Language Model Never Touches a Number

A completed job triggers an invoice draft built from a versioned rules registry that encodes the finance lead's judgment. The draft must clear strict validation before anything is written anywhere, and the language model is permitted exactly two jobs, neither of which is arithmetic.

In buildcurrently in diagnosis0totals set by a modelFail closedvalidation gate before any writeVersionedrules registry, reconstructable

The problem

The worst possible place to put a probabilistic system.

  • Invoicing a completed job is repetitive judgment: which invoice type, which line items, which tax treatment, which account code. It is slow, and it queues behind the one person who holds the judgment.
  • It is also the worst possible place for a probabilistic system. An invoice that is subtly wrong is worse than an invoice that is late, because it reaches a customer and an accounting ledger and is discovered months later.
  • Most attempts at AI invoicing put the model in the middle of the flow and then try to constrain it with instructions. That is the design error, and no amount of prompt engineering repairs an architecture that lets a model near a total.

The solution

Rules produce the numbers. The model produces the sentences. A gate decides if anything is written.

The judgment is encoded once, explicitly, in a versioned rules registry built with the finance lead. A completed job triggers a deterministic draft from that registry, so the same job produces the same draft every time.

The draft is then checked against strict validation, and only a draft that clears it is written anywhere. The language model is used for exactly two things: naming the invoice type and wording the customer-facing lines. It never sets, adjusts, reviews, or influences a total, a tax decision, or an account code.

Determinism where it countsThe boundary around the model is architectural, not instructional. The model is not told to avoid totals. It is never given them, and its output has nowhere to reach them.

How it works

Six components, and a hard boundary.

RegistryJudgment written down once

The finance lead's rules are encoded explicitly and versioned, so a change in policy is a reviewable change to a file rather than a new habit that spreads by word of mouth.

TriggerA completed job, not a schedule

The draft is produced from the event that actually carries meaning. Drafts track reality rather than a clock, and a job that did not complete produces nothing.

DeterministicSame job, same draft

Totals, tax treatment, and account coding come from the registry alone. Given the same job, the draft is byte-identical every time, which is what makes review and diagnosis possible at all.

GateStrict validation before any write

A draft that fails validation is not written, not partially written, and not queued for later. It stops and raises. Fail closed is the only acceptable default when the output is financial.

Model scopeExactly two jobs

The language model names the invoice type and words the customer-facing lines. That is the entire surface. It has no access to totals, tax treatment, or account codes, by construction rather than by instruction.

VersioningReconstructable on any date

Because the registry is versioned, what the system believed on any given date can be reconstructed. That matters the first time someone asks why an invoice from three months ago looks the way it does.

01Job completedThe real event, not a scheduled sweep
02Versioned rules registryThe finance lead's judgment, written down
03Deterministic draftTotals, tax, and coding. No model involved
04Model wordingInvoice type name and customer-facing lines only
05Strict validationFails closed. Nothing partial is ever written
06Written onceTo the accounting system, after the gate
The model enters at step four and touches only text. Steps two, three, and five never see it.

The boundary

Where the model is not allowed, and why that is architecture rather than policy.

What the model never sees

Totals, subtotals, tax treatment, account codes, and rate tables are produced by the registry and are not part of any model input or output. A boundary that depends on the model choosing to respect an instruction is not a boundary.

What the model is genuinely good at

Naming an invoice type sensibly and wording a customer-facing line clearly are language problems, and the model is better at them than a template is. Using it for exactly that, and nothing adjacent, is the point.

Status: in build, in diagnosis

This system is in build and currently in diagnosis against real job data. It is described here at the design level. It is included because the design commitment is the transferable part, and because the boundary it draws is the one most AI invoicing attempts get wrong.

What it demonstrates

The skills behind the system.

Deterministic automation with a narrow model surfaceRules-registry design and versioningValidation gates and fail-closed designFinancial-system integrationEncoding domain expert judgmentEvent-driven triggersAuditability and point-in-time reconstructionKnowing where not to use a modelArchitectural constraints over prompt instructions

Outcome and what is next

In build. The design commitment is already settled.

The system is in build and in diagnosis against real job data. The architecture is not in question: the numbers come from reviewed, versioned rules, the model writes sentences, and nothing is written anywhere until a strict gate permits it. Results will be added here when the system has produced enough of them to report.