Back to Growth Notes

A bookkeeping agent handles 713 purchase invoices a month across eleven administrations. It approves none of them. That refusal is the most valuable thing about it, and most automation projects get it exactly backwards.

Eleven separate administrations. One leisure group. An average of 713 purchase-invoice postings a month, spread across entities that each have their own cost centers, their own VAT quirks and their own opinions on how a supplier should be coded.

So we built an agent for it. It reads the invoice, matches the supplier, proposes the ledger account, fills the cost center, sets the VAT code and writes it into Exact Online. Every working day, without asking.

And it approves exactly zero of them.

That is not a limitation we are still working on. It is the design. Because the interesting question about automating invoice processing is not how much it can do. It is where it must stop, and whether anyone bothered to decide that on purpose.

The myth of the confident agent

Most automation gets sold on confidence. It reads the document, it knows the answer, it books the line. Demos love this. You drop in a PDF, the fields populate themselves, everyone in the room nods.

However, that demo hides the only number that matters in finance: what happens to the cases it got wrong.

An invoice-processing model that is right 95 percent of the time sounds like a strong result. In a classification benchmark it is. In a ledger it is a slow leak. Because the 5 percent does not announce itself. It looks exactly like the 95 percent: a filled field, a plausible account, a posting that sails through review because nothing about it looks odd.

An agent that guesses right 95 percent of the time still puts about 35 wrong lines into your ledger every month at 713 invoices. We’d rather build one that says: these seven, I don’t know.

Thirty-five wrong lines a month is 420 a year. Spread across eleven administrations, none of them big enough to trip an alarm on their own. You find them at year-end, or your accountant does, or you never find them and your cost-center reporting has been quietly wrong for three quarters.

Seven blanks with a reason attached is a ten-minute job on a Tuesday morning.

Design decision one: it drafts, it never signs

The agent creates drafts. It never approves and it never pays. That is in the contract we signed with the client, not in a settings screen where a future admin can flip it for convenience.

This matters more than it sounds. The moment an automation can approve its own work, you have removed the last place where a human reads a number before money moves. And the pressure to flip that switch always comes later, always from a good place: the drafts are piling up, the quality is clearly fine, someone is on holiday.

Therefore we made it structurally unavailable. The agent has write access to draft entries. It does not have the permission to approve. Not because we distrust the model, but because the person who signs off on spend should be the person who is accountable for it. Automation should compress the work in front of that signature, not remove the signature.

Furthermore, this is what makes the whole thing adoptable. A controller will let an agent prepare 713 postings. No controller will let an agent pay 713 invoices, and they are right not to.

Design decision two: unknown is not zero

This is the spine of the build.

When the agent is not sure of the cost center, it leaves the cost center blank. When it cannot determine the VAT treatment with confidence, it leaves the VAT code blank. In both cases it attaches a short reason: new supplier with no booking history, invoice line covers two entities, VAT statement on the document conflicts with the supplier’s usual treatment.

It does not fill in the most likely value.

Most systems do the opposite, because a fully populated record feels finished and a blank field feels like failure. But a blank field gets noticed by a human. A plausible wrong one does not. That asymmetry is the whole game. Your review process is tuned to spot gaps, not to re-audit confident answers.

So the agent’s output is not “713 postings, done”. It is “706 complete drafts, 7 with open fields and here is why”. The human work shifts from checking everything to resolving exceptions, which is the only version of this that actually saves time instead of relocating it.

Every booking carries its reasoning

Each draft records why it was booked the way it was: which supplier rule applied, which historical postings it matched against, which entity it was assigned to and on what basis. Not a confidence score. An audit line a controller can read.

Because “the system did it” is not an answer you can give your accountant, and it is not an answer that lets anyone improve the rules. If you cannot trace a booking back to the rule that produced it, you cannot fix the rule. You can only argue with the output, invoice by invoice, forever. This is the same failure pattern as a weak architecture behind a strong model: the intelligence is fine, the structure around it makes it unusable.

Design decision three: it learns, but only past a quorum

Self-learning is the feature everyone asks for and almost nobody scopes. The pitch is obvious: correct it once, it never makes that mistake again. Efficient. Also, in an eleven-entity administration, dangerous.

Because a single correction is not knowledge. It might be a one-off. It might be someone booking something to the wrong cost center in a hurry. It might be a temporary arrangement that applies to one project and nothing else. Promote that to a rule and you have not taught the agent anything. You have let one person’s exception become policy across eleven administrations.

So we set a quorum. A correction becomes a rule when two different people make the same correction, or when one person makes the same correction three times. Anything below that threshold is logged as a signal, not a rule. It waits.

Plus, rules expire. A learned rule has a 60-day life. If nobody reconfirms it in that window, it drops out and the agent goes back to asking. Suppliers change, contracts change, projects end. A rule learned in March should not be silently governing your bookings in November because nobody thought to go looking for it.

And there is one switch that turns learning off entirely without disabling the agent. Month-end close, an audit, a migration, a period where you want the behaviour frozen: flip it. The agent keeps drafting on the rules it has and stops adopting new ones. That switch exists because “we need to stop it learning” should never require taking the automation offline.

Self-learning without a quorum is not an agent that gets smarter. It is an agent that cements its own mistakes, faster and faster, with every correction it misreads as a lesson.

Seven questions for your own invoice automation

If you already automate purchase invoice processing, or a vendor is about to sell you on it, run through this. The answers tell you more than any accuracy percentage on a slide.

  • Your automation stops at a draft that a human has to sign off, and cannot approve or pay on its own
  • When it is not confident about a field, it leaves that field blank with a reason instead of filling in its best guess
  • Every booking can be traced back to the specific rule or history that produced it, without taking the system’s word for it
  • A correction only changes behaviour past a defined threshold, and you know how many people and how many repeats that takes
  • Learned rules expire on their own if nobody reconfirms them, rather than staying until someone notices they are wrong
  • There is a single switch that stops the agent learning without shutting the automation down
  • You know your accounting platform’s daily API limit and how close your automation runs to that ceiling at peak

Two or more unchecked is a design gap, not a tuning issue. If the first two are unchecked, you do not have invoice automation. You have an unsupervised bookkeeper with no accountability and a very good memory for its own mistakes.

This is not about Exact Online

We built this on Exact Online, and the API limits shaped the batching, the retry logic and how the agent paces itself across eleven administrations. That part is real engineering work and it matters at 713 invoices a month.

But none of the logic above is tied to Exact. Draft-only permissions, blank-over-guess, traceable bookings, quorum-based learning, expiring rules: any accounting package with a decent API can carry that. The integration is the easy half.

The hard half is encoding the client’s own booking rules. Which supplier goes to which account in which entity. How intercompany lines get split. Which cost centers are project-bound and which are permanent. When a VAT treatment deviates and why. That knowledge lives in two people’s heads and a spreadsheet nobody has updated since 2022, and getting it out is most of the project.

That is also why “just plug in an AI” never works here. There is no model that knows your booking rules. There is only the work of writing them down properly for the first time.

The ODB Way

We do not start by asking what the agent can automate. We start by asking where it must stop, who signs, and what happens to the cases it gets wrong. Then we build inward from there.

In practice that means mapping the actual process before touching an API, encoding the client’s real booking rules instead of generic best practice, and shipping something that runs every day in production rather than a pilot that impresses once. The refusals are part of the spec, not a limitation to be removed in version two.

If you want the broader argument for why automation should produce actions and decisions rather than answers, read why your AI should drive action. And if you are wondering why so many finance and back-office pilots look great in a demo and then stall, the four places AI pilots quietly die covers most of the reasons we run into.

Onno de Bel

Onno de Bel

AI Engineer & Architect | ODB Growth

Ready to build something
that moves a number?

Let's find out where AI actually pays in your process, and build it.

Plan a Growth Call