Case study · Public sector, citizen services

Shipped AI agents the audit team approved on day one

Eight-week engagement. Eight agents live, same-week rollout, zero audit gaps.

Columns of a civic building lit at night

The models worked. The sign-off did not come.

Nobody would sign off on the AI, and the model was not the reason.

Not the blocker

Model quality

The models worked. The output looked good in a demo.

Blocker

No traceable basis

No one could answer what the system had used to reach a conclusion.

Blocker

No answer for failure

No one could say what would happen if an agent got one wrong.

Why it counts

Citizen services

That is the whole question here. An agent that cannot be audited cannot be deployed.

Each blocker has a named control

The objections that stopped sign-off on the left, and the control that answers each one on the right.

  1. No traceable basisLineage on every signal
  2. No answer for failureKill switch per agent
  3. A weak basis buried in a confident answerConfidence scores
  4. Checks that run after the output has gone outPolicy checks on the signals
  5. Governance that lapses after go-liveOn-call governance pod
Every control on the right went in during the Weeks 2 to 6 build, not as a pass before go-live.

We started from the approval, not the demo

The order of the work decided what got built and what never got built at all.

  1. Start

    Ask what approval needs

    We began from what the audit team would need to approve, rather than from what the model could do.

  2. Then

    Let that set the scope

    It sounds like a small reordering. It decides what gets built in Week 2.

  3. So

    Controls go in with the build

    Lineage, policy checks, confidence scores and the kill switch were part of the build, not a later pass.

  4. Result

    Nothing to retrofit

    A retrofit is exactly what a reviewer learns to spot. At the presentation there was nothing left to argue about.

The controls that got it approved

Completeness is the proof here, so each of these was part of the build rather than something added before go-live.

  • Lineage on every signal

    What an agent used to reach a conclusion can be traced back to source after the fact.

  • Policy checks

    Run on the signals feeding a decision, not on the output once it has already gone out.

  • Confidence scores

    Attached to every signal, so a weak basis is visible instead of buried inside a confident-sounding answer.

  • Kill switch per agent

    Each of the eight can be stopped on its own.

  • On-call governance pod

    Handed over at the end, so the controls stay live after go-live, which is where most AI governance quietly lapses.

Built first, not retrofitted. Nothing on this list was bolted on after the agents already worked.

Eight weeks, three phases, no backdated heroics

What we built, and in what order.

  1. Week 1

    Audit

    Two calls, fixed fee. We write down what the audit team would need to approve as a one-pager the customer keeps either way.

  2. Weeks 2 to 6

    Build for sign-off

    Lineage, policy checks and confidence scores on every signal, and a kill switch on each agent. Eight agents built to be approved, not demonstrated.

  3. Weeks 7 to 8

    Hand over and operate

    An on-call governance pod so the controls stay live after go-live. Approved and rolled out in the same week it was presented.

What the agents are built on

No reseller layer, and no mystery about which model is doing what.

Partner

Anthropic Claude Partner Network

We build directly on Claude rather than through an intermediary, so the behaviour you sign off is the behaviour that runs.

Every day

Sonnet

Production reasoning. The work that runs every day.

Depth

Opus

The smaller set of calls where depth matters more than cost.

Volume

Haiku

The cost-sensitive long tail. High volume, low complexity, priced for it.

Eight agents live, zero audit gaps

No money figure on this one: the proof is what the audit team could find and what it could not, taken from the customer's own systems.

0Audit gapsAcross every agent running in production.
8Agents liveIn production on citizen services.
Same weekRolloutApproved and rolled out in the week it was presented.
8 weeksEngagementFixed scope, fixed fee, agreed before Week 1.

Before you ask

Why is the customer not named?

Anonymised where we have to be, specific everywhere we can be. Reference calls are possible after the audit, subject to the customer agreeing.

What happens when an agent gets one wrong?

That is what the controls are for. Confidence scores make a weak basis visible, policy checks run before an output counts, and each agent has its own kill switch. After go-live the on-call governance pod is who you call.

What did this cost?

Fixed scope, fixed fee, agreed before Week 1. We give you the number on the audit call once we know the shape of the work, rather than publishing a band that would not apply to you.

Is our situation close enough to this one?

That is exactly what the Week 1 audit answers. Two calls, fixed fee, and you keep a written one-pager whether or not we work together.

Is your shape close to this one?

Two calls, a fixed fee, and a written one-pager you keep either way. We will tell you on the call if we are not the right fit.