Industries · Public sector

AI agents the audit team approved, and the controls that got them there

We build data, analytics and AI for public bodies, starting from what an auditor needs to approve rather than from a demo. Our published public-sector work: eight AI agents in production on citizen services, approved and rolled out in the week they were presented, with zero audit gaps.

Published work in this industry
1 case study
AI agents live
8
Audit gaps
0
Rollout
Same week
On this pageWhat we hear in public sector
  1. 01What we hear in public sector
  2. 02What we have delivered
  3. 03Where to start
  4. 04How the first weeks run
  5. 05Questions we get asked
  6. 06Further reading
01

What we hear in public sector

Each of these comes from our public-sector case study or from an article we wrote for teams in it.

  1. 01The models work. Nobody will sign off.

    The output looks good in a demo, but no one can say what the system used to reach a conclusion, or what happens when an agent gets one wrong.

  2. 02Data that is not allowed to leave the country

    The database is in the right region. The backups, the logs and the vendor's support tooling may not be.

  3. 03Access rules that live in the dashboard

    A row filter protects the dashboard. It stops protecting anything the moment a notebook or an export reaches the same data.

  4. 04A process nobody has written down

    The flowchart describes the good day. The two days a week the job actually takes are spent on everything it leaves out.

  5. 05An app that can change records

    The first write-back button is the cheapest moment to get identity, permissions, an append-only log and a way to reverse a change.

02

What we have delivered

One published case study, anonymised: citizen services, eight weeks, fixed scope.

0Audit gaps

Across every agent running in production.

8Agents live

In production on citizen services.

Same weekRollout

Approved and rolled out in the week it was presented.

8 weeksEngagement

Fixed scope, fixed fee, agreed before Week 1.

The controls that got it approved

Each was part of the build rather than something added before go-live.

  1. 01Lineage on every signal

    What an agent used to reach a conclusion can be traced back to source after the fact.

  2. 02Policy checks

    Run on the signals feeding a decision, not on the output once it has already gone out.

  3. 03Confidence scores

    Attached to every signal, so a weak basis is visible instead of buried inside a confident answer.

  4. 04A kill switch per agent

    Each of the eight can be stopped on its own.

  5. 05An on-call governance pod

    So the controls stay live after go-live, which is where most AI governance quietly lapses.

Read the full account

  1. Case study, public sectorShipped AI agents the audit team approved on day one

    We began from what the audit team would need to approve, rather than from what the model could do.

    Read the case study ↗
03

Where to start

Four places a first engagement can begin. Week 1 tells you which one comes first.

  1. AIGovernance before go-live

    Defensible in a board meeting, auditable on demand.

    AI governance ↗
  2. AIKnow what is worth building

    One week, and you know which two or three AI initiatives are worth building.

    AI audit and roadmap ↗
  3. ApplicationsOff paper and shared spreadsheets

    The forms, approvals and registers that still run on paper, WhatsApp and one shared spreadsheet, moved into a system with a trail.

    Process digitisation ↗
  4. DataOne set of records departments accept

    One set of records two teams both accept, with integration, quality and governance underneath.

    Data integration and governance ↗
04

How the first weeks run

The same three phases in every industry. What changes is what we build in them.

  1. Week 1, fixed feeAudit

    Two calls and a written one-pager. Yours to keep whether or not you continue.

  2. Weeks 2 to 6Build

    Something real in front of real users by week three. Median six weeks to the first production-grade artefact.

  3. Week 7 onwardOperate

    Quarterly reviews and on-call governance. The people who built it pick up the phone.

05

Questions we get asked

Which models do the agents run on?+

Claude, directly, with no reseller layer. Sonnet for the work that runs every day, Opus where depth matters more than cost, and Haiku for the high-volume tail.

Can one agent be stopped without stopping the rest?+

Yes. In the published case each of the eight agents has its own kill switch.

What happens after go-live?+

An on-call governance pod keeps the controls live. Go-live is where most AI governance quietly lapses, so it is part of the build rather than an extra.

Is our data used to train a model?+

No. Not ours, not anyone else's. Where a build needs a model to see your data, it sees it under your agreement. Our page on how we use AI sets out the detail.

06

Further reading

Ready when you are

Start with the Week 1 audit

Two calls and a written one-pager naming the first thing worth building, and who owns it. You keep it either way. NDA-friendly, fixed scope. Write to hello@woodfrog.tech.

Ask Priya

A question this page did not answer?

Priya is woodfrog’s digital representative. Ask her here and get a straight answer, no call needed.