AI governance

Governance you can defend in a board meeting and audit on demand

Risk classification, bias and explainability evidence, privacy controls and automatic rollback, built into the pipeline the model ships from rather than written into a policy document. We start bottom-up from the models you already run, not from a generic mandate.

Framework
4 pillars
Assessed against
EU AI Act, NIST AI RMF, ISO 42001
Week 1
Fixed-fee audit, two calls, a one-pager
Team on the work
A pod of three
Median to first production artefact
Six weeks

The four pillars, in the order they have to happen

We start from your real deployments and build the structure around them, rather than issuing a framework and hoping the models comply.

  1. Pillar 1

    Discovery and risk profiling

    Every model, agent and shadow deployment gets mapped and classified. A customer chatbot rates low, a fraud detection model medium, a credit scoring engine high, an HR screening model critical.

  2. Pillar 2

    Safety engineering

    Bias testing, explainability and privacy controls go into the model's path to production, so they run on every release instead of once for a report.

  3. Pillar 3

    Accountability

    Each classified model gets a named owner, a documented decision path and a sign-off route that reaches the chief executive for the high-risk cases.

  4. Pillar 4

    Continuous compliance

    Live monitoring for drift and policy breach, with automatic rollback. Compliance becomes a state you hold and can show, not a date you once passed.

The eight pieces of work

Most engagements take a subset of these, not all eight.

Risk4 stages

AI risk assessment and classification

Intake, classification, controls, report. Models are scored against regulatory thresholds and given proportionate controls, so leadership signs off on something with a number attached.

How we evaluate
Fairness< 5%

Bias and fairness auditing

Outcomes tested across demographic groups against a stated variance threshold, before launch and again on live models. A pass is a measurement, not an opinion.

XAISHAP

Model explainability

Feature importance and local interpretability attached to individual decisions, so a customer, a regulator or an underwriter can be told why the answer came out that way.

Regulation83%

Regulatory compliance audit

Technical audit and documentation against the EU AI Act, NIST AI RMF and ISO 42001, article by article, ending in a conformity score with the gaps named.

Ethics3 layers

Ethical AI framework design

Your principles written down, turned into policies, then compiled into automated guardrails. The third step is the one most programmes quietly skip.

Privacyε = 0.1

Data privacy and security for AI

PII scanning on every data load with redaction, masking, removal and hashing, plus differential privacy where re-identification is the real risk.

Data engineering
MonitoringLive

Continuous model monitoring

Drift, anomaly and policy-breach detection on production models, with automatic rollback and an incident log written as things happen rather than afterwards.

Programme6 weeks

Governance programme management

Agile delivery with a pod of three, budget control and a maturity level that actually moves. Six weeks is our median to a first production-grade artefact.

Each obligation maps to a thing that exists in your pipeline

An article of the EU AI Act is only satisfied by an artefact someone can open and read.

  1. Risk management system (Art. 9)Risk register with every model classified
  2. Data governance and quality (Art. 10)PII scanning and quality checks on load
  3. Technical documentation (Art. 11)Model cards written and kept current
  4. Record-keeping and logging (Art. 12)Audit trail and incident log
  5. Transparency obligations (Art. 13)SHAP explanations attached to decisions
  6. Human oversight mechanisms (Art. 14)Named owner, sign-off and automatic rollback
On one engagement, five of these six were assessed compliant and human oversight was still in review, which is what an 83% conformity score meant in practice.

The numbers a governance audit actually produces

These come from real assessments and each one is a position on a date rather than a permanent property.

83%Conformity score at last auditFive of six EU AI Act articles assessed compliant, one still in review. Change a model and the score has to be earned again.
4%Widest gap in approval ratesAcross five demographic groups, against a threshold of 5%. Equalized odds came out at 0.97 and calibration error at 0.03.
0.42Largest single SHAP contributionCredit history, on a lending decision that moved from a 0.35 base value to a 0.92 output. Every feature's push is written down and readable.
1,242Email addresses redacted on one loadAlongside 856 phone numbers masked, 124 SSN and medical IDs removed and 463 physical addresses hashed, at a differential privacy budget of 0.1.

Three layers, and only the third one enforces anything

A value becomes a policy, and a policy becomes something that runs on every commit.

  1. Layer 1

    Core values

    Transparency, accountability, fairness and human oversight, written as your organisation's own constitution rather than lifted from a template that fits nobody.

  2. Layer 2

    Governance policies

    The model card policy, the bias testing SLA, the audit trail requirement. This is where a value becomes something a reviewer can check against.

  3. Layer 3

    Automated guardrails

    A bias gate in CI/CD, real-time monitoring, automatic rollback and generated compliance reports. Policy that executes, rather than policy that is filed.

If a policy in layer 2 has no layer 3 equivalent, it is a document. We would rather cut it from the framework than let it sit there implying a control that does not exist.

What layer three caught on a single data load

The personal data found and handled by PII scanning before that load reached a model.

Email addresses redacted1,242
Phone numbers masked856
Physical addresses hashed463
SSN and medical IDs removed124
Bars are counts of records of each type found on one data load, in the same unit throughout, at a differential privacy budget of 0.1.

How a governance programme runs

The same shape as every Woodfrog engagement: audit in week one, build to week six, operate from week seven.

  1. Week 1

    Risk assessment

    Two calls and a fixed fee. Every model, agent and shadow deployment we can find, classified by risk, with the gaps named. You keep the one-pager either way.

  2. Weeks 2-6

    Policy design

    Values become policies with owners and thresholds. Risk classification is agreed with the people who will be putting their name to the sign-off.

  3. Weeks 2-6

    Framework build

    Guardrails go into CI/CD, model cards get written and audit logging is wired in. A pod of three does the work.

  4. Week 6

    Compliance audit

    The estate is assessed article by article and the conformity score is produced. This is the point where the number stops being an estimate.

  5. Week 7 onward

    Go-live and operate

    Monitoring runs, drift alerts route to a named owner, rollback is exercised rather than assumed, and the incident log starts filling up.

What keeps running after go-live

Governance fails in year two, not year one, and it fails because someone had to maintain it by hand.

  • Drift is detected, not discovered later

    Accuracy is tracked continuously. When the curve moves, an alert fires against the model that moved, rather than surfacing in a summary read the following quarter.

  • Rollback is automatic and exercised

    A breach returns the system to the last known-good version without waiting for a meeting. Rollback that has never been tested is a claim, not a control.

  • Every incident lands in a log

    Open, owned, resolved. The log is the first thing a regulator or a board asks for, and it cannot be reconstructed convincingly after the fact.

  • A bias gate sits in the pipeline

    Fairness tests run on the way to production. A model that misses the threshold does not ship, which costs far less than withdrawing it once it has.

  • Compliance reports generate themselves

    From the audit trail rather than from a spreadsheet somebody maintains in their own time. Manual reporting is what quietly kills these programmes.

These are capabilities we build and run, not certifications: nothing here is SOC 2, ISO 27001, HIPAA or GDPR. Where the EU AI Act, NIST AI RMF or ISO 42001 applies, we assess and document; the certificate comes from an accredited body.

Where the conformity score lands in the engagement

The compliance audit sits in week six, one week before the programme moves into operation.

Week 1, fixed-fee risk assessmentWeek 7 onward, operate
Weeks 2 to 6 are policy design and framework build. Six weeks is our median to a first production-grade artefact, which is the point where the conformity score stops being an estimate.

Where this is the wrong answer

Four situations where we would tell you to keep your money.

Your models do not decide anything about a person

Governance earns its cost when a model touches money, eligibility or employment. An internal summariser with a human reading every output does not need a framework wrapped around it.

Nobody is ever going to ask you to defend a decision

If no regulator, auditor, board or customer will put that question to you, most of this is overhead. Buy it because the question is coming, not because the topic is loud this quarter.

What you want is a document, not a control

We can write policy, but we will not hand over a set of principles with nothing in CI/CD behind them and call it governance. If a PDF is the deliverable, a law firm is the better purchase.

The models are not in production yet

Classification and guardrails need something real to attach to. At prototype stage, spend the money on an evaluation harness instead and come back when something is live.

What a fairness audit looks like when it passes

Approval rates for the same model across five demographic groups, from one bias and fairness audit.

Group A (male, 25-34)82%
Group B (female, 25-34)80%
Group C (male, 35-50)84%
Group D (female, 35-50)83%
Group E (non-binary)81%
Bars are approval rate as a percentage of applicants in that group. The widest gap is 4 points against a 5-point threshold, so this model shipped. A failing chart is the useful one, and we show clients those too.

Questions buyers put to us

The five that come up in almost every first call.

Do you hold SOC 2, ISO 27001, HIPAA or GDPR certification?

No, and you should not read one into this page. What we build are capabilities: PII scanning and masking, differential privacy, audit trails and automated guardrails. Where the EU AI Act, NIST AI RMF or ISO 42001 is what you are measured against, we assess your estate article by article and document the evidence. The certificate itself is issued by an accredited body, not by us.

What does an 83% conformity score actually mean?

On one engagement, five of the six EU AI Act articles assessed came out compliant and one, human oversight mechanisms, was still in review. That is a position on a date. Retrain a model or change a data source and the assessment has to be run again. Anyone quoting you a compliance percentage before looking at your models is guessing.

We already have an AI policy. What would you add?

Usually layer three. Most policies we are shown are sound on values and reasonable on policy, and have nothing in the pipeline enforcing them. The work is turning a bias testing SLA into a CI/CD gate that blocks a release, and an audit trail requirement into logging that records what happened while it happens.

What happens in week one?

Two calls, a fixed fee, and a one-pager: every model and agent we can find including shadow deployments, classified by risk level, with the gaps and the order in which to close them. You keep the one-pager whether or not you continue. The checklist we work from is published at /audit-checklist.

Who does the work?

A pod of three. Woodfrog was founded in Pune in 2023, is an Anthropic Build Partner and joined the Claude Partner Network at launch, and has run 50+ projects for 20+ clients. Six weeks is our median to a first production-grade artefact. Examples are on /case-studies.

The rest of the practice

These are genuinely different jobs with different ways of failing. Most engagements start in one of them.

Start with the Week 1 audit

Week one is a fixed-fee audit: two calls, and a one-pager you keep either way, naming every model and agent we can find, classified by risk, with the gaps put in order.