AI agents

One agent that watches your business and acts on what matters

We build a single reasoning agent on Claude, running on your live data, in production in eight weeks. Every signal it raises carries its evidence, a confidence score and a kill switch, and a human decides whether it acts.

Demo to production
8 weeks
Terms
Fixed scope, fixed price, fixed date
No dollar outcome in 90 days
The build is free
Model layer
Claude Sonnet, Opus and Haiku
Team on the engagement
A pod of three

What happens between a number moving and something being done

This is the path of one real signal, from the agent noticing a cost move to a person deciding.

  1. Watch

    The agent reads live data continuously

    Not a scheduled report. It monitors the sources you already run and holds the definitions your business agreed, so a move it flags means the same thing to Finance and to Operations.

  2. Detect

    Delivered cost in APAC is up 12.4 per cent

    Week 18. The agent raises it as a P1 signal rather than waiting for someone to open a dashboard and notice.

  3. Explain

    Six sources, 91 per cent confidence

    The signal arrives with the evidence behind it and the projected margin impact, minus 2.4 points. A number you cannot trace does not survive its first meeting with Finance.

  4. Recommend

    Trigger a logistics RFP with DHL and Yamato

    The agent proposes the action, not just the observation. Estimated saving $240k, first-year return nine times.

  5. Decide

    Adopt, dismiss, or pull the kill switch

    A person approves before anything runs. The kill switch is one click, on every signal, at every point.

What the agent is built to do

Six behaviours, in the order they usually get switched on.

01Watch

Track performance metrics in real time

Continuous monitoring at production volume, up to 2,847 events a second on the reference agent.

02Warn

Flag a signal that falls outside the norm

As it happens, rather than at the next reporting cycle, and with a severity attached to it.

03Explain

Reveal what is driving a trend

The agent explores the data itself and surfaces patterns nobody thought to build a chart for.

04Advise

Guide you toward the effective decision

A recommendation with the evidence and the confidence attached, in plain language, with no technical jargon.

05Act

Launch operations or campaigns

Purchase orders, alerts, campaign adjustments and workflows, each one behind a human approval step.

06Talk

Hold a normal conversation

Ask it a question the way you would ask a colleague who happens to be connected to your data.

Where agents meet your internal tools

Figures from the reference agent

These come from one engagement rather than an average across clients, and we would rather say so than dress them up.

$240kRecovered in 90 daysFrom a single signal on delivered cost, acted on inside the quarter it appeared.
First-year returnThe first-year figure reported on the reference engagement, from the same signal that recovered the $240k.
91%Confidence on the signalReported on the signal itself, with the six sources it was drawn from listed alongside it.
2,847/sEvents monitoredThe agent does not sleep and does not sample. Continuous monitoring is what makes it proactive rather than another report.

Eight weeks, and what happens in each of them

The shape is the same as every Woodfrog engagement, compressed to one agent.

  1. Week 1

    Fixed-fee audit

    Two calls and a one-pager you keep either way. It names the one decision the agent should serve and what your data can honestly support today.

  2. Weeks 2-6

    Build

    A pod of three connects the sources, writes the definitions, builds the agent and wires the approval and kill-switch path.

  3. Week 7

    Operate

    The agent runs on live data with humans on every action. Signals get adopted or dismissed, and the dismissals are the useful ones.

  4. Week 8

    In production

    Demo to production in eight weeks, on the fixed scope, fixed price and fixed date agreed in week one.

  5. Day 90

    The outcome is measured

    If the agent has not produced a measured dollar outcome by then, you do not pay for the build.

The agent gets more accurate because you use it

Accuracy on the reference agent over its first eight weeks, as it learned the questions the business actually asks.

Week 162%
Week 478%
Week 885%
Bars are answer accuracy in percent on the reference agent. Week 1 is the honest number: an agent is not right on day one, and any vendor whose first week looks like week eight is showing you a demo.

Fewer reporting surfaces, not more

The dashboard audit on the reference engagement retired three surfaces and improved one.

4 Four reporting surfaces maintained1 One kept, agent-assisted
Sales Overview v3 and the Weekly KPI Tracker were replaced by the agent, the legacy Ops Report was deprecated, and Finance Monthly stayed and got help. Zero reports missed.

What keeps an acting agent safe to run

An agent that can trigger a purchase order needs more governance than one that can only draw a chart.

  • Evidence on every signal

    The sources behind a recommendation are listed with it. Six sources on the APAC signal, and you can read each of them.

  • A confidence score you can act on

    Stated on the signal, not buried. A high-confidence signal and a marginal one should not be treated the same way, so the agent does not present them the same way.

  • Lineage and an audit trail

    Any figure the agent quotes walks back to the record the source system sent. This is what makes a recommendation defensible after the fact.

  • Human approval before action

    Adopt or dismiss sits with a person. The agent proposes, and nothing runs on your systems until someone with the authority says so.

  • A kill switch one click away

    On the signal and on the agent, at any time. If the behaviour drifts, you stop it yourself rather than raising a ticket with us.

  • A governed data layer underneath

    We build the data infrastructure the agent stands on, because plugging a model into ungoverned data is how you get confident wrong answers.

These are capabilities we build, not certifications we hold. Nothing here should be read as SOC 2, ISO 27001, HIPAA or GDPR certification, and we claim none. If a certification gates your procurement, ask us before you shortlist.

What one signal carries with it

The APAC cost signal, broken into the four things that arrive alongside it.

6Sources behind the signalListed with the recommendation, and you can read each of them. A number you cannot trace does not survive its first meeting with Finance.
12.4%Delivered cost rise in APACRaised in week 18 as a P1 signal, rather than waiting for someone to open a dashboard and notice it.
-2.4 ptsProjected margin impactThe projected impact travels with the signal, so the conversation is about the action rather than about the arithmetic.
1 clickKill switch, every signalOn the signal and on the agent, at any time. If the behaviour drifts, you stop it yourself rather than raising a ticket with us.

Where an agent is the wrong answer

There are cases where we would rather tell you in week one than take the build fee.

No decision is waiting on the other side of the answer

An agent earns its keep by shortening the distance between a signal and an action. If nobody would act differently on the answer, you are buying a faster way to feel informed, and a dashboard does that more cheaply.

Your data layer cannot support it yet

An agent is only as strong as the foundation it is built on. If your sources disagree and nobody owns the definitions, the honest first project is the data work, not the agent. That is a different engagement and it starts at /data-engineering.

Nobody in the business will accept a recommendation from software

Human approval and a kill switch are the right controls, but they only help if a person is willing to press adopt. If every action would go to a committee anyway, you have bought a slower report with extra steps.

You want a general assistant rather than one agent

This build is deliberately narrow: one agent, one decision, one measurable outcome in 90 days, described at /ai-use-case. If what you want is a broad copilot across everything, the guarantee cannot apply and we would not offer it.

You need a reasoning layer that is not Claude

We are an Anthropic Build Partner and we lead with Claude, choosing between Sonnet, Opus and Haiku by workload. We integrate the rest of your stack around it, but the reasoning layer is Claude. If that is a constraint, better you know now.

Three departments asking one thing

The silo is rarely the data, it is that each team has its own route to it.

  1. Finance: what was our burn rate in Q3?One agent, in natural conversation
  2. Operations: show me warehouse utilisation trendsOne agent, in natural conversation
  3. Marketing: which campaign drove the most sign-ups?One agent, in natural conversation
Three questions, one agent, one source of truth. The value is not that each team gets an answer, it is that the three answers agree with each other.

What buyers ask before they commit

The five questions that come up in almost every first call.

What exactly does the guarantee cover?

One reasoning agent, on your live data, in production in eight weeks, on a fixed scope, a fixed price and a fixed date. If it has not produced a measured dollar outcome within 90 days, you do not pay for the build. What counts as that outcome is agreed in week one during the audit and written down, so there is no argument at day 90 about the definition.

Which models does it run on, and what about the rest of our stack?

We lead with Claude and pick between Sonnet, Opus and Haiku per workload, as an Anthropic Build Partner and a member of the Claude Partner Network at launch. Around that we integrate the orchestration layer you already run, including Azure, Databricks, Snowflake and the BI tools your team trusts. Right model, right place, every time.

Do we need to build new pipelines first?

Usually not. The agent works on the data infrastructure you already have, which is why the first useful answers arrive in minutes rather than after a two to three week engineering sprint. Where the existing layer genuinely cannot support a trustworthy answer, we say so in the Week 1 audit instead of building on top of it and hoping.

How do you know the agent is right?

Every signal carries its evidence, its sources and a confidence score, and every number traces back through lineage to the record the source system sent. Beyond that, the agent is tested the way we test any model going into production. The method is set out at /ai-evaluation, and it is the part of this work we would want you to interrogate hardest.

What happens after week eight?

From week seven onward the engagement is operate: the agent runs, signals get adopted or dismissed, and accuracy improves as it learns the questions your teams actually ask, from 62 per cent in week one to 85 per cent by week eight on the reference agent. You can keep us on to run it or take it over. It is your agent, on your infrastructure. Past work is at /case-studies.

Eight weeks to production, ninety days to the measure

The build date and the measurement date sit inside one window, and the stretch between them is where the agent has to earn the fee.

Week 1, the fixed-fee auditDay 90, the outcome is measured
The window is measured in days, from the Week 1 audit to the day 90 measure, with week eight marked at day 56. If there is no measured dollar outcome at the end of it, you do not pay for the build.

The rest of the practice

These are genuinely different jobs with different ways of failing. Most engagements start in one of them.

Start with the audit, not the build

Two calls, a fixed fee, and a one-pager you keep either way: the one decision it should serve, what your data can support today, and the measured outcome at day 90. If an agent is not the right first move, we will say so.