AI agent build

One AI agent, in production in eight weeks, or you do not pay for the build

We build one production agent on Anthropic's Claude that watches your live data, catches what is costing you money, recommends the action with evidence, and acts once you approve. If it has not made or saved real money in 90 days, the build is free.

Agents in the first build
One
Demo to production
8 weeks
Guarantee window
90 days
Built on
Anthropic's Claude
Recovered for one retailer
$240k

A chatbot answers. An agent acts.

Most things sold as AI reply to a prompt and forget it ever happened, which is a different product from the one described on this page.

  1. Answers the question you thought to askFinds the decision that is worth money
  2. Forgets the conversation the moment it closesKeeps an audit log of every call it made
  3. Runs only when a person opens itRuns in the background around the clock
  4. Reads whatever gets pasted into the boxWatches your live data where it already sits
  5. Ends at an answer on a screenActs in your tools once you approve
Left, the assistant pattern. Right, what we build. If the left column is what you actually want, say so in the first call and we will tell you plainly that it is a smaller and cheaper piece of work.

The loop the agent runs, around the clock

Five steps, and the example running through them is the retailer engagement behind the figures on this page.

  1. Watch

    Your live data, without being asked

    The agent sits on the systems that already hold cost, price, mix and volume. Nobody has to remember to open it, and nothing waits for a weekly review.

  2. Catch

    The signal that is costing you money

    At 06:12 it caught APAC logistics cost running 23% above plan. Not a chart that has drifted, a number with money attached to it.

  3. Recommend

    The action, with evidence and confidence

    Switch to Vendor B's RFP, saving $240k this quarter. It showed 94% confidence and cited three sources, so the claim could be checked before anyone acted on it.

  4. Act

    In your tools, only on your approval

    Adopt, hold, or hit the kill switch. The agent has an approval step in front of every action it can take, and a human decides which of the three it gets.

  5. Prove

    The result, tracked in dollars

    It follows the action it recommended and reports what the action was worth. That tracking is what makes the 90-day guarantee something we can offer at all.

One morning of that loop, in numbers

The same retailer catch, pulled out of the five steps above so the figures sit together.

06:12When the signal was caughtNobody opened anything. The agent was already watching the live data, so the catch did not wait for a weekly review.
23%APAC logistics cost above planNot a chart that had drifted. A number with money attached to it, raised while the quarter was still open.
94%Confidence on the recommendationSwitch to Vendor B's RFP. The score sits beside the recommendation so a person can weigh it before approving.
3Sources cited with itThe claim came with its evidence attached, so it could be checked before anyone acted on it.

Live in eight weeks, fixed scope and fixed date

Four phases, and the scope stops moving at the end of week one.

  1. Week 1

    Scope

    Two calls and a one-pager you keep either way. We pick the one job with the clearest payoff and lock the scope around it.

  2. Weeks 2-4

    Build

    Built on Claude, wired to your data and to the one tool the agent is allowed to act in. One job, one tool, no sprawl.

  3. Weeks 5-6

    Guardrails

    The approval step, the confidence score, the cited evidence, the audit log and the kill switch. This is the half that decides whether it survives contact with your auditors.

  4. Weeks 7-8

    Go live

    In production, results tracked in money, your team trained to run it. From week 7 onward the work is operating it, not building it.

Where the guarantee sits on the calendar

The build takes eight weeks and the guarantee runs for ninety days from the day the scope is locked.

Day 0, scope lockedDay 90, the guarantee falls due
Eight weeks is fifty-six days. If the agent has not made or saved real money by day 90, you do not pay for the build. If go-live slips past week 8, the extra weeks are on us.

What we are actually promising

Four commitments, and each one is checkable rather than a matter of trust.

  • Measured

    A dollar outcome inside 90 days, attributable to a specific action the agent recommended and tracked from that action to the result.

  • Defensible

    Every answer cites its sources and carries a confidence score. The audit log records what was recommended, who approved it and what happened next.

  • On time

    In production in eight weeks. If the date slips because of us, the extra weeks are on us rather than on your budget.

  • Under your control

    The agent acts only on approval, in one tool, and there is a kill switch. Nothing is taken automatically because it seemed confident.

The guarantee covers the build fee. The $240k and the 9x figure come from one retailer engagement, not an average across clients, and we would rather you read them that way.

Which Claude model does which part of the work

One agent, three models behind it, chosen by what each step of the loop actually needs.

  1. Sonnet

    Production reasoning

    The everyday work of the loop. Reading the live signal, weighing it against plan, and deciding whether there is a recommendation worth putting in front of a person.

  2. Opus

    The hard analyses

    The cases where the answer is not obvious and the cost of being wrong is high. Used deliberately rather than by default, because it is the expensive option.

  3. Haiku

    The fast long tail

    High-volume, low-difficulty steps such as classification and extraction, where speed and cost per call matter more than depth of reasoning.

Woodfrog is an Anthropic Build Partner and was in the Claude Partner Network at launch. That means we are opinionated about model choice per step, not that any particular model is endorsed for your use case.

Where the eight weeks go

The four phases of the build, drawn to the length each one actually takes.

Week 1, scope1 week
Weeks 2-4, build3 weeks
Weeks 5-6, guardrails2 weeks
Weeks 7-8, go live2 weeks
Bar length is weeks of the eight-week build. Two of those eight weeks go to the approval step, the confidence score, the evidence, the audit log and the kill switch.

Where this is the wrong answer

Five cases where we would rather tell you in week one than take the build fee and find out in week five.

What you want is an assistant, not an agent

If the need is a place to ask questions and get good answers, an agent that watches and acts is heavier and more expensive than the question deserves. Say so and we will scope the smaller thing.

The data it would watch does not exist yet

An agent cannot watch a number that is assembled by hand each month. If cost, price and mix are not landing anywhere reliable, the first job is the pipeline, not the agent.

Nobody will own the approval step

Every action passes a human. If no named person has the authority and the time to approve or hold within hours, the recommendations will queue and the guarantee will fail on your side.

The money will not be visible inside 90 days

Long sales cycles and multi-quarter decisions make attribution honest but slow. If the outcome cannot be measured before day 90, our guarantee is meaningless and you should not buy on it.

You want ten agents in the first engagement

This is deliberately one job, one tool, one number. Breadth in the first build is how these projects end as a demo in a deck. The second agent is a separate decision.

The four figures this page rests on

Two of them come from a single retailer engagement rather than a portfolio average.

$240kRecovered in 90 daysMargin saved for one retailer in the first quarter after go-live, with the lineage to prove every dollar of it.
8 wksDemo to productionScope in week 1, build to week 4, guardrails to week 6, live by week 8. Our median to a first production-grade artefact is six weeks.
9xFirst-year ROIFrom that same retailer engagement. Your number depends entirely on the job you point the agent at, which is what week 1 exists to decide.
90 daysGuarantee windowIf nothing real has been made or saved by the end of it, the build costs you nothing.

What buyers ask before signing

The five questions that come up in almost every first call.

What exactly does the guarantee cover?

The build fee. If the agent has not made or saved real money within 90 days, you do not pay for the build. Separately, if go-live slips past week 8 because of us, the extra weeks are on us rather than on your budget.

Is this the same as the six-week timeline mentioned elsewhere on the site?

No, they measure different things. Six weeks is the median time to a first production result across all our engagements generally, not a guarantee. This page is a narrower, specific offer: exactly one AI agent, on a fixed eight-week date, with the build-fee guarantee above attached. If eight weeks for one agent is not the shape of what you need, the general audit-then-build engagement model is on the home page.

How can you promise a dollar outcome when nobody else does?

Because the agent grades its own homework and shows the marking. It tracks the specific action it recommended through to the result, with the sources it cited and the audit log of who approved what. If the number cannot be traced, we cannot claim it, and that constraint is what makes the promise possible rather than reckless.

What happens after week 8?

From week 7 onward the engagement is operating rather than building: watching how the agent behaves on live data, tuning what it flags, and training your team to run it without us. The 90-day guarantee clock is running through this period, so the first outcome usually lands inside it.

We do not have our data in one place. Can we still start?

Sometimes, and sometimes not. If the numbers the agent would watch already land somewhere reliable, we can point it there. If they are assembled by hand each month, the pipeline is the real first project and an agent on top of it is premature. Our data engineering work covers that case, and the week 1 audit tells you which situation you are in.

Why only one agent?

Because one job with a clear payoff, one tool it can act in and one number it moves is what gets to production in eight weeks and can be measured in ninety days. Several agents at once is how the scope softens, the date slips and the outcome stops being attributable to anything. Once the first one is live and paying, the second is an easier conversation.

The rest of the practice

These are genuinely different jobs with different ways of failing. Most engagements start in one of them.

Start with the Week 1 audit and find out whether an agent is the right build at all

Two calls, a fixed fee, and a one-pager you keep either way: the one job with the clearest payoff, the data it would watch, and the tool it would act in. If your pipeline comes first, we will say so.