One AI agent, in production in eight weeks, or you do not pay for the build
We build one production agent on Anthropic's Claude that watches your live data, catches what is costing you money, recommends the action with evidence, and acts once you approve. If it has not made or saved real money in 90 days, the build is free.
- Agents in the first build
- One
- Demo to production
- 8 weeks
- Guarantee window
- 90 days
- Built on
- Anthropic's Claude
- Recovered for one retailer
- $240k
A chatbot answers. An agent acts.
Most things sold as AI reply to a prompt and forget it ever happened, which is a different product from the one described on this page.
- Answers the question you thought to askFinds the decision that is worth money
- Forgets the conversation the moment it closesKeeps an audit log of every call it made
- Runs only when a person opens itRuns in the background around the clock
- Reads whatever gets pasted into the boxWatches your live data where it already sits
- Ends at an answer on a screenActs in your tools once you approve
The loop the agent runs, around the clock
Five steps, and the example running through them is the retailer engagement behind the figures on this page.
- Watch
Your live data, without being asked
The agent sits on the systems that already hold cost, price, mix and volume. Nobody has to remember to open it, and nothing waits for a weekly review.
- Catch
The signal that is costing you money
At 06:12 it caught APAC logistics cost running 23% above plan. Not a chart that has drifted, a number with money attached to it.
- Recommend
The action, with evidence and confidence
Switch to Vendor B's RFP, saving $240k this quarter. It showed 94% confidence and cited three sources, so the claim could be checked before anyone acted on it.
- Act
In your tools, only on your approval
Adopt, hold, or hit the kill switch. The agent has an approval step in front of every action it can take, and a human decides which of the three it gets.
- Prove
The result, tracked in dollars
It follows the action it recommended and reports what the action was worth. That tracking is what makes the 90-day guarantee something we can offer at all.
One morning of that loop, in numbers
The same retailer catch, pulled out of the five steps above so the figures sit together.
Point it at the bleeding
Four jobs where one agent moves a number that Finance already watches, and a fifth door if yours is not on the list.
Margin you are leaking
Cost, price and mix watched in real time. The agent flags the money walking out the door before the month closes rather than after.
Getting the cost data straight →02Churn before it cancels
The accounts about to leave, scored and ranked, so your team acts while the save is still possible instead of reading about it afterwards.
Case studies →03Anomalies in operations
Faults, spikes and SLA breaches caught the moment they happen and pushed to the person who can fix them, not to a dashboard nobody opens.
Alerting and routing →04Back-office document work
Claude reads the invoices, contracts and forms, extracts what matters, and routes only the exceptions to a human. The routine cases stop reaching a desk.
Document automation →05Something not on this list
The pattern is the same wherever a decision is made late because nobody was watching. Tell us the job and we will say whether it fits this shape.
Tell us the problem →Live in eight weeks, fixed scope and fixed date
Four phases, and the scope stops moving at the end of week one.
- Week 1
Scope
Two calls and a one-pager you keep either way. We pick the one job with the clearest payoff and lock the scope around it.
- Weeks 2-4
Build
Built on Claude, wired to your data and to the one tool the agent is allowed to act in. One job, one tool, no sprawl.
- Weeks 5-6
Guardrails
The approval step, the confidence score, the cited evidence, the audit log and the kill switch. This is the half that decides whether it survives contact with your auditors.
- Weeks 7-8
Go live
In production, results tracked in money, your team trained to run it. From week 7 onward the work is operating it, not building it.
Where the guarantee sits on the calendar
The build takes eight weeks and the guarantee runs for ninety days from the day the scope is locked.
What we are actually promising
Four commitments, and each one is checkable rather than a matter of trust.
Measured
A dollar outcome inside 90 days, attributable to a specific action the agent recommended and tracked from that action to the result.
Defensible
Every answer cites its sources and carries a confidence score. The audit log records what was recommended, who approved it and what happened next.
On time
In production in eight weeks. If the date slips because of us, the extra weeks are on us rather than on your budget.
Under your control
The agent acts only on approval, in one tool, and there is a kill switch. Nothing is taken automatically because it seemed confident.
Which Claude model does which part of the work
One agent, three models behind it, chosen by what each step of the loop actually needs.
- Sonnet
Production reasoning
The everyday work of the loop. Reading the live signal, weighing it against plan, and deciding whether there is a recommendation worth putting in front of a person.
- Opus
The hard analyses
The cases where the answer is not obvious and the cost of being wrong is high. Used deliberately rather than by default, because it is the expensive option.
- Haiku
The fast long tail
High-volume, low-difficulty steps such as classification and extraction, where speed and cost per call matter more than depth of reasoning.
Where the eight weeks go
The four phases of the build, drawn to the length each one actually takes.
Where this is the wrong answer
Five cases where we would rather tell you in week one than take the build fee and find out in week five.
What you want is an assistant, not an agent
If the need is a place to ask questions and get good answers, an agent that watches and acts is heavier and more expensive than the question deserves. Say so and we will scope the smaller thing.
The data it would watch does not exist yet
An agent cannot watch a number that is assembled by hand each month. If cost, price and mix are not landing anywhere reliable, the first job is the pipeline, not the agent.
Nobody will own the approval step
Every action passes a human. If no named person has the authority and the time to approve or hold within hours, the recommendations will queue and the guarantee will fail on your side.
The money will not be visible inside 90 days
Long sales cycles and multi-quarter decisions make attribution honest but slow. If the outcome cannot be measured before day 90, our guarantee is meaningless and you should not buy on it.
You want ten agents in the first engagement
This is deliberately one job, one tool, one number. Breadth in the first build is how these projects end as a demo in a deck. The second agent is a separate decision.
The four figures this page rests on
Two of them come from a single retailer engagement rather than a portfolio average.
What buyers ask before signing
The five questions that come up in almost every first call.
What exactly does the guarantee cover?
The build fee. If the agent has not made or saved real money within 90 days, you do not pay for the build. Separately, if go-live slips past week 8 because of us, the extra weeks are on us rather than on your budget.
Is this the same as the six-week timeline mentioned elsewhere on the site?
No, they measure different things. Six weeks is the median time to a first production result across all our engagements generally, not a guarantee. This page is a narrower, specific offer: exactly one AI agent, on a fixed eight-week date, with the build-fee guarantee above attached. If eight weeks for one agent is not the shape of what you need, the general audit-then-build engagement model is on the home page.
How can you promise a dollar outcome when nobody else does?
Because the agent grades its own homework and shows the marking. It tracks the specific action it recommended through to the result, with the sources it cited and the audit log of who approved what. If the number cannot be traced, we cannot claim it, and that constraint is what makes the promise possible rather than reckless.
What happens after week 8?
From week 7 onward the engagement is operating rather than building: watching how the agent behaves on live data, tuning what it flags, and training your team to run it without us. The 90-day guarantee clock is running through this period, so the first outcome usually lands inside it.
We do not have our data in one place. Can we still start?
Sometimes, and sometimes not. If the numbers the agent would watch already land somewhere reliable, we can point it there. If they are assembled by hand each month, the pipeline is the real first project and an agent on top of it is premature. Our data engineering work covers that case, and the week 1 audit tells you which situation you are in.
Why only one agent?
Because one job with a clear payoff, one tool it can act in and one number it moves is what gets to production in eight weeks and can be measured in ninety days. Several agents at once is how the scope softens, the date slips and the outcome stops being attributable to anything. Once the first one is live and paying, the second is an easier conversation.
The rest of the practice
These are genuinely different jobs with different ways of failing. Most engagements start in one of them.
- Data engineeringJobs that finish, migrations that reconcile, and storage that stops paying for cold data.
- Data integration and governanceOne set of records the finance team and the operations team both accept.
- Data platforms and modernisationOff the platform you outgrew, without a twelve-month freeze on new reporting.
- Apache SupersetSuperset built and run by people who commit to the project, including embedding and Kubernetes.
- Command centresThe one screen an operations floor runs the day from, not another dashboard.
- AI audit and roadmapOne week, and you know which two or three AI initiatives are worth building.
- AI evaluationA graded eval set, because an AI system cannot simply pass or fail a test suite.
- AI governanceGovernance you can defend in a board meeting and audit on demand.
- AI agentsOne reasoning agent on Claude, on your live data, with evidence and a kill switch.
- Data agentsAgents that watch the data, flag what moved, and explain what is driving it.
- Applications and automationThe system your team works in all day, built or replaced in slices.
- Application modernisationThe system nobody wants to touch, replaced a slice at a time rather than rewritten.
- System integrationSystems that stop disagreeing about the same customer, order and item.
- Process digitisationThe process that still runs on paper, WhatsApp and one shared spreadsheet.
Start with the Week 1 audit and find out whether an agent is the right build at all
Two calls, a fixed fee, and a one-pager you keep either way: the one job with the clearest payoff, the data it would watch, and the tool it would act in. If your pipeline comes first, we will say so.