Four cases, four disciplines, no result used twice
Each engagement below is filed under the one discipline that actually produced it. Applications and automation has no published case yet, and we say so rather than borrow one.
The four published cases, each labelled with the discipline behind it
Read the label before the number, because a result only tells you something if it came out of the discipline you are buying.
Fintech, Series C: a dashboard estate cut to eight
Six weeks. Two hundred and forty dashboards down to eight people actually open, load time from 15 seconds to 3, and $72k saved in year one. Definitions agreed once, queries tuned, a smaller surface that stays correct.
Read the case →AI$240kAPAC grocery, $500M+ revenue: a margin agent that flagged a leak on day 14
Six weeks end to end. The agent flagged a margin leak on day 14, $240k was recovered, margin moved 2.4 percentage points, and the customer's own Finance team signed the number off within 30 days.
Read the case →AI governanceZero audit gapsPublic sector citizen services: eight agents live, rolled out the same week
Eight weeks to eight agents in production, rolled out the same week they were cleared, with zero audit gaps found. The controls were lineage, policy checks, confidence scores, a kill switch and an on-call pod.
Read the case →Analytics inside an operating process$480k ARR heldB2B SaaS: churn signals joined across three systems
Churn signals sat in three separate systems and nobody looked at all three. Joining them was the straightforward half. Putting the result into an account manager's Monday morning was the work. CSAT 4.6 to 4.8, $480k of ARR held.
Read the case →The result that kept getting misused
One engagement carries one claim, a rule that costs us a named client story on several service pages, because four results that turn out to be the same job are worth less than one plainly stated gap.
Where each case is filed, and the gap we are not filling
Four disciplines on the left, the evidence each one is allowed to claim on the right.
- AnalyticsFintech Series C, 240 dashboards to 8B2B SaaS, churn signals and CSAT
- AI and evaluationAPAC grocery margin agentPublic sector, eight agents live
- Data engineeringAPAC grocery margin agent
- Applications and automationNo published case yet
Three of the four cases carry a money figure
The dollar results from the cards, drawn at the same scale so you can see how differently sized they are.
What has to be true before a figure is allowed on this page
Four tests, and a result that fails any of them does not get published however good it looked internally.
Someone outside Woodfrog verified the number
For the APAC case that was the customer's own Finance team, within 30 days of the first flag. We do not publish a saving that only our own model believes in.
The figure carries its time window
A number without its window is a claim rather than a result, so every figure on this page is published with the period it covers.
The discipline is named on the card
So you can tell in five seconds whether this is evidence about the thing you are buying, or about a neighbouring practice in the same firm.
The client stays anonymous, the reference does not
Cases are written to be NDA-friendly, which is the condition on which we were allowed to write them. Attributed references are arranged on the audit call instead.
All four cases ran on the same three phases
Different disciplines, same three phases, which is what makes the results comparable to each other.
- Week 1
Audit
Fixed fee, two calls, and a one-pager saying what we would do and what we think it is worth. You keep it whether or not the engagement continues. All four cases began here.
- Weeks 2 to 6
Ship
A pod of three, with something working in your hands by Week 3. In the APAC engagement that working thing raised its first real flag on day 14.
- Week 7 onwards
Operate
Quarterly reviews and on-call governance. The public sector controls, kill switch included, live in this phase. They are an operating commitment, not a slide shown once during a pitch.
Which case to read first, depending on what you are trying to settle
Pick by the question you are actually stuck on, not by the industry that matches yours.
- If the question is reporting
Read the fintech case
You have too many dashboards, they disagree with each other, and they are slow. The case covers what got cut, what survived, and why eight was the right place to stop.
- If the question is trust near money
Read the APAC case
An agent raised a margin flag on day 14 and a Finance team accepted it inside 30 days. Most of the case is about why they accepted it: the lineage built before the agent.
- If the question is your auditor
Read the public sector case
Eight agents live, cleared and rolled out in the same week, zero audit gaps. The interesting part is the five controls that made a same-week rollout possible at all.
- If the question is adoption
Read the B2B SaaS case
Signals joined across three systems are worth nothing until they land in someone's weekly routine. CSAT moved 4.6 to 4.8 in a quarter because the routine changed, not the model.
The figures, each with the period behind it
These are the four numbers the cases turn on, shown with the window they were measured over.
Fair objections to a page like this one
The ones we get asked on the first call, answered here instead.
Four cases is not very many.
It is not. More than 20 companies have worked with us; four agreed to be written up in this detail with an outside team checking the figure. We would rather be short than padded, and you can ask about the rest on a call.
Every one of these is anonymous, so how do I know they are real?
You do not, from a web page, and you should not pretend otherwise about anyone's site. Attributed references are arranged on the audit call under NDA. Anonymity is the condition under which these were allowed to be written at all.
You have nothing published for applications and automation.
Correct, and that pillar is sold on method and capability until a client agrees to be named. Quietly dropping the fintech dashboard result into that page would be the easiest thing to do and the fastest way to lose a serious buyer.
Are these your best results or your normal ones?
The median across more than 20 companies is six weeks to a first production-grade artefact and $240k recovered or saved. The APAC case sits on that median rather than above it. The fintech and public sector cases are stronger on their own axes.
Could you repeat any of this for us?
We do not know until Week 1, and anyone answering before they look at your systems is guessing. The audit is fixed fee, two calls, and ends in a one-pager you keep, including when the honest finding is not to proceed.
Which model sits behind the AI cases?
We build directly on Claude as part of the Anthropic Claude Partner Network, with no reseller layer in between. Sonnet handles production reasoning, Opus the hard analyses, Haiku the cost-sensitive long tail. The choice is an engineering decision we defend, not a badge.
Read the nearest case, then bring us what you do not believe
If one of the four looks like your situation, read it properly and bring the part you find hardest to accept. Week 1 is a fixed fee audit: two calls, a one-pager you keep either way, and no obligation past it. Write to hello@woodfrog.tech. If none of them look like your situation, that is worth saying out loud too, because it usually means the first week is about something we have not written up yet.