The ERP says shipped, the CRM says pending, and the shop floor is still building it
Operational integration between the systems people type into all day, not analytical pipelines. We settle in writing which system is allowed to be right about each field, choose honestly between point to point and a middle layer, make redelivery of the same message harmless, and put failures in a queue a named person works rather than a log nobody opens. Upstream changes then break a test instead of a customer.
- Companies delivered for
- 20+
- Median to first production-grade artefact
- 6 weeks
- Software a real user can open
- By week 3
- Team on a build
- Pod of three, no bench
- Published case study for this pillar
- None yet
On this pageWhen the systems stop agreeing
When the systems stop agreeing
Signs the systems have stopped agreeing about the same record
Nobody books a call about integration. They book a call about one order that three systems describe differently, and a customer who has noticed.
- 01The ERP, the CRM and the shop floor each hold a different version of the same order
Sales quotes one status, production works from another, finance invoices from a third. Everybody is reading their own screen correctly. The screens disagree, and the argument gets settled by whoever is most senior in the room.
- 02A duplicate went out and nobody can say which system sent it
The same invoice, the same dispatch note, the same notification, twice. The retry that produced it looked like a success in every log you have, because from the sending side it was one message, sent once.
- 03The integration works until the upstream team ships on a Friday
A field is renamed, a code list gains a value, a payload gains a level of nesting. Nothing was announced, because from where that team was standing nothing looked breaking.
- 04The file arrived but was empty, and everything downstream treated that as zero
Zero sales, zero stock, zero open jobs. The job reported success, no alert fired, and the first person to notice was outside the company.
- 05A person re-keys the same record into two systems every morning
Somebody's first hour is spent being a message bus. When they are on leave the two systems drift, and the drift is found weeks later by somebody else, in a reconciliation.
- 06Failures go to a log, and the log has no owner
There is a queue somewhere with failed messages in it. Nobody's job description includes reading it, so it gets read after an escalation rather than before one.
Three or more of these and the problem is not a missing connector. It is that nothing between the systems has an owner, a contract, or a way to fail loudly.
Each symptom is a missing mechanism, not a missing connector
Connectors are the part that is easy to buy. What is usually missing is one of these mechanisms, and the symptom tells you which one.
Left is what the business reports. Right is the mechanism that was never built. The row you recognise tells you which mechanism to ask about first.
What one message has to survive
What one message has to survive on the way across
Integration reads as plumbing until you follow a single record end to end. Each of these five points is somewhere it can be lost, duplicated or quietly mangled.
- 01Emitted
Something happened in the source system. The first question is whether that system can tell you it happened at all, or whether you are going to have to go and look on a schedule.
- 02Carried
A queue, a topic, an API call, or a file on a server. This is the part integration products are sold on, and it is the part with the fewest decisions in it.
- 03Understood
The payload is checked against a contract before anything acts on it. A field that changed shape is rejected here, where it is cheap, rather than three systems later where it is not.
- 04Applied
The receiving system changes. This is where a second delivery of the same message either does nothing or creates a duplicate invoice, and which of those happens was decided months earlier.
- 05Acknowledged, or parked
Success is recorded so the work is not repeated. Failure goes somewhere a person will actually look, with the payload attached, and can be replayed once the cause is fixed.
Points 03, 04 and 05 are where a message gets quietly mangled rather than visibly lost. Point 02 is the part that is easy to buy, and buying it settles none of the other four.
The rules that stop the second delivery becoming a duplicate invoice
Retries are not optional, because networks fail halfway through an operation. Everything on this list exists so that a retry is boring.
- 01Every message carries a key the receiver recognises
The receiving system records what it has already applied and matches on that key. Applying the same message twice then changes nothing, which is the property that makes automatic retries safe in the first place.
- 02Retries back off, and they stop
An immediate retry loop against a system that is already struggling is an outage you caused. Delays widen, attempts are capped, and whatever is left after the cap goes to the error queue rather than round again.
- 03Ordering is stated, not assumed
Two updates to the same order can arrive in the wrong order. Either the receiver refuses anything older than what it already holds, or we write down plainly that this flow does not need ordering, and why.
- 04Late arrivals have a decided destination
The record that turns up after the period closed goes somewhere agreed in advance, with somebody told. Silently dropping it and silently applying it are both wrong, and the second is worse, because it moves a number that was already signed off.
- 05Empty and partial count as failures until proved otherwise
A file with no rows, a response with an empty list, a payload cut short. Each needs a check before anything downstream reads it, because zero is a value the rest of your estate will happily believe and act on.
- 06One bad message cannot stop the queue
A single malformed record should not hold up everything behind it. It gets parked, the queue keeps moving, and the parking is visible on a screen rather than inferred from a gap in the numbers.
This list gets written for your specific flows during the Week 1 audit. Two calls, a fixed fee, and you keep the one-pager whether or not the rest of the work goes ahead.
What happens when a message fails, in order
The difference between an integration people trust and one they work around is almost entirely what happens after something fails.
- 01It fails loudly, in one place
The failure is recorded with the payload, the destination, the attempt count and the reason. Not a stack trace in a log file. A row somebody can read without being an engineer.
- 02It is retried, within written limits
Transient failures resolve themselves. The retry policy is written down per flow, so nobody has to guess whether a message is still in flight or gone.
- 03It lands in an error queue with a named owner
Not a team. A person, with a deputy for when they are away. An error queue with no name against it is a log with extra steps and a nicer interface.
- 04Somebody is told, on a channel they actually watch
The alert says which flow, which record, and what a human is expected to do about it. Alerts that report only that something went wrong get muted, and then the queue is unwatched again.
- 05It is fixed, then replayed
Replay is built at the same time as the flow, not added after the first incident. Reprocessing is safe precisely because of the idempotency rules above.
- 06The cause goes back into the contract
If a payload shape broke the flow, the contract test is updated so the same shape breaks a build next time instead of breaking a customer.
Choosing the shape
Point to point, or a middle layer, and when each is honestly right
Both answers are defensible. The wrong one is expensive in a way that takes a long time to become visible.
- Point to pointTwo systems, one direction, a stable contractOften right
If you have two systems, one flow, and no plan to add a third, a direct connection is the honest answer. Fewer moving parts, easier to debug at seven in the morning, and a middle layer here would be a tax with no service attached.
- Point to pointWhen the connections start multiplyingStops being right
Each new system adds connections rather than one connection. The tell is not the count. It is the day somebody asks which flows would break if a field were renamed, and nobody can answer without opening code in four repositories.
- Middle layerWhen you need one place to see what movedEarns its place
A broker, a hub or a small integration service gives you routing, retries, replay, and one screen where every message can be seen. That visibility is usually the real reason to build it. The routing is the easy part.
- Middle layerIt is a system, and systems need ownersThe cost
It needs deployment, monitoring, someone on call and a version story of its own. We will say plainly when an estate is not large enough to pay that back, because an unowned hub is worse than the point-to-point flows it replaced.
Events, batch, or the nightly file drop nobody wants to touch
The question is never which of these is modern. It is how stale the receiving system is allowed to be, and who is harmed while it is.
- EventsWhen staleness costs something measured in minutes
Stock sold twice, a credit limit checked after the order was taken, a dispatch that leaves before the cancellation lands. Events are worth their complexity when the gap between two systems is where the money leaks. They also bring ordering and duplicate handling with them, and that is real work rather than a checkbox.
- BatchWhen the receiving system only acts once a day anyway
If the finance run happens overnight, a stream feeding it changes nothing except your operational burden. A well-instrumented batch job, with a completion check and an alert when it does not finish, is frequently the correct answer and a great deal cheaper to run and to staff.
- The file dropThe nightly file that has run for years
It is easy to be rude about it. It has also run for years, the operations team knows its shape, and the person who wrote it may have left the industry. We do not remove it because it is old. We put an arrival check, a row count check and a named owner around it, and we replace it when there is a reason beyond taste.
This gets decided per flow, not per estate. An estate that ends up running all three at once is a correct outcome rather than a compromise.
The system of record decision is yours, and it cannot be skipped
This is the one part of an integration project that is not a technical decision, and it is the part that stalls most of them.
For every field that lives in more than one system, one system has to be allowed to be right. Not usually right. Right, so that when two values disagree there is no discussion about which one wins.
Until that is decided, an integration can only move the disagreement around faster than before.
Whether the warehouse system or the ERP owns stock on hand is a decision about how your business runs and who is accountable when the figure is wrong. We can set out the consequences of each answer, show what breaks either way, and write down what we would choose and why. Signing it is yours.
The other copies stop being second opinions. They go read-only, or they are switched off, or they are kept and clearly marked as a copy with a staleness people can see on the screen.
What we will not do is leave two systems both able to write the same field and hope the flows keep them level. That is the arrangement that produced the problem you called about.
Contract testing
The most common way a working integration stops working is that somebody upstream made a change they had every reason to believe was safe.
- The problemNobody upstream knows you are there01
The team that owns the source system has a roadmap and no list of who consumes their payloads. From where they stand, adding a field or tightening a validation is not a breaking change. From where you stand it can be.
- The mechanismThe contract is written down and executed02
The shape each side depends on becomes a test that runs in both builds. Required fields, types, code lists, and what is allowed to be absent. It is not a document, because documents do not fail a pipeline.
- The effectThe break moves to where it is cheap03
The change fails a build on a Tuesday afternoon, with a named owner and the reason attached, instead of arriving on a Friday evening and being discovered by whoever answers the phone on Monday.
- The limitWhere the upstream is not yours to change04
Third-party APIs and vendor systems will not run your tests. There we check payloads against the contract in flight, quarantine what no longer matches, and alert instead of quietly coercing it into shape. That is a weaker guarantee, and we say so rather than imply otherwise.
How the first six weeks run
What the first six weeks look like on an integration
Three phases, always. The interesting decision is which flow gets picked first.
- Week 1The audit, fixed fee
Two calls, and one real record followed end to end across every system that touches it. You get a one-pager naming the flows, who owns which field, the mechanisms that are missing and the order to fix them in. Yours whether or not we go further.
- Week 2One flow is chosen, and it is not the easy one
We take the flow that carries real transactions and causes real arguments, because that is where the payback shows and where the team will notice the difference. The demo-friendly flow teaches you nothing about whether the design holds.
- Week 3Something runs
The flow moves real records, with its contract test, its error queue and its replay path attached. This is the week a wrong assumption about ordering or ownership surfaces, when correcting it costs a conversation rather than a rebuild.
- Weeks 4 to 6It runs alongside, then takes over
The new flow runs beside the current path and the outputs are compared before anyone depends on the result. Switching is then a configuration change rather than a weekend and a war room.
- Week 7 onwardsOperate
Quarterly reviews and on-call governance. Somebody answers when the error queue fills at seven in the morning, and the fix goes back into the contract rather than into a workaround.
Why a pod of three moves quickly on this
Integration work has always been mostly reading: undocumented payloads, exception branches nobody wrote down, and the reason one field carries two meanings. That reading is the part that has changed.
- Already builtPatterns that have run in production
Idempotent receivers, retry and back-off policies, error queues with replay, arrival checks and contract test harnesses. Fitted to your systems rather than written again from an empty repository.
- AI first passReading payloads nobody documented
Claude reads sample messages, legacy interface code and the exception handling, then drafts the candidate contract and flags what it could not explain. The flags are where we start asking questions.
- AI first passDrafting the tests alongside the flow
Contract tests and reconciliation checks are generated from the agreed contract, so the thing that proves the integration is built with it rather than after it, when nobody has budget left.
- Human callOwnership, ordering and what is allowed to be quiet
Which system owns which field, whether a flow needs ordering, what a late arrival does, and which failures may pass without waking anyone. Those decisions make an integration trustworthy, and they stay with the engineer.
- What you keepYours to run without us
Contracts, tests, runbooks and the error queue, in your repositories. We are an Anthropic Build Partner and joined the Claude Partner Network at launch, so we build on Claude directly with no reseller sitting in between.
What we have delivered
There is no named integration case study on this site yet, so these are firm-wide delivery figures. We will say on the call which of them came from work that looked like yours.
Founded in 2023, based in Pune, delivering across India, APAC, the UK, Europe and the US.
Not a diagram of the target architecture. A flow moving real records with its tests attached.
Early enough that your team can disagree with us while disagreeing is still cheap.
The people on the calls are the people writing the flows and carrying the pager.
Questions operations and platform teams ask
The ones that come up on almost every first call about this work.
Further reading
- Some questions need a button, not another chartA report can tell you fourteen orders are stuck. It cannot let anyone unstick them, which is why the work quietly moves into a spreadsheet you never see.
- You cannot automate a process nobody has written downThe flowchart you were handed describes the good day. The two days a week the job actually takes are spent on everything the flowchart leaves out.
- Most of the work you want to automate does not need a modelA model earns its place on three conditions. A great deal of what gets scoped as an AI project fails all three, and a form, a rule and a lookup would do it better.
Start with the Week 1 audit
Two calls, a fixed fee, and one real record followed end to end across every system that touches it. You get a one-pager naming the flows, who owns which field, and the order to fix things in. It is yours whether or not we build anything. NDA-friendly, fixed scope. hello@woodfrog.tech, Pune.
