Redefining Technology

LogisticsFuture of AI & Visionary Thinking

Supply fusion in logistics: merging demand, capacity, transit, inventory, finance and risk into one estimate

Supply fusion is the practice of combining supply-chain signals that differ in latency, reliability and coverage — demand, supply and capacity, in-transit visibility, inventory position, finance, and risk — into a single estimate carrying a stated confidence, so one commitment can be made against all six rather than six commitments made against one each.

Illustration of a warehouse floor where separate demand, capacity, transit and inventory views converge into one operating picture on a shared display
Logistics · Future of AI & Visionary Thinking

Key takeaways

  1. Supply fusion is combining signals of different latency, reliability and coverage into one estimate with a stated confidence. Putting six dashboards on one screen is not fusion — it renders the disagreement at higher resolution and leaves the resolution to whoever is looking.
  2. A fused estimate only beats the best single source when the sources' errors are not copies of each other. Two feeds derived from the same upstream reduce nothing and inflate confidence, which is why a correlation register belongs in the design before any weighting does.
  3. Most apparent disagreement between supply-chain sources is frame error — different event definitions, identifiers, time zones or units — not signal error. Resolve the frame first; weighting incompatible claims is arithmetic on nonsense.
  4. The moment a fused number drives a commitment it needs provenance: the contribution vector — which source, what value, what weight, how stale — stored alongside the result and retained for as long as the commitment can be disputed.
  5. The characteristic failure of fusion is one confident wrong number replacing three honest uncertain ones. The cheapest defence is to make confidence a first-class field the system of record refuses to write without.

Abbreviations used on this page

TMS
Transport management system
WMS
Warehouse management system
YMS
Yard management system
ERP
Enterprise resource planning — the finance and order ledger
S&OP
Sales and operations planning
EDI
Electronic data interchange — e.g. the 214 shipment status and 315 ocean status messages
ETA
Estimated time of arrival
AIS
Automatic identification system — broadcast vessel position and heading
EPCIS
The GS1 event-sharing standard: the what, when, where, why and how of a supply-chain event
SSCC
Serial shipping container code — the GS1 key that identifies a logistics unit
POS
Point of sale — retail sell-through, the fastest demand signal
OTIF
On-time in-full delivery rate

Free · 8 questions · ~3 minutes

Score your operation on the fusion ladder

Eight questions, one at a time, about three minutes. Answer them and we build your personalised fusion report — your rung on the ladder, your score on each of the four dimensions, and the specific thing standing between your signals and a single defensible estimate — and send it to your inbox.

0 of 8 answered

Question 1 of 8Domain coverage

How many of the six domains — demand, supply and capacity, in-transit, inventory position, finance, risk — feed the estimate behind your biggest recurring commitment?

Coverage sets the ceiling. An estimate blind to a domain will be confidently wrong exactly when that domain moves.

How the score maps to a stage
  • 04 — Stage 1, Siloed. Siloed is when each decision domain keeps its own number and nobody is required to reconcile them, so disagreement is invisible rather than absent.
  • 59 — Stage 2, Aligned. Aligned is when the domain numbers are put side by side on a calendar and one of them is chosen, so reconciliation happens — but as an event, by seniority, and without a record of the reasoning.
  • 1015 — Stage 3, Integrated. Integrated is when all six signals land in one place against shared identifiers and a deterministic rule picks the winner — but the output is still a chosen source, not a combined estimate.
  • 1621 — Stage 4, Fused. Fused is when the estimate is a weighted combination of sources with a stated confidence, and the commitment it drives carries that confidence and its provenance with it.
  • 2224 — Stage 5, Self-reconciling. Self-reconciling is when the fusion measures its own sources: weights are re-estimated from realised outcomes, weak feeds are demoted automatically, and cross-party feeds carry quality terms in the contract.

What supply fusion is — and what it borrows from sensor fusion

A definition, the five disciplines that make fusion rigorous rather than decorative, and the path a signal takes from a feed to a commitment.

Supply fusion is the practice of combining supply-chain signals that differ in latency, reliability and coverage into a single estimate with a stated confidence. The six domains being merged are demand signal, supply and capacity, in-transit visibility, inventory position, finance and working capital, and risk. The product of fusion is not a screen; it is one number, an interval around it, and a record of how both were produced.

The discipline worth borrowing is sensor fusion — the problem of turning radar, lidar and camera readings into one estimate of where an object is. It is worth borrowing because it is rigorous about exactly the thing supply chains are vague about: what to do when two instruments disagree. A sensor-fusion engineer never averages raw readings. They put every reading into a common frame of space and time, attach a measured uncertainty to each instrument, combine so that the result is sharper than any input, gate readings that fall too far outside what the model expected, and re-estimate the weights from the residuals — the gap between what each instrument claimed and what turned out to be true.

Each of those five transfers directly. The common frame is identifiers and event definitions — the work that standards bodies have already done in GS1's identification keys (opens in a new tab) and the EPCIS event standard (opens in a new tab). The uncertainty model is a measured error distribution per source, per horizon, not a reputation. The combination is weighted, so the fused interval is narrower than any single source's. The gate is what stops a large disagreement being quietly averaged into invisibility. And the residual loop is what keeps the weights honest as the network changes. Everything on this page is one of those five, applied to freight.

How six supply signals become one estimate with a confidence

The path from feed to commitment, in three regimes. The top lane is where most operators sit: six views, one human picking, and a commitment that carries no record of the uncertainty it was made under. The middle lane is fusion proper. The bottom lane closes the loop by scoring the sources against what actually happened.

  • Data & feeds
  • Where value leaks
  • System-of-record action
  • AI / model
  • Human in the loop

The process, in words

  • In the top lane the six domain views arrive independently, a planner picks whichever they trust today, and the commitment is made against a single value with no record of the uncertainty or the discarded claims. When the commitment is questioned months later, the evidence that would explain it no longer exists.
  • In the middle lane every source first enters a common frame — resolved identifiers, defined events, normalised time and units — then meets a store of its own measured error by horizon, lane and counterparty. The estimator produces a value and an interval, with feeds sharing an upstream weighted as one cluster rather than as independent witnesses.
  • The gate is the part that distinguishes fusion from averaging. When two sources disagree by more than their own uncertainty allows, the estimate is not quietly split down the middle: the disagreement escalates to a named owner with both numbers and their history attached, because a large residual usually means the world moved, not that a sensor is noisy.
  • In the bottom lane the realised outcome — gate-in, receipt, sell-through, invoice — is scored against what each source claimed, and the residuals re-fit the weights. This is what makes the arrangement self-reconciling, and it is only safe once you can prove the outcome was not influenced by the estimate itself.
Step-by-step insights
The common frame — where most 'model' problems actually live
Before any weighting is legitimate, two sources must be proved to describe the same event. In freight that is rarely trivial: a carrier's 'arrival' may be berthing, the terminal's may be first-lift, and the receiving team's may be gate-in at the DC — three different instants, hours or days apart, all called ETA. Add time zones stored inconsistently, units that differ between pallets and cases, and identifiers that only resolve through a reference field, and a large share of what teams experience as disagreement between sources turns out to be frame error. This is why the standards work matters operationally rather than theoretically: shared identification keys and a shared event grammar remove a class of disagreement instead of arbitrating it.
The source error store — a reputation replaced by a distribution
Every source needs a measured error distribution, and it must be conditional. The same carrier's ETA is a different instrument at fourteen days than at thirty-six hours; a terminal event is superb inside two days and silent beyond it; a lane's historical dwell is weak in normal weeks and the only thing left standing during a disruption. Storing error as a single number per source throws away the structure that makes fusion work. The practical minimum is error by source × horizon bucket × lane or trade, refreshed as outcomes land, with the sample size stored alongside so a thin cell can be treated as thin rather than as confident.
Why the fused interval is narrower — and when that is a lie
Combining independent estimates reduces variance: that is the whole mathematical case for fusion, and it is why a weighted blend beats even the best single source. The condition is independence. If two of your feeds are both derived from the same upstream — a visibility provider reselling the same carrier messages you already receive, or two portals rendering one terminal system — then they are one witness speaking twice. Treating them as two shrinks the computed interval without shrinking the real error, which is the precise recipe for a confident wrong number. A correlation register that caps the combined weight of each cluster is not an optimisation; it is the thing that keeps the confidence honest.
The gate — a large disagreement is information, not noise
When a source's claim falls outside what the fused estimate expected, given both uncertainties, the disagreement is itself the most valuable signal in the pipeline. Averaging it away destroys that. Gating means the estimate refuses the outlier, records it, and — if gate firings cluster — escalates, because a rising gate rate almost always means the world has moved outside the model's validity: a new lane, a diverted vessel, a strike, a carrier system migration. Operators who monitor gate rate get a leading indicator of network change for free; operators who average get a lagging one, expressed as complaints.
The contribution vector — what makes the commitment defensible
A fused value destroys its inputs unless you deliberately keep them. The contribution vector is the small record stored with every fused number: which sources contributed, what each claimed, what weight each carried, how stale each was at fusion time, and whether the gate fired. It costs a few hundred bytes per estimate and it is the difference between explaining a commitment months later and asserting it. Retention should be set by the commercial dispute window — demurrage, detention, customs and customer service-level claims all run far longer than a typical data-retention default — not by storage convenience.
Closing the loop safely — the contamination trap
Re-fitting weights from realised outcomes is only valid when the outcome is independent of the estimate. Supply chains break that assumption constantly: publish your fused ETA to a partner portal and it may come back tomorrow as that partner's own estimate, which your scorer then reads as independent corroboration and rewards with weight. The loop then converges on your own opinion with rising confidence and falling accuracy. Tag the lineage of every inbound value, exclude any source whose value could have been influenced by yours, and audit for round-tripping before automatic re-weighting is switched on.
  • Fused estimates beat the best single source because errors partially cancel

    Combining sources whose errors are not perfectly correlated produces an estimate with lower variance than any input. This is the entire mathematical case, and it holds only under that condition — which is why identifying correlated feeds matters more than choosing an estimator.

  • …because each source is blind somewhere different

    Carrier messaging stops at the gate, terminal events start at the quay, the warehouse system begins at the door, and finance only sees the movement weeks later on an invoice. No single feed covers the journey. The union does, and the seams are exactly where commitments fail.

  • …and because a fused estimate degrades gracefully when one input dies

    A feed that stops updating takes a single-source estimate down silently. In a fused estimate with staleness weighting, the same failure widens the interval and shifts weight to the surviving sources — a visible, proportionate degradation rather than an invisible one.

  • But only if the confidence travels with the number

    A fused point estimate with the interval stripped off is worse than the three honest sources it replaced, because it has destroyed the information that they disagreed. Confidence is not a nice-to-have on a fused number; it is the part that makes fusion defensible.

The five stages in detail: Siloed to Self-reconciling

For each rung: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.

The ladder runs Siloed → Aligned → Integrated → Fused → Self-reconciling, and each rung is defined by how disagreement between sources is resolved. At Siloed it is not resolved because it is not visible; at Aligned it is resolved in a meeting by authority; at Integrated by a fixed rule; at Fused by measured weights with a stated confidence; and at Self-reconciling by weights the system re-estimates from what actually happened.

The rungs are cumulative in a specific way: each one is mostly the previous one plus a new artefact. Aligned adds a cadence, Integrated adds a frame and a rule, Fused adds an error store and an interval, Self-reconciling adds a residual loop and a contract. Skipping an artefact does not accelerate the climb — it produces a system that computes a confidence from beliefs nobody measured, which is the most expensive failure on this page.

Decision quality released as fusion matures

The value is close to flat across the first two rungs, because a reconciled number produced weekly by a meeting is not usable by any downstream system. It inflects when the frame exists and the estimate starts carrying an interval — the point at which other systems can begin to reason about the number rather than merely display it.

Decision quality released by stage

  • Stage 1 · Siloed — 22% of operators. Siloed is when each decision domain keeps its own number and nobody is required to reconcile them, so disagreement is invisible rather than absent.
  • Stage 2 · Aligned — 34% of operators. Aligned is when the domain numbers are put side by side on a calendar and one of them is chosen, so reconciliation happens — but as an event, by seniority, and without a record of the reasoning.
  • Stage 3 · Integrated — 27% of operators. Integrated is when all six signals land in one place against shared identifiers and a deterministic rule picks the winner — but the output is still a chosen source, not a combined estimate.
  • Stage 4 · Fused — 14% of operators. Fused is when the estimate is a weighted combination of sources with a stated confidence, and the commitment it drives carries that confidence and its provenance with it.
  • Stage 5 · Self-reconciling — 3% of operators. Self-reconciling is when the fusion measures its own sources: weights are re-estimated from realised outcomes, weak feeds are demoted automatically, and cross-party feeds carry quality terms in the contract.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with GS1's published position on shared supply-chain data standards.

Select a rung

Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Siloed

22% of operators sit here

Siloed is when each decision domain keeps its own number and nobody is required to reconcile them, so disagreement is invisible rather than absent.

Stage 1 is not an absence of data. Most operators at this stage have more supply-chain data than they have ever had: carrier messages, terminal events, telematics, warehouse scans, sell-through, invoices, weather. What is missing is any obligation for those sources to agree, and any mechanism that would notice if they did not. Each domain is internally consistent and externally unreconciled.

The tell is the arrival date. Ask three people in the same building when a specific container lands and you will get three answers — the carrier's published ETA, the terminal's berth plan, and the receiving team's working assumption based on how that lane behaved last month. Every one of them is defensible. None of them is the number, because there is no such thing as the number yet: there are only claims, and the claims have never been asked to stand next to each other.

This is a cheap stage to leave and an expensive stage to remain in, and the expense is invisible on any budget line. It shows up as buffer: extra days of cover, extra labour called in because nobody trusted the arrival window, extra premium freight because the working-capital view and the transport view were reconciled after the decision rather than before it. The buffer is the price of unreconciled signals, and it is paid every week.

In practice

The three arrival dates

A grocery importer's inbound container has an ETA of Tuesday in the carrier portal, a berth window starting Wednesday in the terminal's schedule, and a Thursday assumption in the DC's labour plan, because that terminal has been running two days late all quarter. Each system is correct about what it knows. The receiving manager books labour for Thursday, the planner promises the customer Tuesday, and the finance team accrues on the carrier date. When the box lands on Wednesday, all three are wrong in different directions and nobody can explain why the commitments disagreed.

What it looks like

  • Demand, capacity, in-transit, inventory, finance and risk each have their own screen and their own owner
  • No shared identifier links a purchase order, a container, a receipt and an invoice
  • The same shipment carries three different arrival dates in three systems, and none of them is wrong
  • An hour after a number is quoted, nobody can say which system it came from

Diagnostic signals you can check this week

  • Ask three teams for the arrival date of one specific container and write down all three answers
  • Ask which identifier links the carrier booking to the warehouse receipt — if the answer is 'the reference field, usually', you are here
  • Count the screens a planner opens before making one commitment; above four, no fusion is happening
  • Ask what happens when the carrier ETA and the terminal event disagree. If the answer is a person's name, the rule is not written down

Anti-pattern · Buying a control tower to fix a definition problem

The instinctive fix is a visibility platform that shows all six domains on one screen. It will, and the disagreement will still be there — now rendered in higher resolution and refreshed more often. A screen cannot resolve a conflict between two claims that were never defined in comparable terms, and every hour spent admiring the new view is an hour not spent writing down what 'arrival' means. Write the event dictionary for one decision first. The platform's real requirements are visible after that, not before it.

What holds you here

There is no common identifier and no shared event definition, so the domain numbers cannot be compared — let alone combined.

Highest-leverage next move

Pick one recurring commitment and write the event dictionary behind it: what each source claims, in what units, keyed to what identifier, refreshed how often.

Cost of leaving

Effort
2–4 months
Team
One data engineer and one planner, part-time
Risk
Low — the work is documentation and instrumentation; nothing in production depends on it yet
To next stage
2–4 months

If this is you, the next step is

A two-week exercise: list the sources, define the events, find the identifier gaps.

Map the signals behind one decision

Stage 2

Aligned

34% of operators sit here

Aligned is when the domain numbers are put side by side on a calendar and one of them is chosen, so reconciliation happens — but as an event, by seniority, and without a record of the reasoning.

Stage 2 is where most logistics operators are, and it is a real achievement over stage 1: the claims are now made to stand next to each other on a regular cadence, and someone is accountable for producing a single answer at the end of it. The weekly S&OP call, the daily exception review and the pre-peak alignment session are all this stage in different clothing.

The structural weakness is that reconciliation is an event rather than a function. Between meetings the domains diverge again, quietly and at their own rates: the demand view moves with every order, the transit view moves with every milestone message, the finance view moves once a month. By Thursday the Tuesday consensus is describing a network that no longer exists, and the next meeting rebuilds it from scratch rather than updating it.

The second weakness is that the meeting resolves disagreement by authority. That is not a criticism of the people — in the absence of measured source quality, seniority is a reasonable proxy for judgement. But it means the resolution is unreviewable, unversioned and untransferable. When the person who always called it correctly changes role, the operation loses a capability that was never written down, and nobody can say by how much the estimates got worse, because nothing was ever scored.

In practice

The Tuesday call

A regional 3PL runs a Tuesday alignment call: planning brings the forecast, transport brings the carrier ETAs, warehousing brings the receiving capacity and finance brings the cost view. It takes ninety minutes and it works — a single plan leaves the room. It is also the only ninety minutes in the week when the six views agree. By Thursday, two carriers have revised, an inbound has been rolled, and the plan in the system reflects a Tuesday world. The following week the same ninety minutes is spent rebuilding rather than refining.

What it looks like

  • A recurring call compares the demand, supply, inventory and transport views
  • Identifiers mostly resolve, via a mapping table one person maintains
  • The chosen number is recorded; the numbers not chosen are discarded
  • Confidence is expressed verbally — 'we think Thursday' — and never stored anywhere

Diagnostic signals you can check this week

  • Ask how long the reconciled plan stays valid; if the honest answer is under 48 hours, the cadence is the constraint
  • Ask whether the losing numbers are stored anywhere after the call. Usually they are not
  • Ask who resolves a demand-versus-capacity conflict, and whether the rule they use is written down
  • Check whether the mapping table between systems lives in one person's spreadsheet

Anti-pattern · Adding a seventh source to settle a six-source argument

When the meeting cannot resolve a disagreement, the reflex is to buy another feed — a second visibility provider, a market-rate index, a weather service — on the theory that more evidence produces more certainty. It usually produces more argument, and often less certainty, because the new source is frequently derived from an upstream you already consume. You have not added an independent opinion; you have added a louder echo. Before buying anything, score the sources you already have against what actually happened.

What holds you here

Reconciliation is an event rather than a function, so the domains diverge between meetings and the reasoning behind each chosen number is lost.

Highest-leverage next move

Move the tie-breaks out of the meeting and into written rules: for each pair of sources, which one wins under which condition, and what evidence would change that.

Cost of leaving

Effort
3–6 months
Team
A data engineer, an analyst, and the person who currently runs the reconciliation call
Risk
Low to medium — the change is procedural, and the meeting stays in place while the rules are written
To next stage
3–6 months

If this is you, the next step is

We sit in the call, capture the tie-breaks people already make, and write them down as testable rules.

Turn your reconciliation call into rules

Stage 3

Integrated

27% of operators sit here

Integrated is when all six signals land in one place against shared identifiers and a deterministic rule picks the winner — but the output is still a chosen source, not a combined estimate.

Stage 3 is the first stage where the machine, rather than a meeting, produces the answer. The identifiers resolve, the events are defined, and a precedence ladder decides what wins. Operators arriving here report an immediate and durable relief: the arguments about whose number is right simply stop, because the rule is visible and the same rule ran last night as runs tonight.

What precedence buys is consistency. What it costs is information. When the terminal event overrides the carrier ETA, the carrier's claim is discarded rather than used — and the carrier's claim carried signal, particularly about the future, which the terminal event does not. Precedence is a lossy summary of the evidence, and its loss is largest exactly when the sources disagree most, which is exactly when the decision matters most.

The second cost is that precedence never learns. It is a set of static beliefs about which source is better, usually written by whoever understood the systems best in the year the ladder was designed. Networks change; a carrier's data quality improves after an integration project, a terminal's feed degrades after a system migration, a lane changes mode. Nothing in a precedence ladder notices. The move to stage 4 begins when someone starts measuring how often each source was actually right.

In practice

The precedence ladder that never learned

A retailer's inbound estimate applies a fixed ladder: terminal event first, then carrier ETA, then the lane's historical average. It ran unchanged for three years. In year two one carrier rebuilt its messaging and its ETAs became the most accurate signal in the estate at horizons over five days — better than the terminal, which only exists inside 48 hours. The ladder had no way to express that, so the best available evidence at the horizon where decisions were actually made was systematically discarded in favour of a source that had nothing to say yet.

What it looks like

  • One data layer holds all six domains against common keys and a shared event dictionary
  • Precedence rules resolve conflicts deterministically — terminal event beats carrier ETA, carrier ETA beats plan
  • Every displayed number is traceable to the source it came from
  • There are no weights and no confidence: the answer is a pick, not a blend

Diagnostic signals you can check this week

  • Ask to see the precedence ladder, then ask when it was last changed and on what evidence
  • Ask whether the losing sources' values are retained alongside the winning one — if not, no fusion is possible later
  • Ask for the measured error of each source, split by horizon. If it does not exist, the ladder is a belief
  • Check whether the estimate is a single value with no interval anywhere in the pipeline

Anti-pattern · Treating precedence as fusion

Because the pipeline is now automated, deterministic and traceable, it is easy to describe it internally as fused. It is not: it is a well-governed choice between claims. The distinction matters commercially, because a picked value cannot carry a defensible confidence, and a commitment made against it inherits an uncertainty nobody has quantified. The honest position at stage 3 is that you have a single number with unknown error. Say that out loud before someone downstream builds an automated decision on it.

What holds you here

Precedence discards information: the losing sources still carried signal, and nobody measures how often the winner was actually right.

Highest-leverage next move

Score every source against realised outcomes by horizon, lane and counterparty. That error table is the raw material of every weight you will ever set.

Cost of leaving

Effort
6–12 months
Team
A data engineer, an ML or estimation engineer, and a named operations owner for the decision
Risk
Medium — the error store and the shadow estimate can be built without touching any live commitment
To next stage
6–12 months

If this is you, the next step is

Twelve weeks of realised outcomes scored against every source, by horizon and lane.

Build your source-error store

Stage 4

Fused

14% of operators sit here

Fused is when the estimate is a weighted combination of sources with a stated confidence, and the commitment it drives carries that confidence and its provenance with it.

Stage 4 is where the arithmetic starts paying. A weighted estimate with a measured error distribution is not merely a nicer number: it is a number a downstream system can reason about. A labour planner can book to the 80th percentile of an arrival window; a customer-facing promise can be made at the confidence the contract requires; an automated replenishment can decline to act when the interval is too wide. None of that is possible with a point estimate, however accurate it happens to be on average.

The engineering is unglamorous and mostly bookkeeping. The estimator itself — an inverse-variance weighting, a small state-space model, a gradient-boosted residual correction — is rarely the hard part and is usually a week's work once the error store exists. What takes the quarter is the frame: proving that two sources are describing the same event, normalising time and units, resolving identifiers, and discovering which of your six feeds are secretly two.

The stage's characteristic risk is that the interval gets dropped at the last metre. The estimate is computed with a confidence, and then written into a system-of-record field that takes exactly one value, and the confidence dies at the boundary. Everything downstream then behaves as if the number were certain — which makes the fused estimate strictly worse than the three honest sources it replaced, because the disagreement information has been destroyed rather than summarised. Refusing to write a value without its interval is the single cheapest control on this page.

In practice

The 72-hour labour booking

A distribution centre books agency labour 72 hours ahead. The fused arrival estimate combines the carrier ETA, the terminal berth plan, an AIS-derived vessel position and the lane's realised dwell history, weighted by each source's measured error at the 72-hour horizon. The output is a window and a confidence. The booking rule is explicit: book to the 80th percentile of the window, and when the window is wider than eight hours, book the flexible shift instead. Reversals still happen — but they happen inside the stated interval, which is the definition of the estimate working.

What it looks like

  • Weights derive from measured per-source error at the relevant horizon, not from contract tier or seniority
  • Correlated sources are registered and weighted as a single cluster
  • Every fused value carries an interval and a contribution vector
  • Disagreements beyond tolerance fire a gate and escalate to a named owner instead of being averaged away

Diagnostic signals you can check this week

  • Open the system of record and check whether the confidence is a field next to the value, or a slide in a deck
  • Ask how the weights were set and when they were last re-fitted
  • Ask for the correlation register — the list of which feeds share an upstream
  • Ask what happens when two sources disagree beyond tolerance: silent average, or gate and escalation?

Anti-pattern · Shipping the point estimate and dropping the interval

The interval is the first casualty of integration, because the destination field takes one value and adding a second column is a change request. Teams promise to add it later and never do, and within a quarter the organisation has forgotten the number was ever uncertain. The failure surfaces as a category of incident nobody can explain: commitments that were reasonable given the evidence, judged afterwards as errors, because the evidence of uncertainty was not retained. Add the interval column in the same change as the value, or accept stage 3 honestly.

What holds you here

Weights are fitted once and reviewed rarely, so the fusion stays correct for the network as it was on the day it was built.

Highest-leverage next move

Close the loop: score every source against realised outcomes continuously, re-fit the weights on a cadence, and review the change like code.

Cost of leaving

Effort
9–18 months
Team
An estimation engineer, an integration engineer, an operations owner, and commercial support for partner feeds
Risk
Medium to high — the first commitment made automatically on a fused number needs a drilled rollback to the previous source
To next stage
12–24 months

If this is you, the next step is

We audit the weights against your realised outcomes and stress-test the gate on a real disagreement.

Review your weighting and gating

Stage 5

Self-reconciling

3% of operators sit here

Self-reconciling is when the fusion measures its own sources: weights are re-estimated from realised outcomes, weak feeds are demoted automatically, and cross-party feeds carry quality terms in the contract.

Stage 5 is narrower than the name suggests. It does not mean the supply chain reconciles itself; it means the estimation layer maintains its own beliefs about its sources without a human re-fitting them. A carrier whose messaging degrades after a system migration is down-weighted within days rather than at the next annual review. A new terminal feed earns weight by being right, not by being procured.

The engineering that makes this safe is the ablation harness: periodically recompute the estimate with each source removed and measure what the estimate loses. It answers two questions no dashboard can. First, which feeds are actually contributing — operators routinely discover they are paying for a source whose removal changes nothing, because it is a repackaging of something they already have. Second, which feed the whole estimate is quietly leaning on, which is the one to hold a contract conversation about before it fails.

The binding constraint at this stage stops being technical. Re-weighting a partner's feed downwards is a commercial act: it changes what you pay them for, what you promise your own customers, and what you will say in a dispute. Operators who reach here find that the work is contract design, published methodology and review cadence — and that the discipline most worth importing is the one used by institutions that publish uncertain estimates for a living: state the method, state the revision policy, and let the counterparty see their own score.

In practice

The demoted feed

An operator's fusion layer scores every in-transit source weekly against realised gate-in times. After a carrier's platform migration, that carrier's ETA error at the five-day horizon roughly doubles, and the weighting shifts automatically towards the terminal feed and lane history. Nothing breaks and no incident is raised — the estimate simply degrades gracefully. What the scorecard actually triggers is a commercial conversation, held with a measured series rather than an anecdote, and the carrier's own engineering team uses the same series to find the regression.

What it looks like

  • Source scorecards update continuously from realised outcomes and drive the weights directly
  • Ablation runs on a schedule, so each source's marginal contribution is a known number
  • Calibration is monitored: of commitments made at a stated confidence, the hit rate is tracked
  • Partner feeds carry contractual quality terms and a scorecard both parties can see

Diagnostic signals you can check this week

  • Ask when the weights last changed and whether a human changed them
  • Ask for the last ablation report and what it said about the least useful feed
  • Ask for the calibration curve: of commitments made at 80% stated confidence, how many held?
  • Ask whether any partner can see their own score, and whether the contract says what happens when it falls

Anti-pattern · Closing the loop on contaminated data

Automatic re-weighting is only safe if the realised outcomes are independent of the estimate. They frequently are not. If your fused ETA is published to a portal, a partner may ingest it and post it back as their own estimate — which your scorer then reads as an independent source agreeing with you, and rewards with more weight. The loop converges on your own opinion with rising confidence. Tag the lineage of every inbound value, refuse to score a source against an outcome it influenced, and audit for round-tripping before switching automatic re-weighting on.

What holds you here

The remaining constraint is commercial and legal rather than technical: down-weighting a partner's feed is a contract conversation before it is a code change.

Highest-leverage next move

Put the score into the contract — agreed definitions, a published methodology, a revision policy and a review cadence both parties can see.

Cost of leaving

Effort
Continuous
Team
A standing estimation team, an operations owner, and commercial and legal partners for the feed contracts
Risk
Concentrated — low frequency, high consequence, and increasingly contractual rather than technical

If this is you, the next step is

We stress-test the scorer, the ablation harness and the contamination controls against a real feed.

Audit a self-reconciling estimate

Where logistics operators actually sit on the fusion ladder

The distribution across the five rungs, and why the Integrated → Fused step loses the most operators.

Most logistics operators are at Aligned — reconciling their domain views on a cadence, in a meeting, with the outcome recorded and the reasoning discarded. A substantial minority have reached Integrated, where a common frame and a precedence rule produce the answer automatically. The number producing a genuinely weighted estimate with a stated confidence is small, and the number whose weights re-fit themselves from realised outcomes is very small indeed.

Distribution of logistics operators across the five rungs

Illustrative distribution. Aligned is the mode; the largest single drop is Integrated → Fused, because that step requires an artefact nobody has by accident — a measured error distribution per source, per horizon.

Share of operators

  • 22% — 1 · Siloed
  • 34% — 2 · Aligned (the mode)
  • 27% — 3 · Integrated
  • 14% — 4 · Fused
  • 3% — 5 · Self-reconciling

Source: Illustrative distribution, synthesised from published work by MIT CTL, Gartner and GS1 on supply-chain data sharing

The Integrated → Fused step is where the distribution thins, and the reason is that it is the first step requiring an artefact no organisation acquires by accident. A common frame can be built as a side effect of an integration programme; a precedence ladder can be written in an afternoon. A measured error distribution per source, per horizon, per lane can only come from deliberately recording what each source claimed and then scoring it against what happened — which nobody does unless someone decides to. Research centres working on supply-chain analytics, including MIT's Center for Transportation & Logistics (opens in a new tab) and the World Economic Forum's supply-chain centre (opens in a new tab), have made the same point from different directions: the constraint on multi-party supply-chain decision-making has moved from data availability to data comparability.

It is also worth being precise about what the top of the ladder is not. Self-reconciling does not mean a supply chain that runs itself; it means an estimation layer that maintains its own beliefs about its sources. That is a much narrower claim than most futures writing about supply chains makes, and it is the claim this page is prepared to defend. Gartner's supply-chain research (opens in a new tab) tracks the broader adoption picture; what it consistently shows is a gap between organisations that have connected their data and organisations that have made a decision differently as a result.

The signal ledger: six domains, six failure modes, six weight rules

What each fused domain actually claims, how fast it moves, how it fails quietly, where it is blind — and the rule that should set its weight.

Every source in a supply-chain estimate is a different instrument, and the ledger below is the specification sheet for all six. Read it as a sensor datasheet rather than as a data catalogue: the columns that matter are not what a feed contains but how fast it moves, how it degrades, what it cannot see, and therefore how much weight it has earned. A fusion design that skips this table will weight sources by contract value or by which vendor presented most recently, which is how confident wrong numbers get built.

DomainWhat it actually claimsRefresh and lagHow it fails quietlyWhere it is blindWeight rule
Demand signalWhat customers will order, and whenPOS daily; orders continuous; forecast weekly or per cycleKeeps its shape while its level drifts — promotions, channel shifts and price moves all arrive looking like demandDemand it never saw: substitutions, lost sales, unlisted channels, and the customer's own inventory positionInside the replenishment lead time, firm orders outweigh any forecast; beyond it, the forecast is the only instrument in the room
Supply and capacityWhat a supplier, carrier or terminal says it can doConfirmations per booking cycle; spot capacity hourly; contract capacity per seasonConfirmations are commitments, not observations — they stay green while the counterparty quietly falls behind, and correct all at onceThe counterparty's own upstream: their raw-material position, their labour, their subcontractorsWeight by the counterparty's realised fulfilment rate, never by contract tier or relationship seniority
In-transit visibilityWhere the goods are now and when they will arriveEDI 214/315 at milestone events; AIS in minutes; terminal events per moveA stale ETA is re-served as fresh, and the same upstream is resold by two providers as two opinionsThe gaps between milestones — the long silences where nothing is emitted and interpolation is doing the workWeight by measured error at the specific horizon, and cluster every provider sharing an upstream into one witness
Inventory positionWhat is where, and what can be promisedWMS in near real time; ERP nightly; in-transit stock inferred rather than observedCounts drift between cycle counts, and 'available' diverges from 'on hand' as reservations, holds and damage accumulateQuality holds, damaged stock, and inventory sitting in a partner's building under someone else's systemPrefer the system that owns the physical movement over the system that owns the ledger; treat in-transit stock as an estimate, not a count
Finance and working capitalWhat the decision costs, and when cash actually movesInvoices per settlement cycle; rate cards per contract; FX dailyPrices lag reality — the rate card is the last artefact updated after a market move, so cost views stay stale for weeksAccessorials, demurrage and detention, which are invisible until they are invoiced long after the decisionFor anything spot, weight recently realised invoices over rate cards; for contract lanes, the reverse
RiskWhat could invalidate the plan, and roughly whereWeather hourly; port congestion daily; labour, geopolitical and carrier-health signals episodicHigh recall and low precision — acted on literally it generates constant churn, so teams learn to ignore it entirelyThe specific: it can tell you the region, the port or the corridor, but almost never your boxNever a value. Risk widens the interval and lowers the gate threshold; it does not move the point estimate
The six fused domains as instruments. 'Fails quietly' is the degradation that produces no error message — the dangerous class, because the feed keeps returning values. 'Weight rule' is the default; measured error at the decision horizon always overrides it.

The last row is the one most often got wrong, and it is worth stating as a rule: risk signals are interval-wideners, not estimate-movers. A storm forecast does not tell you that this container will be four days late; it tells you the distribution of possible arrivals has a longer tail this week. Feed it in as a shift in the point estimate and you generate churn — every plan moves, most of the moves are wrong, and within two cycles the operation stops believing risk signals at all. Feed it in as a widening of the interval and a lowering of the gate threshold, and the effect is exactly what you want: commitments get more conservative and more disagreements get escalated, precisely during the period when the network is least predictable. Public forecasting institutions have handled this correctly for decades — ECMWF's operational forecasts are ensemble-based (opens in a new tab), describing a range of scenarios and their likelihood rather than a single answer, and NOAA (opens in a new tab) publishes probabilistic products for the same reason.

  • The frame: identifiers, events, time and units

    One key that resolves an order, a logistics unit, a location and an invoice to the same physical thing; one written definition per event; one time base; one unit per measure. GS1's identification keys (opens in a new tab), the EPCIS event standard (opens in a new tab) and master-data syndication through GDSN (opens in a new tab) exist precisely so this does not have to be invented per trading relationship. Build it before the estimator, not after.

  • The error store: what each source claimed, and what happened

    A table of source × horizon × lane holding the measured error distribution and the sample size behind it. It is built by writing down every claim at the time it is made and joining it to the realised outcome later — cheap to start, impossible to backfill, which is why it should start the week you decide fusion is the direction.

  • The correlation register: which feeds are secretly one feed

    A documented list of which sources derive from a shared upstream, with a cap on the combined weight of each cluster. Most estates contain at least one cluster nobody has noticed — typically a visibility provider reselling carrier messages the operator already receives directly, counted twice as agreement.

Those three artefacts are the whole prerequisite list. Notice what is not on it: a new platform, a data lake migration, or a model. The estimator is genuinely the easy part — an inverse-variance weighting over four sources is a short function, and the sophisticated alternatives buy less than practitioners expect once the frame and the error store exist. What buys the accuracy is knowing which instrument to believe, at which horizon, on which lane. That is bookkeeping, and it is the bookkeeping this page is about.

Resolving disagreement: the order that makes fusion defensible

Frame, then correlation, then weight, then gate. Getting the order wrong is what produces a confident wrong number.

Disagreement is resolved by rule, in a fixed order: reconcile the frame, cluster the correlated sources, apply measured weights, and gate what remains. The order is not stylistic. Weighting before the frame is reconciled produces confident nonsense, because the arithmetic is being applied to claims about different events. Weighting before correlation is registered inflates confidence, because echoes are counted as witnesses. Gating last is what catches the residual — the disagreement that survives all three steps and therefore means something.

The disagreement-resolution order

  1. 1 · Prove the sources are describing the same event

    Before any comparison, confirm that both claims refer to the same instant, the same object and the same unit. Berthing is not first-lift; first-lift is not gate-out; gate-out is not receipt. A large share of what teams experience as source disagreement dissolves here, and the resolution is permanent — a definition fixed once stops generating disagreements forever, whereas a weight tuned around a definitional mismatch has to be re-tuned every time the mismatch moves.

  2. 2 · Cluster sources that share an upstream

    Trace each feed to its origin and group the ones that are re-serving the same underlying data. Cap the weight of the cluster, not the members. Two providers rendering one terminal's events are one witness with two microphones, and the most common way an estate ends up with a fused interval far narrower than its actual error.

  3. 3 · Weight by measured error at the decision horizon

    Use the error store, conditioned on horizon and lane. The horizon condition matters more than practitioners expect: sources rank differently at fourteen days than at thirty-six hours, and a fixed ranking is wrong at one end of that range by construction. Where a cell is thin, widen the interval rather than pretending to a precision the sample does not support.

  4. 4 · Gate what still disagrees, and never silently average it

    If a source's claim lies outside what the fused estimate expected given both uncertainties, exclude it, record it, and route it. A gate firing is an event worth a person's attention; a cluster of gate firings on one lane or one counterparty is a network change worth acting on before it becomes an incident.

  5. 5 · Escalate with both numbers, not with an answer

    When a disagreement reaches a human, it should arrive as two claims, their histories and their measured reliabilities — not as a pre-averaged value with an amber icon. The person is being asked to exercise judgement about which instrument to believe in an unusual situation, and stripping the evidence to make the alert tidy removes the only thing that makes their judgement better than a coin toss.

  6. 6 · Write down who owns the tie-break before you need one

    Every fusion needs a named owner per decision class, with the authority to override and the obligation to record why. The record is the point: an override without a reason is indistinguishable from noise six months later, and the accumulated reasons are the highest-quality training data any fusion layer will ever get.

What to do when sources disagree

Plot how much the sources agree against how expensive it is to be wrong. Only one quadrant genuinely needs a human, and the top-right is the one that quietly catches people out — agreement is only reassuring when the agreeing sources are independent.

Commit and move on

  • Sources agree; the commitment is cheap to reverse
  • Automate; sample-audit rather than review
  • Do not spend governance budget here

Commit, but check independence

  • Agreement is only evidence if the sources are independent
  • Check the correlation register before trusting the narrow interval
  • Store the contribution vector — this is the quadrant that gets disputed

Pick a rule and log it

  • Sources disagree, but being wrong is cheap
  • Deterministic tie-break, no human in the path
  • Log the disagreement — it is free error-store data

Escalate with both numbers

  • Disagreement plus expensive commitment: the only quadrant needing a person
  • Send both claims, their histories and their reliabilities
  • Record the decision and the reason; it sets tomorrow's rule
Source agreement — top: Sources agree, bottom: Sources disagree
Cost of being wrong — left: Cheap to reverse, right: Expensive or irreversible

Different data owners are involved per port call: it is always a combination of both the port and the terminal, yet also nautical service providers and services related to vessels and cargo.

That sentence is the whole cross-party problem in one line. No single participant in a port call holds the truth about it, and each holds a fragment that is authoritative for their own part and speculative about everyone else's. The same structure repeats at every handover in a supply chain — factory to forwarder, forwarder to carrier, carrier to terminal, terminal to haulier, haulier to warehouse. Fusion across companies is not a data-engineering exercise with a contract attached; it is a contract with a data-engineering exercise attached, and the sequencing follows from that.

Which frame you adopt depends on the mode, and in every mode somebody has already done the work. Container shipping has DCSA's data standards (opens in a new tab); air cargo has IATA's ONE Record (opens in a new tab), which defines a single record view of a shipment shared through a standardised API rather than a chain of per-party message copies; the message layer between trading partners has GS1's EDI standards (opens in a new tab); and the assurance wrapper around all of it — how a management system is documented, reviewed and audited — sits in the ISO catalogue (opens in a new tab). Adopting an existing frame is almost always cheaper than negotiating a bilateral one, and it carries a second-order benefit that compounds: a counterparty already emitting the standard can be added to your fusion without a project, which is what turns cross-party fusion from a series of integrations into a capability.

The trust machinery that has to sit on top of the frame is short and specific. A partner's number can carry weight only when four things are written down: what the field means, how often it updates, how far it may be revised after the fact, and what happens commercially when it is persistently wrong. Add a fifth for fusion specifically — a scorecard the counterparty can see. Making the score visible is what converts a source-quality measurement from surveillance into a contract term, and it is the difference between a supplier engineering team fixing a regression you detected and a supplier account manager disputing it.

What cross-party fusion looks like in public

Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the organisation's own published material.

The clearest public evidence for the fusion thesis is in what large operators chose to standardise before they tried to combine anything. In each programme below the differentiator was not an estimator — it was that the parties agreed what an event meant, who owned it, and on what terms it would be shared, and only then merged the signals into decisions.

Three programmes read against the fusion ladder

Outcomes as published by the organisations themselves; we have not independently audited them. The card images are illustrative industry scenes from our own library — they are not photographs of these operations and imply no endorsement or involvement.

Illustration of a container terminal berth with automated handling equipment and a live data cabinet on the quaysidePort of RotterdamEurope's largest seaport · multi-party port calls24
Challenge
A single port call involves many independent data owners — the port, the terminal, nautical service providers, and services attached to the vessel and the cargo — each authoritative about their own fragment and speculative about the rest. Combining their claims into one arrival and departure picture is impossible while each defines depth, berth and time differently.
Approach
Port Call Optimisation standardised the process first and then the data, adopting international definitions for elements such as berths, depths and arrival and departure times, and making information available through an international API on the principle that data comes from the data owner so that it stays current.
Reported outcome
The port publishes Port Call Optimisation as a programme built on international definitions and standardised nautical, operational and administrative data, and states that the purpose is to make vessel visits as safe and efficient as possible from departure port to destination port, with lower costs for shipping lines, shippers, terminals and ports.
What it shows about the curveThis is the Integrated rung earned properly: the frame before the fusion. Rotterdam did not start by weighting the participants' claims — it started by making the claims comparable, which is the step most operators skip and then spend years compensating for with tie-break rules.

Port of Rotterdam — Port Call Optimisation (opens in a new tab)

Illustration of a high-bay distribution centre with a shared network view projected across the racking and a control room above the aisleWalmartGlobal retailer · supplier-facing data programme34
Challenge
A supplier planning replenishment sees its own shipments and its own forecast, while the retailer holds the sell-through and store-level inventory signal that would sharpen it. Each side fuses what it has; neither can fuse the other's, and the gap shows up as buffer stock on both balance sheets.
Approach
Walmart established Data Ventures as a commercial channel through which its first-party consumer data and insights are made available as a product, so that suppliers can bring the retailer's own demand and inventory signal into their planning under defined terms rather than through informal exchange.
Reported outcome
Walmart publishes Data Ventures as its first-party consumer data and insights business, positioning retailer-held sales and inventory information as a licensed product for the suppliers who sell through it.
What it shows about the curveCross-party fusion survives when the sharing is a product with terms rather than a favour. Voluntary reciprocity decays because the party that shares most gains least in the short run; a priced, contracted feed with defined fields and cadence is the arrangement that is still running in year three.

Walmart Data Ventures (opens in a new tab)

Illustration of a planning team working around a shared table-top view linking factories, distribution centres and storesInditexGlobal fashion retailer · short-cycle supply model34
Challenge
Commercial information arrives continuously from stores and online, while manufacturing and distribution commitments have to be made ahead of it. A fused demand picture is only worth having if a commitment follows it quickly enough for the fusion to matter.
Approach
Inditex's published description of its group model centres on an integrated and flexible operating model in which information from its stores and online channels feeds design, manufacturing, distribution and sales decisions on a short cycle rather than through a long annual planning process.
Reported outcome
Inditex publishes an integrated business model in which commercial information flows continuously into product and distribution decisions, with the group's own reporting describing the tight coupling of its commercial, design and supply functions.
What it shows about the curveFusion pays where a commitment follows it quickly. A fused estimate that changes nothing this week is a dashboard with better mathematics; the value is created at the moment the merged signal shortens the distance between what was observed and what was committed.

Inditex — the group (opens in a new tab)

The fusion stack, layer by layer

What actually has to exist at each rung — and which layer you can defer without lying about your confidence.

A defensible fused estimate requires six layers, and the order in which they are built determines whether the confidence you publish is meaningful. The stack below is deliberately unfashionable: no layer names a vendor, and every layer is defined by what it must guarantee rather than by what product provides it. The layers annotated from rung 4 upwards are the ones that separate a number people believe from a number people can defend.

Layers required by rung

Each layer is annotated with the rung that first requires it. A programme claiming a fused estimate without the source-characterisation layer is publishing a confidence derived from beliefs nobody measured.

  1. Sources and frame

    Stage 1+

    • Domain feedsPOS and orders, capacity confirmations, EDI 214/315, AIS and terminal events, WMS and ERP, weather and risk
    • Identifier resolutionGS1 keys and internal surrogates resolving order, unit, location and invoice to one thing
    • Event dictionaryOne written definition per event, with time base and units fixed
  2. Source characterisation

    Stage 3+

    • Claim logEvery source's claim recorded at the time it was made, not reconstructed later
    • Error storeMeasured error by source × horizon × lane, with sample sizes
    • Correlation registerWhich feeds share an upstream, and the weight cap per cluster
    • Staleness detectionA feed still returning its last value is treated as absent, not as current
  3. Fusion layer

    Stage 4+

    • EstimatorWeighted combination conditioned on horizon; simple until proven insufficient
    • GateTolerance test per source, with firings recorded and routed
    • Confidence computationInterval derived from source variances and cluster caps, widened by risk signals
    • Weight policyVersioned, reviewed like code, with the fitting data referenced
  4. Commitment and write-back

    Stage 4+

    • Value plus interval in the system of recordTwo fields, written together — the TMS, WMS or planning tool the planner already uses
    • Commitment ledgerWhat was promised, on which estimate, at what stated confidence
    • Fallback sourceThe previous single-source value, one switch away and drilled
    • Escalation queueGate firings with both claims attached and a named owner
  5. Provenance and assurance

    Stage 4+

    • Contribution vector storeSource, value, weight and staleness kept with every fused number
    • Retention aligned to disputesDemurrage, detention, customs and service-level windows, not storage defaults
    • Calibration monitorStated confidence compared with realised outcomes, continuously
    • Ablation harnessRecompute without each source on a schedule; publish the marginal contribution
  6. Inter-party layer

    Stage 5+

    • Feed contracts with quality termsDefinitions, cadence, permitted revision, accuracy review and remedy
    • Attestation and lineage tagsWho asserted this value, when, and whether it originated with us
    • Reciprocal scorecardsThe partner can see their own score and the method behind it
    • Competition-law screenPrice and capacity signals between competitors kept out of the shared frame without counsel sign-off

Pipeline described

  1. Sources and frame (stage 1+) — Domain feeds: POS and orders, capacity confirmations, EDI 214/315, AIS and terminal events, WMS and ERP, weather and risk; Identifier resolution: GS1 keys and internal surrogates resolving order, unit, location and invoice to one thing; Event dictionary: One written definition per event, with time base and units fixed
  2. Source characterisation (stage 3+) — Claim log: Every source's claim recorded at the time it was made, not reconstructed later; Error store: Measured error by source × horizon × lane, with sample sizes; Correlation register: Which feeds share an upstream, and the weight cap per cluster; Staleness detection: A feed still returning its last value is treated as absent, not as current
  3. Fusion layer (stage 4+) — Estimator: Weighted combination conditioned on horizon; simple until proven insufficient; Gate: Tolerance test per source, with firings recorded and routed; Confidence computation: Interval derived from source variances and cluster caps, widened by risk signals; Weight policy: Versioned, reviewed like code, with the fitting data referenced
  4. Commitment and write-back (stage 4+) — Value plus interval in the system of record: Two fields, written together — the TMS, WMS or planning tool the planner already uses; Commitment ledger: What was promised, on which estimate, at what stated confidence; Fallback source: The previous single-source value, one switch away and drilled; Escalation queue: Gate firings with both claims attached and a named owner
  5. Provenance and assurance (stage 4+) — Contribution vector store: Source, value, weight and staleness kept with every fused number; Retention aligned to disputes: Demurrage, detention, customs and service-level windows, not storage defaults; Calibration monitor: Stated confidence compared with realised outcomes, continuously; Ablation harness: Recompute without each source on a schedule; publish the marginal contribution
  6. Inter-party layer (stage 5+) — Feed contracts with quality terms: Definitions, cadence, permitted revision, accuracy review and remedy; Attestation and lineage tags: Who asserted this value, when, and whether it originated with us; Reciprocal scorecards: The partner can see their own score and the method behind it; Competition-law screen: Price and capacity signals between competitors kept out of the shared frame without counsel sign-off
Step-by-step insights
Sources and frame — the layer that decides how long everything else takes
Every week spent on identifier resolution and event definitions removes months of downstream tie-break engineering. The reason is structural: a definitional mismatch generates a new disagreement every time the data flows, so the cost of not fixing it is recurring, while the cost of fixing it is paid once. Sequence this layer around a single decision rather than the whole estate — the identifiers and events behind one commitment are a fortnight's work, whereas the same exercise across a network is the kind of programme that gets cancelled in year two.
Source characterisation — the artefact nobody has by accident
This is the layer that separates rung 3 from rung 4, and it is almost never present unless someone deliberately built it. The claim log is the hard part, and it is hard for an unglamorous reason: it must be written at the time the claim is made, because a claim reconstructed later has already been contaminated by hindsight. Start it before you need it. An error store built from six months of logged claims is a competitive asset that cannot be bought or backfilled, and the operators who have one can price a source, negotiate a feed and defend a commitment in ways their competitors simply cannot.
Fusion layer — keep the estimator boring
The estimator attracts disproportionate attention and rarely deserves it. An inverse-variance weighting over well-characterised sources beats a sophisticated model over poorly characterised ones, every time, and it has the practical advantage of being explainable to a counterparty in a dispute. Spend the sophistication budget on the conditioning — horizon buckets, lane segmentation, cluster caps, staleness decay — rather than on the functional form. Where a more expressive model does earn its place, it is usually as a residual correction on top of the simple estimate, which keeps the explanation intact.
Commitment and write-back — two fields, written together
The single highest-leverage engineering decision on this page is that the value and its interval are written in the same transaction into the same record. It sounds trivial; it is the control that prevents the entire failure class this page is named after. It also forces a useful conversation with whoever owns the destination system, because adding an interval column requires stating what downstream consumers should do with it — and that conversation is where booking rules like 'commit to the 80th percentile' get written down rather than improvised per planner.
Provenance and assurance — retention set by the dispute window
Provenance is cheap to keep and impossible to recreate. The contribution vector for a fused estimate is a few hundred bytes; the demurrage dispute that needs it arrives fourteen months later. Set retention from the longest commercial window the commitment can be challenged in — customs, detention, service-level claims and audit all run longer than most default data-retention policies — and store the vector with the commitment rather than in a general log, so that retrieving it is a lookup rather than an investigation.
Inter-party layer — where fusion becomes a contract
The last layer is legal and commercial rather than technical. Its components are contract terms, not services: what a field means, how often it updates, how much it may be revised, what accuracy review applies, and what happens when the score falls. Two further items belong here and are easy to forget. First, lineage tags, so a value you published cannot return as independent evidence. Second, a competition-law screen — fusing capacity and price signals across competing carriers has an antitrust surface, and the shared frame should exclude those fields unless counsel has explicitly cleared them.

The layer most often skipped is source characterisation, and skipping it is what produces the failure this page is named after. Without an error store, a weight is a belief; without a correlation register, a confidence is a guess; and a system that publishes a guess as an interval is more dangerous than one that publishes a bare number, because it has borrowed the authority of measurement without doing any.

A 90-day plan: one fused arrival estimate that books DC labour

The Integrated → Fused step made concrete on one logistics problem — inbound container arrival, fused from four sources, driving a single 72-hour labour commitment.

Ninety days is enough to fuse one decision properly, and nowhere near enough to fuse a function. To make that concrete, the plan below runs the Integrated → Fused transition on a specific, common logistics problem: the inbound container arrival estimate that a distribution centre uses to book agency labour 72 hours ahead. Four sources feed it — the carrier ETA, the terminal berth plan, an AIS-derived vessel position and the lane's realised dwell history — and the quarter contains no new source procurement and no model research. The work is frame, measurement, weighting and write-back.

Integrated → Fused on inbound arrival, in one quarter

One DC, one trade lane, one commitment. If a phase overruns, narrow the scope — fewer lanes, one carrier — rather than extending the window. The deliverable at day 90 is a calibration curve, not a demo.

  1. Days 1–20

    Fix the frame and log the claims

    Choose one DC and one trade lane. Write the event dictionary: what each of the four sources means by arrival, in which time base, keyed to which identifier. Start logging every claim at the moment it is made, alongside the realised gate-in when it lands. Name the inbound operations manager as owner — labour cost and dock congestion are their numbers, so the estimate has to answer to them.

    An event dictionary, resolved identifiers and a live claim log

  2. Days 21–45

    Characterise the sources and find the echoes

    Score each source's error against realised gate-in, split by horizon bucket — 10 days, 5 days, 72 hours, 24 hours — and by carrier. Trace each feed to its upstream and build the correlation register; expect to find at least one pair that is really one source. Set initial weights from the measured errors and cap each cluster.

    An error store by source and horizon, plus a correlation register

  3. Days 46–70

    Run the estimate in shadow, with an interval

    Produce the fused arrival window and its confidence every hour, and store the contribution vector with each one — but commit nothing. Compare the shadow estimate against the incumbent carrier ETA on realised outcomes. Define the gate tolerance and the escalation owner, and let gate firings route to a real person so the alert volume is known before anything depends on it.

    A shadow estimate with a measured calibration curve

  4. Days 71–90

    Let it drive one commitment, and measure the reversals

    Write the window and its confidence into the appointment board and the labour planning tool as two fields, and let the 72-hour booking follow an explicit rule — book to the 80th percentile, take the flexible shift when the window exceeds eight hours. Keep the carrier ETA one switch away as the fallback. Report calibration, sharpness, and reversal rate split by whether the reversal fell inside or outside the stated interval, against a holdout DC still booking on carrier ETA alone.

    A calibrated estimate, an attributable labour delta and a reversal split

The order matters

  1. Frame before weights

    Do not tune a weight around a definitional mismatch. A weight fitted to compensate for the fact that one source means berthing and another means gate-in will be wrong the moment either party changes anything, and it will be wrong invisibly, because the estimate keeps producing plausible numbers.

  2. Shadow before commitment

    Run the estimate in parallel with the incumbent for at least three weeks before anything commits on it. The purpose is not to build confidence in the model; it is to measure the calibration curve, which is the only evidence that lets you write a booking rule with a percentile in it.

  3. Confidence before automation

    Never automate a commitment against a value with no interval. The interval is what lets the rule decline: too wide, take the flexible shift; too wide and expensive, escalate. Automation without that branch converts every uncertain estimate into a confident commitment, which is the failure mode, not the goal.

  4. Ablate before you add

    Before buying a fifth source, recompute the estimate without each existing one and see what the estimate actually loses. Operators routinely discover a feed contributing nothing because it is a repackaging of one they already have — and the money is better spent widening horizon coverage than adding another echo.

Measuring fusion: calibration, sharpness and marginal source value

Accuracy is not the metric. The eight readings that tell you whether a fused estimate is honest, useful and worth what its feeds cost.

A fused estimate is measured on two axes at once: calibration, meaning the stated confidence is honest, and sharpness, meaning the interval is narrow enough to act on. Either alone is worthless. An arrival window of 'somewhere between Tuesday and Friday, 95% confident' is perfectly calibrated and operationally useless; a two-hour window that holds a third of the time is beautifully sharp and actively dangerous. Every metric below exists to keep those two in tension, and none of them is model accuracy.

ReadingHow it is computedSourceCadenceHonest from
CalibrationRealised hit rate within the stated interval, bucketed by stated confidenceCommitment ledger joined to realised outcomesWeeklyRung 4
SharpnessMedian interval width at the decision horizonEstimate logWeeklyRung 4
Frame-error rateShare of investigated disagreements traced to definition, identifier, unit or time-base mismatchEscalation queueMonthlyRung 2
Gate-firing rateSource claims rejected as out-of-tolerance ÷ claims evaluatedGate logDaily, trended weeklyRung 4
Marginal source valueChange in error when the estimate is recomputed without that sourceAblation harnessQuarterly, and before any renewalRung 5
Staleness exposureShare of fused values whose largest weight came from a source older than its refresh intervalContribution vectorsWeeklyRung 4
Provenance completenessShare of commitments whose contribution vector is retrievableCommitment ledgerMonthlyRung 4
Out-of-interval reversal rateCommitments revised where the outcome fell outside the stated interval ÷ all commitmentsCommitment ledgerMonthlyRung 4
The fusion measurement sheet. 'Honest from' is the rung at which the reading first measures something real — asking for calibration before an interval exists produces a number with no referent.

The last reading deserves emphasis because it reframes what counts as a failure. Reversals inside the interval are not errors — they are the estimate working exactly as advertised, and an operation that treats them as failures will pressure the team into narrowing intervals dishonestly. Reversals outside the interval are the real defect, and they should trigger a review of the weights, the correlation register and the staleness handling in that order. Keeping those two categories separate in reporting is the single most useful governance habit in this whole area; the discipline is the same one that underpins NIST's AI Risk Management Framework (opens in a new tab), which treats measurement of a system's stated uncertainty as a first-class control rather than a research nicety. The same habit is visible wherever institutions publish fused numbers for a living: the New York Fed's Global Supply Chain Pressure Index (opens in a new tab) combines several underlying transport-cost and survey series into one published figure, and it publishes the method alongside it — which is exactly the standard a fused ETA should be held to when it is the basis of a commercial commitment.

Fusion readiness checklist

If you cannot tick all eight, the confidence you publish is not yet defensible — regardless of how good the estimate looks on backtest. Tick as you go; this list works without JavaScript.

0 of 8 ticked

Nothing ticked — start with one decision's event dictionary

An empty list is not a bad result; it is an honest one, and it is where most operators genuinely are. Do not begin with tooling. Pick the single commitment that costs you most when it is wrong, write down what each of its sources means by the event in question, and start logging claims this week. Everything else on this list is downstream of those two moves.

Failure modes: how fusion produces one confident wrong number

Five mechanisms account for almost every case where a merged estimate turned out worse than the sources it replaced.

The signature failure of supply fusion is one confident wrong number replacing three honest uncertain ones. It is worse than the problem it solved, because the disagreement between the original sources was itself information — a warning that the situation was unusual — and the fused number destroyed that warning while inheriting the authority of having been computed. Five mechanisms produce it, and each has a cheap, specific preventive measure.

Likelihood: highImpact: high

Correlated sources sold as independent evidence

Two feeds derived from the same upstream — a visibility provider reselling carrier messages you already receive, or two portals rendering one terminal system — are counted as two agreeing witnesses. The computed interval narrows while the true error does not, so confidence rises precisely as accuracy stalls. This is the most common cause of overconfidence in a real estate, and it is almost never visible from the feeds themselves.

PreventionTrace every feed to its upstream, register the clusters, and cap the combined weight of each cluster rather than each member.

Likelihood: highImpact: high

The interval is dropped at the destination field

The estimate is computed with a confidence, then written into a system-of-record field that takes exactly one value, and the uncertainty dies at the boundary. Everything downstream behaves as if the number were certain, which makes the fused estimate strictly worse than the sources it replaced. The failure is invisible until a category of unexplainable incidents accumulates — commitments that were reasonable on the evidence, judged afterwards as mistakes.

PreventionWrite value and interval in the same transaction, and make the write fail if the interval is missing.

Likelihood: mediumImpact: high

Feedback contamination — your own estimate returns as a partner's

You publish a fused ETA to a portal; a counterparty ingests it and posts it back as their own estimate; your scorer reads it as independent corroboration and rewards it with weight. The loop converges on your own opinion with rising confidence and falling accuracy, and every diagnostic looks healthy because agreement is high and gate firings are rare.

PreventionTag the lineage of every inbound value and refuse to score any source against an outcome your own estimate could have influenced.

Likelihood: highImpact: medium

The stale feed that keeps its full weight

A source stops updating but continues returning its last value, so nothing errors and nothing alerts. In a precedence system this is catastrophic and obvious; in a fused system it is subtle, because the stale value keeps voting at full strength and drags the estimate towards a world that ended on Tuesday.

PreventionDecay every source's weight by its age against its own expected refresh interval, and treat a value older than two refresh intervals as absent.

Likelihood: mediumImpact: medium

The tie-break that lives in one person's head

The rules cover the common disagreements and a named expert handles the rest — until they change role. What leaves with them is not documentation but a calibrated sense of which source to believe in unusual conditions, which is exactly the judgement the escalation path depends on. Estimates degrade slowly and nobody can say by how much, because overrides were never recorded with reasons.

PreventionRequire a recorded reason on every override; the accumulated reasons are both the succession plan and the best training data the fusion layer will get.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Supply fusion
Combining supply-chain signals that differ in latency, reliability and coverage — demand, supply and capacity, in-transit, inventory, finance and risk — into a single estimate with a stated confidence, so one commitment can be made against all of them.
Frame
The common basis that makes two sources comparable: resolved identifiers, one written definition per event, a single time base and consistent units. Reconciling the frame precedes any weighting, because most apparent disagreement is frame error rather than signal error.
Contribution vector
The record stored with every fused value listing which sources contributed, what each claimed, what weight each carried, how stale each was at fusion time, and whether the gate fired. It is what makes a commitment explainable months after it was made.
Inverse-variance weighting
Combining estimates in proportion to the inverse of their measured error variance, so more reliable sources count for more and the fused variance is lower than any input's. Valid only to the extent that the sources' errors are not correlated.
Correlated source cluster
A group of feeds derived from the same upstream, which must be weighted as one witness rather than several. Registering clusters and capping their combined weight is the primary defence against a fused interval that is narrower than the true error.
Gate
The tolerance test that rejects a source's claim when it lies further from the fused estimate than both uncertainties allow. Gate firings are recorded and escalated rather than averaged away, because a cluster of them usually means the network has changed.
Calibration
Whether a stated confidence is honest: of the commitments made at 80% stated confidence, roughly 80% should hold. Measured by joining the commitment ledger to realised outcomes and bucketing by stated confidence.
Sharpness
How narrow the stated interval is at the decision horizon. Calibration without sharpness is useless — a very wide interval is trivially honest and operationally worthless — so the two are always reported together.
Source ablation
Recomputing the fused estimate with one source removed to measure what that source actually contributes. Run on a schedule and before any feed renewal, it identifies both redundant purchases and the single source the estimate is quietly leaning on.
Residual
The gap between what a source claimed and what actually happened, measured per source, horizon and lane. Residuals are the raw material of every weight: a fusion that does not collect them cannot improve, only be re-guessed.
Commitment ledger
The log of what was promised, on which estimate, at what stated confidence, and by whom. It is the artefact a demurrage, detention or service-level dispute is answered from, so its retention period is set by the commercial window rather than by storage policy.
Feedback contamination
When a value you published returns to you through a counterparty and is scored as independent evidence, causing the fusion to converge on its own opinion with rising confidence. Prevented by tagging the lineage of every inbound value.

Frequently asked questions

The questions operators ask most often when they start merging supply-chain signals into single decisions.

What is supply fusion?

Supply fusion is combining supply-chain signals of different latency, reliability and coverage into one estimate with a stated confidence. The six domains normally fused are demand signal, supply and capacity, in-transit visibility, inventory position, finance and working capital, and risk. The output is not a screen but a number, an interval around it, and a record of how both were produced — so that a single commitment can be made against all six sources rather than six separate commitments each made against one.

How is supply fusion different from a control tower or a single source of truth?

A control tower shows the sources together; fusion combines them into one estimate with a confidence. The distinction is operational rather than semantic: a screen leaves disagreement resolution to whoever is looking, so the outcome depends on who is on shift, and nothing downstream can consume the result automatically. A single source of truth goes further and picks a winner, which is deterministic but lossy — the discarded sources carried signal, and nobody measures how often the winner was actually right.

Why does a fused estimate beat the best single source?

For three reasons. Errors partially cancel, so a weighted combination of sources whose errors are not perfectly correlated has lower variance than any input. Coverage is complementary — carrier messaging stops at the gate, terminal events start at the quay, warehouse systems begin at the door — so the union covers a journey no single feed does. And a fused estimate degrades gracefully: when one feed dies, the interval widens and weight shifts, rather than the number silently going stale.

How do you weight sources that disagree?

In a fixed order. First reconcile the frame, because a large share of apparent disagreement is definitional — different events, time bases, identifiers or units. Second, cluster feeds that share an upstream and cap the cluster's combined weight. Third, weight by measured error at the decision horizon, since sources rank differently at fourteen days than at thirty-six hours. Fourth, gate what still disagrees: exclude it, record it, and escalate rather than silently averaging it into the estimate.

What is a contribution vector and why does it matter?

It is the small record kept with every fused value, listing which sources contributed, what each claimed, what weight each carried, how stale each was, and whether the gate fired. It matters because a fused number destroys its inputs unless you deliberately keep them, and the questions that need those inputs — a demurrage dispute, a service-level claim, a customs query — arrive months later. It costs a few hundred bytes per estimate and cannot be reconstructed after the fact.

How do you fuse data across companies without giving away commercial advantage?

By making the sharing a contract rather than a favour. Three things need to exist: an identifier and event agreement so the data is comparable, quality terms specifying what a field means, how often it updates and what happens when it is wrong, and purpose and retention terms limiting use. Reciprocity or payment resolves the asymmetry that otherwise kills voluntary sharing, since the party sharing most gains least in the short run. Keep price and capacity signals between competitors out of the shared frame without legal sign-off.

What breaks when a fused number starts driving a commitment?

Provenance and retention break first. While a fused estimate is only informing people, nobody asks how it was produced; the moment it books labour, promises a customer date or releases a purchase order, someone will eventually ask why that commitment was made. The answer must be reconstructable at the length of the commercial dispute window, which is typically far longer than default data retention. Store the contribution vector with the commitment, version the weights, and set retention from the contract rather than from storage cost.

How do you avoid producing one confident wrong number?

Make confidence a first-class field the system of record refuses to write without, so the interval cannot be dropped at the destination. Then remove the three mechanisms that inflate it: register correlated feeds and cap their combined weight, decay each source's weight by its age so a stale feed stops voting at full strength, and tag lineage so a value you published cannot return as independent evidence. Finally, monitor calibration — if commitments made at 80% confidence hold far less often, the confidence is not honest.

Which domain should we fuse first in a logistics operation?

In-transit visibility, in almost every case. It has the most sources, the clearest realised outcome to score against — the gate-in actually happened at a recorded time — and the shortest feedback loop, so an error store accumulates in weeks rather than quarters. It also touches the commitments that cost most when wrong: labour bookings, appointment slots and customer promises. Demand fusion pays more in the long run but scores slowly, because the outcome you are predicting arrives on a much longer cycle.

How do you measure whether fusion is working?

On two axes together: calibration, meaning the stated confidence is honest, and sharpness, meaning the interval is narrow enough to act on. Either alone is meaningless — a very wide window is trivially calibrated and useless, and a very narrow one that holds a third of the time is dangerous. Around those, track gate-firing rate as a network-change indicator, staleness exposure, provenance completeness, and the out-of-interval reversal rate. Model accuracy is not on the list.

Do we need new systems, or can this run on our existing TMS and WMS?

It runs on the systems you have, with two additions. The first is an estimation service that reads your existing feeds and writes back — it does not replace the TMS, WMS or ERP, it produces a field they display. The second is a second column: the destination record needs to hold the interval next to the value. That schema change is usually the largest single piece of integration work in the whole programme, and it is worth doing first, because everything downstream depends on it existing.

How long does it take to get from Aligned to Fused?

Roughly two to three quarters when it is scoped to one decision, and several years when it is scoped to a function. The elapsed time is dominated by two things that cannot be compressed: fixing the frame for the sources behind that decision, and accumulating enough realised outcomes to characterise each source's error by horizon. The estimator itself is typically a week. Scoping to one commitment — one DC, one lane, one booking rule — is what keeps the programme inside a year.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for logistics, manufacturing and energy operators — forecasting, estimation, optimisation and decision support running against live operational data and written back into the TMS, WMS, YMS and planning layer rather than delivered as dashboards.

  • · Estimation and fusion pipelines running against live carrier, terminal and warehouse feeds
  • · Source-quality instrumentation designed jointly with operations and commercial teams
  • · Integration-first delivery: system-of-record write-back, provenance, monitoring, rollback
  • · 18 cited sources on this page

Sources

  1. GS1GS1 standards (opens in a new tab)
  2. GS1EPCIS event-sharing standard (opens in a new tab)
  3. GS1GS1 identification keys (opens in a new tab)
  4. GS1Global Data Synchronisation Network (GDSN) (opens in a new tab)
  5. GS1GS1 EDI standards (opens in a new tab)
  6. DCSAContainer shipping data standards (opens in a new tab)
  7. IATAONE Record data-sharing standard (opens in a new tab)
  8. ISOISO standards catalogue (opens in a new tab)
  9. ECMWFEnsemble-based operational forecasts (opens in a new tab)
  10. Federal Reserve Bank of New YorkGlobal Supply Chain Pressure Index (opens in a new tab)
  11. NISTAI Risk Management Framework (opens in a new tab)
  12. NOAAProbabilistic weather and environmental products (opens in a new tab)
  13. Massachusetts Institute of TechnologyCenter for Transportation & Logistics (opens in a new tab)
  14. GartnerSupply chain research and insights (opens in a new tab)
  15. World Economic ForumCentre for Advanced Manufacturing and Supply Chains (opens in a new tab)
  16. Port of RotterdamPort Call Optimisation (opens in a new tab)
  17. WalmartData Ventures (opens in a new tab)
  18. InditexThe group and its business model (opens in a new tab)

Get one number your whole operation can commit against

We run the assessment with your engineering and operations leads, build the signal ledger and correlation register for one decision, score your sources against twelve weeks of realised outcomes, and leave you with a costed plan to put a calibrated estimate into the system your planners already use. You keep the ledger and the error table either way.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.