LogisticsAI Adoption & Maturity Curve
AI adoption KPIs in logistics: measuring every stage of the maturity curve
AI adoption KPIs are the metrics that prove an AI programme is actually maturing — not just running. In logistics they split into adoption KPIs (decision latency, write-back coverage, acceptance rate) and value KPIs (dock dwell, OTIF, cost per shipment), and the right set changes at every stage of the maturity curve.

Key takeaways
- AI adoption KPIs come in two families: adoption KPIs that prove the programme is maturing (decision latency, write-back coverage, acceptance rate) and value KPIs that prove it is paying (dock dwell, OTIF, cost per shipment). Reporting only one family is how programmes mislead themselves.
- The right KPI set changes at every stage of the maturity curve: at stage 1 you can only measure data readiness, at stage 2 accuracy and usage, from stage 3 acceptance and operational deltas, and at stage 5 automation and escalation rates.
- Decision latency is the single most predictive adoption KPI. If a recommendation reaches a planner more than one planning cycle after the data that produced it, the programme is capped at stage 2 regardless of model quality.
- Model accuracy is a model property, not an adoption KPI. Every value claim needs a holdout — a lane, door bank or shift kept on the old process — or the number will not survive a budget review.
- Baseline before the pilot, not after: six months of dwell, OTIF and cost-per-shipment history from the TMS, WMS and YMS is what makes every later KPI delta attributable.
Abbreviations used on this page
- TMS
- Transport management system
- WMS
- Warehouse management system
- YMS
- Yard management system
- APS
- Advanced planning system
- ERP
- Enterprise resource planning
- S&OP
- Sales and operations planning
- 3PL
- Third-party logistics provider
- OTIF
- On-time in-full delivery rate
- ETA
- Estimated time of arrival
- EDI
- Electronic data interchange (e.g. the 214 shipment status message)
- AEO
- Authorised Economic Operator (customs programme)
- GLEC
- Global Logistics Emissions Council framework
Free · 8 questions · ~3 minutes
Score your operation on the curve
Eight questions, one at a time, about three minutes. Answer them and we build your personalised maturity report — your stage on the curve, your score on each of the four dimensions, and the specific blocker standing between you and the next stage — and send it to your inbox. Your result doubles as your first adoption-KPI baseline.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, how you compare with operators of similar network shape, and the 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Ad hoc
AI exists as individual experiments with no shared data, no owner and no route into an operational decision.
Your next movePick one decision and instrument the data behind it end to end — one route, one lane, one facility. Not a platform.
Stage 2 · Pilot
One or more models work in a bounded proof of value, but their output reaches operations through a human reading a dashboard.
Your next moveWrite the model's output back into the system that already runs the decision — the TMS, WMS or planning tool — even if a human still approves it.
Stage 3 · Repeatable
AI output is delivered into the operational system on a schedule, with monitoring, retraining and a named owner.
Your next moveExtract the shared layer — feature store, monitoring, deployment — so use case four is configuration rather than a project.
Stage 4 · Integrated
AI is a shared platform capability: multiple decisions are served from common data and infrastructure, with value measured in operational terms.
Your next moveDefine the confidence and risk thresholds under which a decision executes without approval, and the audit trail that makes that safe.
Stage 5 · Autonomous
Defined decisions execute without human approval inside agreed risk bounds, with humans handling exceptions and setting policy.
Your next moveTreat the decision policy as a versioned, reviewable artefact with the same rigour as the model itself.
0 / 24
Data foundation
— / 6
Operational integration
— / 6
Governance & ownership
— / 6
Value measurement
— / 6
Your score maps to a stage on the curve. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your maturity, and it is where the next investment belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the curve. The dimension breakdown matters more than the total: the lowest dimension is what actually caps your maturity, and it is where the next investment belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this benchmarked against comparable operators?
We will walk your engineering and operations leads through the dimension scores, compare them against operators of similar scale and network shape, and leave you with a costed 90-day plan for the weakest dimension. No obligation, and you keep the plan either way.
How the score maps to a stage
- 0–5 — Stage 1, Ad hoc. AI exists as individual experiments with no shared data, no owner and no route into an operational decision.
- 6–11 — Stage 2, Pilot. One or more models work in a bounded proof of value, but their output reaches operations through a human reading a dashboard.
- 12–16 — Stage 3, Repeatable. AI output is delivered into the operational system on a schedule, with monitoring, retraining and a named owner.
- 17–21 — Stage 4, Integrated. AI is a shared platform capability: multiple decisions are served from common data and infrastructure, with value measured in operational terms.
- 22–24 — Stage 5, Autonomous. Defined decisions execute without human approval inside agreed risk bounds, with humans handling exceptions and setting policy.
What AI adoption KPIs are — and why they change by stage
A definition, the two KPI families, and the maturity curve that decides which of them you can honestly measure.
AI adoption KPIs are the metrics that measure how far AI has travelled from experiment to execution inside a logistics operation. They split into two families that must be reported together: adoption KPIs, which prove the programme is maturing — decision latency, write-back coverage, recommendation acceptance — and value KPIs, which prove it is paying, in the units the operation already runs on: dock dwell, OTIF, cost per shipment, picks per labour hour.
Which KPIs you can honestly measure depends on where you sit on the AI adoption maturity curve — the five-stage model beneath every number on this page. A stage-2 operator has no attributable value KPI yet, only accuracy; a stage-3 operator can put dwell minutes against a holdout; only a stage-5 operator has an automation rate worth quoting. Reporting a KPI before its stage is how programmes mislead themselves — the number exists, but nothing connects it to the operation.
Value released against time on the curve
The curve is not linear. Value stays close to flat through stages 1 and 2 — where most operators are — and inflects at stage 3, when output starts reaching the system that runs the decision. This is why programmes that measure progress in models built rather than decisions changed report activity without results.
Operational value released by stage
- Stage 1 · Ad hoc — 18% of operators. AI exists as individual experiments with no shared data, no owner and no route into an operational decision.
- Stage 2 · Pilot — 41% of operators. One or more models work in a bounded proof of value, but their output reaches operations through a human reading a dashboard.
- Stage 3 · Repeatable — 27% of operators. AI output is delivered into the operational system on a schedule, with monitoring, retraining and a named owner.
- Stage 4 · Integrated — 11% of operators. AI is a shared platform capability: multiple decisions are served from common data and infrastructure, with value measured in operational terms.
- Stage 5 · Autonomous — 3% of operators. Defined decisions execute without human approval inside agreed risk bounds, with humans handling exceptions and setting policy.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with McKinsey's supply-chain AI research.
How AI output reaches a logistics decision at each stage
The software path, stage by stage. The stage is determined by where the arrow ends: stages 1–2 terminate at a human reading something, stage 3–4 write into the TMS/WMS with monitoring and rollback around the model, and stage 5 executes within a versioned policy. Most operators are in the top lane.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- At stages 1–2, data leaves the TMS and WMS as manual CSV exports, feeds a model on an analyst's laptop, and surfaces on a BI dashboard behind a separate login. Whether the planner acts on it is optional — adoption decays exactly when the model matters most, at peak. This is where value leaks.
- At stages 3–4, EDI 214 messages, telematics and WMS events stream continuously into a governed feature store. A served, drift-monitored model writes its recommendation into the TMS/WMS field the planner already works in, and the planner approves each one with a one-switch fallback to the previous source.
- At stage 5, a versioned decision policy lets routine decisions — slot assignment, replenishment triggers — execute automatically inside agreed bounds. Anything outside the bounds escalates to a human, and every automated action carries a reconstructable audit trail.
Step-by-step insights
- CSV exports — the habit that caps everything above it
- The manual export is the single most predictive artefact of a stalled programme. Every extract is stale the moment it lands, carries no lineage, and encodes one analyst's private filter choices — so two people 'using the same data' quietly are not. Nothing downstream of a hand-pulled CSV can ever be more current, more governed or more repeatable than the export habit itself, which is why fixing ingestion is always the first move out of stage 1, ahead of any modelling work.
- The dashboard dead end
- A dashboard requires no integration approval, which is exactly why pilots ship one — and exactly why they stall there. The recommendation lives behind a separate login, outside the planner's working screen, so acting on it is a voluntary extra step. Voluntary steps are the first thing dropped under operational pressure, and peak season — when the model is worth most — is when adoption reliably collapses. No standard operating procedure changes because a chart exists.
- Event streams and the feature store
- EDI 214 status messages, telematics pings and WMS events arrive continuously, so the data layer should too. The feature store's real contribution is not technology but agreement: one definition of dwell time, one definition of on-time, versioned, with named owners. Operators consistently report that the reconciliation meetings — arguing about whose number is right — simply end once a governed shared layer exists, and every later use case inherits that agreement for free.
- The served model — monitored where logistics actually drifts
- Logistics models do not drift randomly; they drift on schedule. A carrier-bid cycle changes the carrier mix, a new DC changes flow paths, peak changes everything. Drift monitoring should therefore be keyed to those known events, not just to statistical thresholds — and retraining should run on a calendar the operation recognises. A model nobody retrains after bid season is a model quietly answering last year's network.
- Write-back and the approval log
- Writing the recommendation into the TMS/WMS field the planner already reads removes the voluntary step: the default action becomes the informed one. The approval step is not a concession — it is data collection. Every accept and override, with context, is the training set for tomorrow's autonomy thresholds. Operators who skip straight to automation have no such log and end up setting bounds by guesswork.
- Policy, bounds and the escalation rate
- Stage 5 is a policy artefact, not a model artefact: versioned thresholds that state which decisions may execute unattended and within what limits. The most useful operational signal is the escalation rate — the share of decisions falling outside bounds. When it rises, the world has moved outside the policy's validity (new lanes, new carriers, new seasonality) and the policy needs review before an incident forces one. The audit trail on every automated action is what makes the whole arrangement defensible to a customer or regulator.
The five stages in detail
For each stage: what it actually looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.
Each stage below is written for a practitioner rather than a buyer. The hallmarks describe observable conditions, the diagnostic signals are checks you can run against your own systems this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Ad hoc
18% of operators sit here
AI exists as individual experiments with no shared data, no owner and no route into an operational decision.
Stage 1 is not the absence of capability — it is the absence of a supply chain for that capability. The models are often genuinely good. What is missing is any repeatable path from operational reality to a model and back again, so every piece of work begins by reconstructing the world from CSV exports.
The tell is where the data lives. At stage 1 the authoritative version of 'last quarter's dwell times' is a spreadsheet on somebody's laptop, reconciled by hand, and two analysts asked the same question will produce two different numbers for defensible reasons. Nobody is wrong; there is simply no shared definition to be right about.
This is a cheap stage to leave and an expensive stage to stay in. The cost is not the failed experiments — it is that every experiment amortises nothing, so the tenth pilot costs exactly what the first one did.
In practice
The recurring dwell-time question
A regional 3PL's operations director asks quarterly why dock dwell is rising at three sites. Each quarter an analyst pulls WMS exports, reconciles them against the yard management spreadsheet, builds a model, and produces a deck. The deck is good. Next quarter the work starts again from zero, because nothing from the last run was written down anywhere a system could read.
What it looks like
- Analysts build models in notebooks against exported spreadsheets
- No single system of record for operational data
- Results are presented, not consumed by any workflow
- No named owner accountable for AI outcomes
Diagnostic signals you can check this week
- Ask two teams for the same metric and compare the numbers
- Ask where a model's training data came from — if the answer is a filename, you are here
- Check whether any AI output has a scheduled refresh
- Ask who gets paged if a model is wrong. If there is no answer, there is no model in production
Anti-pattern · Buying the platform first
The instinctive fix is a data platform programme — eighteen months, a lakehouse, a governance council. It is the most reliable way to spend a year without reaching stage 2, because the platform's requirements are being guessed rather than observed. Instrument one decision end to end first; the platform's real shape is visible after use case three, not before use case one.
What holds you here
There is no reliable, current dataset to build on, so every experiment starts by rebuilding the data.
Highest-leverage next move
Pick one decision and instrument the data behind it end to end — one route, one lane, one facility. Not a platform.
Cost of leaving
- Effort
- 3–6 months
- Team
- One data engineer, one operations analyst, part-time
- Risk
- Low — the work is additive and nothing in production depends on it yet
- To next stage
- 3–6 months
If this is you, the next step is
A 2-week engagement: pick the decision, map the data path, size the build.
Stage 2
Pilot
41% of operators sit here
One or more models work in a bounded proof of value, but their output reaches operations through a human reading a dashboard.
Stage 2 is the most dangerous stage on the curve, because it looks like success. The model beats the baseline, the deck lands well, the sponsor is pleased, and the programme has produced exactly zero operational change. Every quantitative measure of the pilot is green and the P&L is untouched.
The structural reason is that a pilot is optimised to answer 'does this work?' and the organisation's real question is 'will anyone act on it?'. Those have different success criteria and different builds. Answering the first well can make the second harder, because the fastest path to a demonstrable result is a dashboard, and a dashboard is precisely the artefact that leaves the decision unchanged.
Time spent at stage 2 is not neutral. Planners learn that AI output is advisory, sponsors learn that AI does not move numbers, and the next proposal is funded against that memory. Operators who sit at stage 2 for three years are usually harder to move than operators at stage 1, because the organisational antibodies are established.
In practice
The forecast nobody used
A freight operator built a volume forecast that cut error against the planners' own baseline by a meaningful margin on backtest. It shipped as a Power BI page. Six months later, adoption analytics showed the page was opened a handful of times a week, always by the analytics team. Planners had a working process and a screen they already lived in; the forecast lived somewhere else, and asking them to check it was asking them to add a step during peak.
What it looks like
- A model demonstrably beats the baseline on historical data
- Output lands in a BI dashboard or a weekly report
- Scope is one lane, one site or one customer
- Nobody has changed a standard operating procedure yet
Diagnostic signals you can check this week
- Count how many SOPs changed because of the pilot. Usually zero
- Check dashboard usage analytics against the planning team roster
- Ask a planner to show you where they see the model output in their normal day
- Ask what happens to the pilot's output during peak week — if the honest answer is 'nobody looks', integration is the gap
Anti-pattern · Improving the model to drive adoption
When a pilot is not adopted, the reflex is to make it more accurate on the theory that trust follows precision. It rarely does. Adoption is a function of where the output appears, not how good it is: a 70%-accurate recommendation inside the planning tool changes more decisions than a 90%-accurate one behind another login. Spend the next quarter on the write-back path, then revisit accuracy when you can measure what an accuracy point is worth in shipments.
What holds you here
The pilot proves accuracy but never earns a place in the execution system, so its value depends on a human choosing to act on it.
Highest-leverage next move
Write the model's output back into the system that already runs the decision — the TMS, WMS or planning tool — even if a human still approves it.
Cost of leaving
- Effort
- 6–12 months
- Team
- One integration engineer, one ML engineer, a named operations owner
- Risk
- Medium — the first write into a system of record needs a rollback path
- To next stage
- 6–12 months
If this is you, the next step is
The stage 2→3 transition is our most common engagement. Typically 90 days.
Stage 3
Repeatable
27% of operators sit here
AI output is delivered into the operational system on a schedule, with monitoring, retraining and a named owner.
Stage 3 is the first stage where the programme survives its founders. Output has a destination, a schedule, an owner and an alarm, which together mean the capability continues working when the person who built it moves teams. That is the actual definition of production, and most organisations reach it later than they think they have.
The character of the work changes here. Stage 1 and 2 problems are analytical; stage 3 problems are operational, and the discipline that solves them is closer to site reliability engineering than to data science. What is the error budget for this prediction? What is the paging policy? What is the rollback? These questions have well-established answers in software operations, and importing them wholesale is faster than rediscovering them.
The constraint that emerges is throughput. Each use case is still built and operated as its own thing, so the team's capacity is consumed by maintenance at roughly the fourth deployment. Programmes that miss this plateau spend a year adding use cases and wondering why velocity fell.
In practice
The fourth use case that never shipped
A parcel operator shipped three integrated models in eighteen months — ETA prediction, sort-centre volume, and driver assignment scoring. Each had its own ingestion, its own monitoring cron and its own on-call rota entry. The fourth was scoped, approved and never delivered: the team's entire capacity had been absorbed by keeping the first three healthy through a network redesign and a peak season.
What it looks like
- Predictions are written into the TMS/WMS, not just displayed
- Retraining runs on a schedule, and drift is monitored
- A second use case reuses the first one's data pipeline
- Someone is accountable for the model's operational performance
Diagnostic signals you can check this week
- Count the distinct monitoring implementations. More than one means the platform layer is missing
- Measure elapsed time from idea to production for the most recent use case, then the one before
- Ask what fraction of the team's week is maintenance. Above 50% and you have hit the plateau
- Check whether use case two reused use case one's features or rebuilt them
Anti-pattern · Declaring victory and scaling headcount
The plateau reads as a resourcing problem, so the response is to hire. Adding engineers to a set of independently-operated use cases raises the operating burden roughly linearly and buys less than expected. The leverage is in extracting the shared layer — feature definitions, serving, monitoring, deployment — so the fourth use case is configuration. Do that with the team you have before growing it.
What holds you here
Each use case is still built and operated separately, so the fourth costs as much as the first and the team becomes the bottleneck.
Highest-leverage next move
Extract the shared layer — feature store, monitoring, deployment — so use case four is configuration rather than a project.
Cost of leaving
- Effort
- 9–18 months
- Team
- Platform engineer, ML engineer, SRE-style on-call rotation
- Risk
- Medium — the refactor competes with new use-case demand for the same people
- To next stage
- 9–18 months
If this is you, the next step is
We map your existing use cases and identify what is genuinely shareable.
Stage 4
Integrated
11% of operators sit here
AI is a shared platform capability: multiple decisions are served from common data and infrastructure, with value measured in operational terms.
At stage 4 the marginal cost of a new decision collapses. Because features, serving, monitoring and deployment are shared, the work of adding a use case is mostly specification: which decision, which metric, which threshold. Operators at this stage stop talking about 'AI projects' and start talking about which decisions are on the roadmap, which is a different and much healthier conversation.
The measurement discipline is what distinguishes stage 4 from a well-engineered stage 3. Every served decision has a named operational metric and, ideally, a holdout — a lane, a region or a shift kept on the previous process so the difference is attributable rather than asserted. This is unglamorous and it is the thing that keeps the programme funded through a budget cycle where someone asks what the AI actually did.
The remaining constraint is human. Every decision still passes through an approver, so total throughput is bounded by planner capacity rather than by the system. That is often the correct place to stop — the question of whether to go further is a risk-appetite decision, not a technical one.
In practice
The roadmap conversation
An operator with a shared feature and serving layer reached the point where adding a new decision — carrier scoring for a newly acquired region — took three weeks, most of it spent agreeing the target metric and the holdout design with operations. The engineering was two days. That ratio, specification-heavy and build-light, is the signature of stage 4.
What it looks like
- Common feature and serving layer across use cases
- New use cases ship in weeks, not quarters
- Model performance is tracked against a business metric, not accuracy alone
- Exception handling and escalation paths are defined
Diagnostic signals you can check this week
- Time from decision agreed to decision served, for the last three use cases
- Whether a single monitoring dashboard covers every served model
- Whether any served decision has a live holdout group
- Whether an operations leader, not an engineer, can name what each model is worth
Anti-pattern · Automating because you can
Stage 4 makes autonomy technically easy, which is exactly when it gets extended past the evidence. Thresholds derived from a quarter of approval logs on one decision type get applied to decision types that were never in those logs. The first bad automated decision then results in all automation being switched off, and the programme loses more ground than autonomy ever gained.
What holds you here
Humans still approve every action, so throughput is bounded by planner capacity rather than by the system.
Highest-leverage next move
Define the confidence and risk thresholds under which a decision executes without approval, and the audit trail that makes that safe.
Cost of leaving
- Effort
- 18+ months
- Team
- Platform team, operations product owner, risk/compliance partner
- Risk
- Higher — governance and audit evidence become the binding constraint
- To next stage
- 18+ months
If this is you, the next step is
Which decisions should execute unattended, and the evidence to prove it is safe.
Stage 5
Autonomous
3% of operators sit here
Defined decisions execute without human approval inside agreed risk bounds, with humans handling exceptions and setting policy.
Stage 5 is narrower than it sounds. It is not an autonomous supply chain; it is a specific, enumerated set of decisions that execute unattended inside stated bounds, with everything outside those bounds escalating to a person. Slot assignment, replenishment triggers and routine carrier selection qualify. Anything with regulatory exposure, safety implications or high commercial variance is correctly held at stage 4 forever.
The engineering is largely solved by the time an operator arrives here. The hard part is the evidence: demonstrating to an auditor, a customer or a regulator that a decision made without a human was made correctly, with a reconstructable trail, and that the policy under which it was made was reviewed and versioned. Treat the decision policy with the same rigour as the model — it is the artefact that will be examined.
Sustaining stage 5 is a governance discipline rather than a technical one, and it is the stage most likely to regress. Networks change, thresholds drift out of validity, and the audit trail that satisfied last year's review does not satisfy this year's.
In practice
The bounded decision set
A high-volume operator runs unattended slot assignment across its network inside explicit bounds — value ceiling, customer tier, exception rate. Roughly one decision in twenty escalates to a planner. The escalation rate itself is monitored: a rise indicates the world has moved outside the policy's validity, and it triggers a review before it triggers an incident.
What it looks like
- Routine decisions execute automatically within thresholds
- Humans manage exceptions and policy, not individual decisions
- Full audit trail on every automated action
- Rollback and kill-switch procedures are tested, not theoretical
Diagnostic signals you can check this week
- Whether the decision policy is versioned and reviewed like code
- Whether the kill switch has been exercised in the last six months
- Whether escalation rate is monitored as a leading indicator
- Whether an auditor could reconstruct any single automated decision from logs
Anti-pattern · Treating the policy as configuration
Thresholds get tuned in a settings screen with no review, no version history and no record of who changed what on which date. The system works right up until someone has to explain a decision made eight months ago, at which point neither the model nor the threshold that produced it can be reconstructed. Version the policy, review changes, keep the trail.
What holds you here
Sustaining autonomy is a governance problem — the constraint becomes regulatory evidence and change control, not engineering.
Highest-leverage next move
Treat the decision policy as a versioned, reviewable artefact with the same rigour as the model itself.
Cost of leaving
- Effort
- Continuous
- Team
- Platform team plus a standing governance forum
- Risk
- Concentrated — low frequency, high consequence, regulatory in nature
If this is you, the next step is
We stress-test the policy, the trail and the rollback against a real scenario.
The KPI catalogue: what to measure at every stage
Adoption KPIs prove you are maturing; value KPIs prove it is paying. The honest set for each stage of the curve — and the trap each stage reports instead.
Every stage of the maturity curve has KPIs it can honestly report and KPIs it cannot yet support. The catalogue below is the full measurement ladder for a logistics operation: read your stage's row, instrument those metrics from the systems named in the decision map further down, and treat the right-hand column as the early-warning list — each trap is the number that stage most often reports in place of the truth.
| Stage | Adoption KPIs — is it maturing? | Value KPIs — is it paying? | The trap |
|---|---|---|---|
| 1 · Ad hoc | Share of target decisions with instrumented data; export-to-insight lead time; count of conflicting metric definitions | Baselines only: six months of dock dwell, OTIF and cost per shipment from the TMS, WMS and YMS | Reporting model demos as adoption |
| 2 · Pilot | Accuracy against the planning baseline (forecast bias, error); planner sessions on pilot output; lanes and doors covered | None attributable yet — value claims at stage 2 are projections | Celebrating accuracy while no SOP changes |
| 3 · Repeatable | Decision latency within one planning cycle; write-back coverage; recommendation acceptance vs override; freshness-SLA compliance; retrain cadence adherence | Dwell minutes and detention charges vs holdout; picks per labour hour; OTIF delta on covered lanes | Acceptance quietly falling after go-live, with nobody watching it |
| 4 · Integrated | Time-to-production per use case; share of use cases on shared features; portfolio model-health composite | Attributed savings per quarter, holdout-verified; cost per served decision | A portfolio dashboard hiding one decaying model |
| 5 · Autonomous | Automation rate within policy bounds; escalation-rate trend; policy version age; audit reconstruction time | Planner hours redeployed; decisions handled per planner; incident-free automated volume | Extending thresholds to decisions the approval log never covered |
Two disciplines make the whole catalogue trustworthy. First, adoption and value KPIs are always reported as a pair — an acceptance rate without a dwell delta is theatre, and a dwell delta without an acceptance rate is unexplainable. Second, every value KPI is measured against a holdout — a comparable lane, door bank or shift left on the previous process — because in a live logistics network, seasonality and carrier-mix changes will otherwise claim the credit or take the blame.
Where AI lands in a logistics network
Warehouse, yard, linehaul, last-mile, planning and compliance — the decisions worth wiring, the system each one lives in, and the KPI it moves.
AI value in logistics concentrates in six operating domains, and each domain has a natural home on the curve. A decision is a good first candidate when three things are true: the system of record is already yours, the decision cycle is short enough to measure inside a quarter, and the KPI it moves is one a budget holder already tracks. The map below is how we scope first and second use cases with operators.
| Domain | High-value decisions | System of record | KPI it moves | Sweet spot |
|---|---|---|---|---|
| Warehouse & fulfilment | Slotting and re-slotting, labour planning, wave release | WMS / LMS | Dock-to-stock, picks per labour hour | Stage 3–4 |
| Yard & dock | Trailer slot assignment, door scheduling, dwell prediction | YMS / WMS | Dock dwell, detention charges | Stage 4–5 |
| Linehaul & network | Load consolidation, carrier selection, dynamic routing | TMS | Cost per shipment, empty miles | Stage 3–4 |
| Last-mile | Route sequencing, in-day re-optimisation, promised ETA | Dispatch / route planner | Stops per hour, first-attempt delivery | Stage 4–5 |
| Planning & S&OP | Demand forecast, capacity and workforce planning | APS / ERP | Forecast bias, OTIF | Stage 2–3 |
| Compliance & customs | Document classification, HS coding, dangerous-goods checks | Customs / broker platform | Clearance time, audit findings | Stage 3–4 |
Warehouse AI enablement is where most operators should begin. Slotting, labour planning and wave release run on systems the operator already controls, the feedback loop is measured in shifts rather than quarters, and dock-to-stock time and picks per labour hour are KPIs nobody disputes. Linehaul and last-mile decisions carry larger absolute savings — empty miles and failed first attempts are expensive — but they touch carrier contracts and customer promises, so their approval paths are longer and they belong later in the sequence.
The compliance domain is the sleeper. Logistics operates under ISO 28000 (opens in a new tab) security management, GDP for pharmaceutical lanes, AEO and C-TPAT customs programmes, and — increasingly — emissions accounting under ISO 14083 and the GLEC Framework (opens in a new tab). Every one of these asks for the same artefact a stage-3 write-back produces as a by-product: a reconstructable trail of what was decided, by what rule, on what data. Operators who reach stage 3 typically find audit preparation getting cheaper, because the evidence is generated by the system instead of assembled for the audit.
Where logistics operators actually sit today
The distribution across the curve, and why the stage 2 → 3 drop is the largest transition loss.
Most logistics operators are at stage 2. The distribution is heavily weighted toward pilot-stage work: a majority have at least one model that demonstrably beats their planning baseline, and a small minority have that model changing what happens on the ground without a person in the loop.
Distribution of logistics operators across the five stages
Stage 2 is the mode and the plateau. The drop from stage 2 to stage 3 is the largest single transition loss on the curve.
Share of operators
- 18% — 1 · Ad hoc
- 41% — 2 · Pilot (the plateau)
- 27% — 3 · Repeatable
- 11% — 4 · Integrated
- 3% — 5 · Autonomous
Source: Illustrative distribution, synthesised from McKinsey, BCG and MHI adoption research
This is not a logistics-specific failure. Cross-industry research has consistently found the gap between organisations experimenting with AI and organisations reporting material bottom-line impact to be wide and persistent — see McKinsey's State of AI (opens in a new tab) and BCG's analysis of where AI value lands (opens in a new tab). What is logistics-specific is the shape of the blocker, which is almost always the write-back path into the TMS or WMS rather than the model itself. MHI's annual industry survey (opens in a new tab) tracks the same adoption-versus-impact gap across material handling and supply chain.
What the transitions look like in public
Two publicly reported deployments, read against the curve. None is an Atomic Loops engagement — each links to the operator's own published material.
The clearest evidence for the integration thesis is in what large operators chose to build. In each case below the differentiator was not model sophistication — it was that the output was wired into the system that dispatches, routes or plans, and that the operating discipline around it was built at the same time.
Two deployments read against the curve
Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them.
UPSGlobal parcel network · 500k+ employees24
- Challenge
- Route sequencing decisions were made by drivers and dispatchers using experience and static route plans, with no systematic way to apply network-level optimisation to the daily dispatch.
- Approach
- ORION embedded route optimisation directly into the dispatch and driver-facing systems rather than presenting recommendations separately, and was later extended toward continuous, in-day re-optimisation rather than a fixed morning plan.
- Reported outcome
- UPS has publicly reported ORION delivering annual mileage reductions in the region of 100 million miles and associated cost savings in the hundreds of millions of dollars per year.
- What it shows about the curveThe value came from the decision moving into the execution path, not from the optimiser being novel. Route optimisation as a research problem was decades old; putting it in the dispatch loop was the change.
DHLGlobal 3PL · contract logistics & express34
- Challenge
- Scaling AI capability across a very large number of facilities with heterogeneous systems, where a per-site build would never amortise.
- Approach
- Treating AI as a repeatable capability rolled out across sites — standardised deployment patterns and shared infrastructure — rather than as bespoke per-site projects.
- Reported outcome
- DHL publishes ongoing research and deployment reporting on applying AI across warehousing, transport and customer operations at network scale.
- What it shows about the curveThe stage 4 signature is the marginal cost of the next site or use case falling. Where each deployment is bespoke, the programme plateaus regardless of how good any single model is.
The four dimensions that set your stage
Maturity is not one number. Four dimensions gate each other, and the lowest is the real stage.
Maturity is not a single number. An operation is scored on four dimensions — data foundation, operational integration, governance and ownership, and value measurement — and the lowest of the four is the real stage, because each one gates the others. A stage-4 model served from a stage-1 data foundation degrades silently and nobody notices.
Data foundation
Freshness, shared definitions and lineage. The binding question is whether two teams asking for the same field get the same number. Until they do, every use case pays a data-rebuilding tax and no result is comparable across the network.
Operational integration
Where the output lands and how fast it gets there. This is the dimension that separates stage 2 from stage 3, and it is overwhelmingly the lowest-scoring dimension in the assessments we run.
Governance and ownership
Whether a named person is accountable for a model's production behaviour, and whether there is a tested path to roll it back. The operational patterns here are borrowed almost wholesale from site reliability engineering — error budgets, paging policy, blameless review — and Google's SRE book (opens in a new tab) remains the most useful reference for teams importing them.
Value measurement
Whether success is defined in operational terms — cost per shipment, on-time delivery rate, dock dwell — and whether that improvement can be attributed. Accuracy is a model property, not a business outcome, and programmes measured on accuracy alone lose funding at the first budget review.
Diagnosing the real constraint
Plot your data foundation against your operational integration. The quadrant tells you what the next investment should be — and three of the four common answers are not 'build a better model'.
Blocked at the last mile
- Good data, no route into operations
- Highest-leverage position on the matrix
- Fix: build the write-back path, not another model
Scaling
- Both foundations in place
- Constraint is now delivery throughput
- Fix: extract the shared platform layer
Experimenting
- Neither foundation in place
- Common at stage 1
- Fix: instrument one decision end to end
Fragile automation
- Integrated but built on unstable data
- The most dangerous quadrant
- Fix: freshness monitoring before any further automation
Why stage 2 is where programmes stall
Three structural patterns account for most of the plateau, and none of them is a modelling problem.
Stage 2 stalls because the pilot was scoped to prove accuracy, and accuracy was never the constraint. A pilot that beats the planning baseline by a meaningful margin has answered a question nobody was really asking; the open question is whether a planner under time pressure will act on it, and the answer is usually no unless the recommendation appears inside the tool they already have open.
The output has no home
The pilot ships a dashboard because a dashboard is the fastest thing to build and needs no integration approval. But a dashboard shifts the burden of action onto the planner, who already has a process that works. Adoption depends on discipline rather than on the system, and discipline decays — fastest during peak, which is exactly when the model is worth most.
The pilot was scoped to a lane, and the value case needs a network
Single-lane pilots are chosen because they are easy to isolate, but the savings from a single lane rarely clear the threshold for a platform investment. The pilot succeeds and the business case fails, which reads internally as the AI having failed.
Nobody owns the operational outcome
The pilot has a data science owner and an executive sponsor, and no owner in operations. When the pilot ends there is no one whose targets improve if it continues, so it does not continue. The fix is structural and cheap: name the operations owner before the build, not after the demo.
The hardest part of scaling AI in supply chain is not the model — it is redesigning the decision process the model is supposed to serve.