Redefining Technology

LogisticsAI-Driven Disruptions & Innovations

Disruptive and sustaining AI in logistics: telling the two apart before you fund them

Sustaining AI improves the performance your existing customers already buy on. Disruptive AI changes who can serve the job, on what asset base and at what price. In logistics almost every pound of realised AI value has been sustaining, and the discipline that matters is separating the two before they compete for one budget line.

Freight operations scene with a quoting desk, dock activity and linehaul yard shown side by side
Logistics · AI-Driven Disruptions & Innovations

Key takeaways

  1. Sustaining AI improves a metric your existing customers already buy on — cost per shipment, OTIF, dwell, picks per labour hour. Disruptive AI changes who can serve the job, through what channel, on what asset base. The technology does not decide which one you have; the customer and the asset base do.
  2. Almost all realised AI value in logistics has been sustaining. The largest publicly reported wins — UPS's ORION mileage reduction, Amazon's robot fleet — came from making an existing network cheaper, not from changing who moves the freight. Treating that as a disappointment is the most expensive mistake on this page.
  3. The disruptive foothold in freight is the business you decline. Small-shipper LTL, one-pallet loads, sub-scale lanes and anything that needs a firm price in seconds are invisible to every KPI you report, because every KPI is computed over volume you accepted.
  4. Every disruptive path in logistics is bounded by an asset, a certification or a piece of physics — ADS rules and operational design domains for driver-out linehaul, airspace and payload for drones, building geometry and capital for robotics. A bet that has not named its bound cannot be timed and is not an option.
  5. Sustaining work is funded from the operating budget against an operational KPI and a holdout. Disruptive probes are funded from a capped, separate pot against a learning milestone and a written kill rule. One business-case template for both is why speculative work either overstates its payback or loses the budget round.

Abbreviations used on this page

TMS
Transport management system
WMS
Warehouse management system
YMS
Yard management system
3PL
Third-party logistics provider
LTL
Less-than-truckload freight
FTL
Full-truckload freight
OTIF
On-time in-full delivery rate
ETA
Estimated time of arrival
EDI
Electronic data interchange (e.g. the 204 load tender)
API
Application programming interface
ADS
Automated driving system
ODD
Operational design domain — the conditions an automated system is certified to operate in

Free · 8 questions · ~3 minutes

Score how your operation handles the split

Eight questions, one at a time, about three minutes. Answer them about your AI portfolio as it is actually governed — not as the strategy deck describes it — and we build your personalised report: your rung on the disruption-response ladder, your score on each of the four dimensions, and the specific thing standing between you and the next rung.

0 of 8 answered

Question 1 of 8Classification discipline

How are the initiatives on your current AI portfolio list described?

If everything is grouped as "AI transformation", the label carries no information and cannot change a funding decision.

How the score maps to a stage
  • 05 — Stage 1, Undifferentiated. Every AI initiative is called transformation, so nothing is either sustaining or disruptive — the label carries no consequence for funding, metrics or governance.
  • 611 — Stage 2, Sustaining by default. AI is used competently to make the existing network cheaper and faster, and the operator has no mechanism at all for noticing a disruptive move.
  • 1216 — Stage 3, Classified. Initiatives are explicitly labelled sustaining or disruptive, and the label changes how each one is funded, measured and stopped.
  • 1721 — Stage 4, Instrumented. The operator watches named leading indicators of disruption in its own transactional data, with thresholds that trigger a costed decision rather than a discussion.
  • 2224 — Stage 5, Optioned. The operator holds a small number of priced, dated options on the disruptive paths, sized so that being wrong is affordable and being right is exercisable.

What disruptive and sustaining AI mean in logistics

The definition, the three paths an initiative can take through a freight business, and why the technology never decides which one you have.

Sustaining AI improves the performance your existing customers already buy on; disruptive AI changes who can serve the job, through what channel, and on what asset base. In logistics that distinction is concrete rather than philosophical. A model that cuts empty miles on lanes you already run, or predicts trailer dwell so the dock scheduler turns doors faster, is sustaining: the customer is the same, the asset is the same, and the improvement lands in the TMS or WMS as a better version of a decision you already make. A quoting path that gives a two-pallet shipper a firm, binding price in four seconds is not a better version of anything you do — it serves freight your desk currently declines, through a channel your desk does not staff.

The theory is Clayton Christensen's, and the Christensen Institute's own statement of it (opens in a new tab) is worth reading before applying it to freight, because the popular usage has drifted a long way from the original. Disruption in the technical sense is not "a big change" or "a scary competitor". It describes a process that begins at the bottom of a market — in applications the incumbent finds unattractive, typically because they are less profitable and more awkward — and then moves upmarket. That definition has an uncomfortable implication for a logistics operator: the disruptive threat is not the thing your best customers are asking for. It is the thing your commercial team is quietly turning down.

Disruptive Innovation describes a process by which a product or service takes root in simple applications at the bottom of the market — typically by being less expensive and more accessible — and then relentlessly moves upmarket, eventually displacing established competitors.

The three paths an AI initiative takes through a freight business

The path is set by who buys the improvement and what it needs to work, not by how new the technology is. The top lane is where nearly all realised value sits. The bottom lane is not a real path at all — it is what happens when a sustaining initiative is funded on a disruptive story, and it is where budget dies quietly.

  • Data & feeds
  • AI / model
  • System-of-record action
  • Human in the loop
  • Where value leaks

The process, in words

  • On the sustaining path, an existing customer's metric is the target. The model trains on the operator's own TMS, WMS and YMS history, its output is written back into the field a planner already reads, the planner approves with a one-switch fallback to the previous source, and unit cost falls across the same customers and the same assets. This path is unglamorous, measurable against a holdout, and responsible for nearly all the logistics AI value anyone has publicly reported.
  • On the disruptive path, the starting point is freight the operator declines — one-pallet LTL, sub-scale lanes, anything needing a firm price in seconds. The initiative changes the channel or the unit of sale rather than the cost of an existing decision, and it then hits a gate that has nothing to do with model quality: an asset it does not own, a certification it does not hold, airspace, capital or hours-of-service physics. That gate is what the board is really deciding about, so the correct output is a priced, dated option with a kill rule — not a programme.
  • The third lane is not a path through the business at all. It is sustaining work funded on a disruptive story: no operational metric because it has been declared strategic, no bound and no kill rule because it is meant to be visionary, and a quiet cancellation two budget cycles later. The damage outlives the spend, because the next genuinely speculative proposal is funded against the memory of this one.
Step-by-step insights
Why the technology never decides the lane
The same model can sit in either of the top two lanes depending on who buys the result. A price-prediction model that helps your pricing desk quote contract lanes faster is sustaining — same shipper, same lanes, a decision you already make, better. The same model exposed as a public API that returns a binding price to a shipper you have never served is on the disruptive path, because it serves a job you currently decline through a channel you do not staff. Teams argue about whether generative AI or reinforcement learning is 'disruptive' as though the answer were a property of the algorithm. It is a property of the customer and the asset base, and you can settle it in a meeting by asking who pays and what they were doing instead.
The write-back is what makes the sustaining lane pay
The sustaining lane only produces value at the third node. A dwell prediction on a separate dashboard changes nothing, because acting on it is a voluntary extra step and voluntary steps are the first thing dropped at peak — exactly when the model is worth most. Writing the recommendation into the appointment board, the tender screen or the rating table makes the informed action the default action. Operators who skip this and go straight to arguing about model quality typically spend a year improving accuracy that nobody consumes. The approval step is not a concession either: the accept-and-override log is the dataset that later sets any autonomy threshold you might want.
The decline register is the entry point to the middle lane
Every KPI in a logistics reporting pack is computed over accepted volume. Cost per shipment, OTIF, dwell, picks per labour hour, empty miles — all of them require a shipment you actually moved. The consequence is structural blindness at exactly the place Christensen's theory says to look. A decline register fixes it cheaply: twelve months of inbound requests from the quoting engine, the CRM and the phone log, classified by lane, weight break, requested service and channel, with a reason code on every decline. It is two weeks of work, it needs no new system, and it is the single highest-value artefact on this page for a stage-2 operator.
The gate node is the whole decision
Node d2 is where most freight disruption theses actually terminate, and it is the node that strategy decks skip. Driver-out linehaul is gated by automated-driving rules and a narrow operational design domain, not by perception quality. Drone delivery is gated by airspace approval, payload and energy density. Unattended commercial commitment is gated by liability and by whether you can reconstruct a decision for an auditor. Writing the gate down converts an argument about the future into a question with an owner: what would have to change, who would tell us it had changed, and what would we do that week?
The mislabelled lane is the expensive one
The bottom lane costs more than its budget line, and the extra cost is organisational. Every cancelled 'strategic' initiative teaches the business that speculative work is expensive theatre, which raises the evidential bar for the next probe — including the one that would have been right. It also consumes the scarcest input in this whole process, which is not money but executive attention in the forum where classification happens. The cheapest defence is procedural: apply the classification test out loud, in the room, before the funding conversation, and make the test one sentence long so nobody can negotiate with it.

One further clarification is worth making early, because it saves a recurring argument. Sustaining does not mean small, incremental or unambitious. UPS's ORION programme is a sustaining innovation by this definition — same customers, same parcels, same trucks — and UPS has publicly reported annual mileage reductions in the region of 100 million miles from putting route optimisation into the dispatch loop. Amazon's robotics fleet is a sustaining investment inside a network Amazon already owns, and it is the largest deployment of mobile robots in the world. If your classification test tells you that most of your best AI opportunities are sustaining, the test is working correctly, and the correct response is to fund them faster.

The five rungs of the disruption-response ladder

How operators actually get better at this: from one business-case template that flattens everything, to a small register of priced, dated options with named bounds.

Operators improve at telling the two apart along a recognisable five-rung ladder, and the rungs are governance changes rather than technology changes. Each rung below is written for a practitioner: the hallmarks describe observable conditions, the diagnostic signals are checks you can run against your own paperwork and quoting data this week, and the anti-pattern is the specific mistake most often made trying to leave that rung. Note that the ladder measures how you handle the split — not how much AI you have. A stage-2 operator can be running excellent production models and still be structurally unable to see a low-end entrant.

Select a rung

Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Undifferentiated

22% of operators sit here

Every AI initiative is called transformation, so nothing is either sustaining or disruptive — the label carries no consequence for funding, metrics or governance.

Stage 1 is not ignorance of the distinction. Most leadership teams can define disruptive innovation on request, and many can name the canonical examples. What is missing is any point at which the definition touches a decision. A linehaul consolidation model, a dock-scheduling model and a speculative autonomous-yard trial arrive in the same slide pack, compete for the same line, and get judged on the same criteria — which in practice are enthusiasm, vendor credibility and how well the story survives a board slide.

The tell is the business-case format. At stage 1 there is one template and it asks every initiative for a payback period. A sustaining initiative can answer that honestly: it improves a number the operation already tracks, on volume it already handles, and the arithmetic is checkable. A probe cannot. Its output is information about whether a market exists, not a return. So it either invents a payback line or it loses to the initiative that can produce one. Both outcomes are damaging, and both are caused by the form rather than by anyone's judgement.

This is a cheap stage to leave and an expensive one to occupy, because it starves both ends at once. The sustaining work is under-funded relative to its near-certain return, because it competes against better stories. The genuinely speculative work is over-specified, because it has to promise a return nobody can know. The organisation ends up with a portfolio optimised for narrative quality, which is not a property that shows up in cost per shipment.

In practice

The transformation deck

A regional 3PL's annual planning pack listed eleven AI initiatives, from an ETA model to a drone-feasibility study, each with a three-year payback line in the same column. Two were funded — the two with the best-looking payback. One was a sustaining ETA project that delivered on schedule. The other was the drone study, whose payback line had been reverse-engineered from the budget it needed. Nobody was dishonest. The form had only one shape, and the drone study filled it in.

What it looks like

  • Every initiative is described as AI or digital transformation regardless of what it changes
  • One business-case template, and it asks every entry for a payback period
  • Sustaining improvement and speculative work compete for the same budget line
  • No initiative has ever been stopped on a rule agreed in advance

Diagnostic signals you can check this week

  • Ask for the list of live AI initiatives and count how many carry a label other than "AI" or "transformation"
  • Check whether any two initiatives on that list are held to different success criteria
  • Read the business-case template — if every entry needs a payback period, speculative work cannot survive it honestly
  • Ask who would have to be wrong for an initiative to be stopped; at stage 1 there is no such person

Anti-pattern · Standing up an innovation lab and calling it classification

The reflex fix is an innovation function: move everything speculative into it and let the operating business get on with efficiency. What usually moves is the label, not the discipline. The lab's projects are still judged on payback at the next budget round, still have no kill rule, and are now also cut off from the quoting and tender data that would tell anyone whether the market thesis is real. Classification is a change to how decisions are recorded and reviewed, not a change to the org chart, and doing the org chart first tends to postpone the governance work by a year.

What holds you here

There is one business-case template and it asks every initiative for a payback period, so speculative work either overstates or loses.

Highest-leverage next move

Split the business-case template in two: an operational-metric case for sustaining work, and a learning-milestone case with a written kill rule for probes.

Cost of leaving

Effort
1–2 months
Team
One operations leader, one finance partner, half a day a week
Risk
Low — the work changes how decisions are recorded, not any production system
To next stage
1–3 months

If this is you, the next step is

A half-day working session: every live initiative labelled, with the metric it should be judged on.

Classify your current AI portfolio

Stage 2

Sustaining by default

38% of operators sit here

AI is used competently to make the existing network cheaper and faster, and the operator has no mechanism at all for noticing a disruptive move.

Stage 2 is where most competent logistics operators sit, and it is a legitimate place to be. Linehaul consolidation, dwell prediction, appointment scheduling, tender-acceptance scoring, carrier performance models and document extraction are all sustaining innovations. They all pay. They all improve a metric the operator's existing customers already buy on, delivered into the TMS or WMS where the decision actually happens. Judged on realised value rather than on column inches, this is where nearly all logistics AI value has come from.

The blind spot is structural rather than intellectual. The measurement system points exclusively at the freight you serve. Cost per shipment, OTIF, dock dwell, picks per labour hour and empty miles are all computed over accepted, executed volume. That means the requests you declined, the quotes you lost because the price took four hours, and the shippers who never called because you do not quote in seconds are invisible by construction. An operator can improve every number in the pack for three years while its addressable market quietly narrows, and no report will contain a row that says so.

Time at stage 2 is not neutral either, because the vocabulary hardens. The organisation learns that "AI" means efficiency — which is true and incomplete — and after a few cycles an "AI project" is understood internally to be a cost-reduction project. A proposal that is not a cost-reduction project then has nowhere to go: it cannot be written on the form, it cannot be argued in the forum, and the person who would have argued it stops trying.

In practice

The quote nobody counted

A mid-sized LTL carrier improved cost per shipment for three consecutive years on sustaining AI: better linehaul consolidation, better dock scheduling, better carrier scoring. Over the same period its share of one-pallet, next-day, small-shipper freight fell steadily, because those requests arrived by phone and took most of a working day to price. The monthly pack showed three years of improvement. It contained no row for freight the company did not haul, so the trend that mattered was never on a page.

What it looks like

  • AI is delivering real, measured savings on existing lanes, sites and doors
  • Every live initiative improves a metric existing customers already buy on
  • No reporting line exists for declined, lost or never-requested freight
  • The response to a disruption story in the trade press is a one-off pilot, not a measurement

Diagnostic signals you can check this week

  • Ask for your decline rate on inbound quote requests, split by weight break — if nobody can produce it, the low end is invisible
  • Check whether a single reported KPI is computed over volume you did not accept
  • Ask how long it takes to return a firm price on non-contract freight — measured from the quoting system, not estimated
  • Ask who won your last five lost tenders; if the answer is "the market", nobody is watching

Anti-pattern · Buying a moonshot to prove you are not complacent

The classic stage-2 response to a disruption story is to fund something visibly futuristic — a drone trial, an autonomous yard tractor, a pilot with a research partner — with no thesis about which customer it serves and no rule for stopping it. It reads as boldness and functions as insurance against the accusation of complacency. It also consumes precisely the budget and executive attention that building a decline register would have needed, and it teaches the organisation that "disruptive" means "expensive and eventually cancelled", which makes the next honest probe harder to fund.

What holds you here

Every reported KPI is computed over accepted volume, so the freight you decline — the classic low-end foothold — is invisible by construction.

Highest-leverage next move

Instrument the freight you say no to. Build a decline register from your quoting and tender records before you fund anything speculative.

Cost of leaving

Effort
3–6 months
Team
A pricing or commercial analyst, one data engineer, a named commercial owner
Risk
Low — the work is measurement of transactions you already touch
To next stage
3–6 months

If this is you, the next step is

Two weeks: every quote you lost or declined for twelve months, classified by lane, weight break and channel.

Build your decline register

Stage 3

Classified

24% of operators sit here

Initiatives are explicitly labelled sustaining or disruptive, and the label changes how each one is funded, measured and stopped.

Stage 3 begins the moment the label does work. A sustaining initiative is funded from the operating budget, owned by the operations leader whose KPI it moves, and judged against a holdout — a lane set, door bank or shift left on the previous process so the improvement is attributable rather than asserted. A disruptive probe is funded from a separate capped pot, owned by whoever would have to build the resulting business, and judged on whether it answered a stated question by a stated date. Same organisation, two different definitions of success, and neither one borrowed from the other.

Two artefacts appear at this stage, and they are what a reviewer should ask for. The first is a portfolio list with a label on every row and a different metric type by label. The second is a probe register with a kill rule per row. The kill rule is the hard part and the one most often fudged. "We will stop if it is not working" is not a rule; it is a sentence. "We stop if fewer than forty shippers accept an instant quote on this lane set within eight weeks" is a rule, because it can be false.

What stage 3 still does not have is any independent signal about whether a disruptive thesis is correct. The classification is applied to whatever ideas the operator happened to have — which came from vendors, conferences and the trade press. That is a very large improvement on stage 2, and it is still a portfolio of opinions with better paperwork around it. The next move is not more ideas; it is instrumentation.

In practice

Two forms, two owners

A 3PL rebuilt its investment paperwork into two forms. The operational form asks for the KPI, the holdout design and the operations owner. The probe form asks for the question, the evidence that would answer it, the cost cap, and the date the answer is due. In its first cycle, two probes were stopped on their own rules at a combined cost below what the single unstopped pilot of the previous era had consumed in two years. Nobody had to argue either one down; the rule had been written by the sponsor.

What it looks like

  • Every portfolio row carries a written sustaining or disruptive label
  • Two business-case forms exist, with genuinely different success criteria
  • The probe budget is capped, separate, and cannot borrow from operations
  • At least one probe has been stopped on a rule agreed before it started

Diagnostic signals you can check this week

  • Ask to see the portfolio list — every row should carry a label, and the metric type should differ by label
  • Ask for the last probe that was stopped, and which pre-agreed rule stopped it
  • Check that the sustaining and probe budgets are separate lines that cannot borrow from each other mid-year
  • Ask a probe owner what question their work answers; a description of a technology is the wrong answer

Anti-pattern · Relabelling sustaining work as disruptive to dodge the payback test

Once two forms exist, the probe form becomes the easier one to complete, and initiatives that are plainly sustaining start migrating onto it to escape a payback number. The symptom is a probe register full of rows whose stated "question" is really a delivery plan with a question mark added. The countermeasure is a single test applied at classification, in the forum, out loud: does this improve a metric an existing customer already buys on? If yes, it is sustaining, however novel the technology inside it happens to be.

What holds you here

The classification is applied to whatever ideas the operator already had, with no independent signal about whether any disruptive thesis is real.

Highest-leverage next move

Instrument the market: pick the signals that would move first if an entrant were taking your low-end freight, and put agreed thresholds on them.

Cost of leaving

Effort
3–6 months
Team
Commercial owner, operations owner, finance partner, quarterly review forum
Risk
Medium — the first stopped probe is a political event and needs visible cover from the top
To next stage
6–12 months

If this is you, the next step is

We draft both forms, the classification test and the kill-rule template against your live portfolio.

Design the two-track review

Stage 4

Instrumented

12% of operators sit here

The operator watches named leading indicators of disruption in its own transactional data, with thresholds that trigger a costed decision rather than a discussion.

At stage 4 the operator stops depending on the trade press to notice that the market has moved. The signals are in data it already owns: the decline register, the channel mix on inbound bookings, the time-to-firm-price distribution, the win rate banded by shipper size, and the identity of whoever won the last twenty lost tenders. External market series — spot-to-contract spread, tender rejection — are joined in as context that tells you whether a loss was cyclical or structural, not as the primary instrument. Own-data first is not a stylistic preference; published indices lag the shipper behaviour they summarise.

The discipline that separates this from a dashboard is the threshold. Each signal has a number attached, agreed before the first reading, and crossing it triggers a specific action with a named owner and a date. Without that, the panel becomes a reading exercise, and it will be read charitably — because everyone on the call has an interest in the current strategy being correct, and a chart with no threshold can always be described as noise for one more quarter.

The second thing stage 4 buys is calibration on speed, which is where most disruption arguments actually go wrong. In freight, structural change has consistently been slower than its advocates predicted and faster than incumbents' planning cycles assumed. Instrumented operators stop arguing about whether something will happen and start arguing about the rate — and the rate is a question their own quote log can settle, which makes the argument shorter and considerably less enjoyable for everyone.

In practice

The channel-mix threshold

A carrier's commercial team agreed one number in advance: if the share of inbound bookings arriving through an API or self-serve quote channel — as opposed to phone, email or an EDI 204 tender from a contracted shipper — passed a stated level in any region, the board would fund a productised quoting path within one quarter. The threshold was crossed in a single region eighteen months later. Because the response had been costed when the threshold was set, the argument at that meeting was about sequencing, not about whether the data meant anything.

What it looks like

  • A named signal panel built from the operator's own quote, tender and booking data
  • Every signal carries a threshold agreed before it was first read
  • Crossing a threshold triggers a pre-costed action with an owner and a date
  • Someone in the business can name who won the last twenty lost tenders

Diagnostic signals you can check this week

  • Ask to see the signal panel and its thresholds — a panel without numbers is a reading exercise
  • Ask what actually happened the last time a threshold was crossed
  • Check whether the decline register refreshes automatically from the quoting system or is assembled by hand each quarter
  • Ask whether anyone can name the counterparty that won the last twenty lost tenders, coded incumbent or entrant

Anti-pattern · Watching the market instead of the customer

Instrumented operators tend to over-invest in external market series — rate indices, tender rejection, capacity surveys — because they are easy to buy and comfortable to discuss in a room. They are also lagging by construction: by the time a published index reflects a channel shift, the shippers who moved did so quarters earlier. The leading signals sit in the operator's own quote log, its decline reasons and its loss codes, and they are unglamorous enough that nobody volunteers to present them. Fund the boring panel first and buy the index as context.

What holds you here

Signals exist and get read, but the organisation holds no priced option it could actually exercise when one of them crosses.

Highest-leverage next move

Convert the two or three live theses into priced, dated options with an explicit exercise cost, a named binding constraint and a retirement test.

Cost of leaving

Effort
6–12 months
Team
Commercial analytics owner, one data engineer, the forum that owns the thresholds
Risk
Medium — the real risk is a crossed threshold that nobody honours
To next stage
12–18 months

If this is you, the next step is

Eight signals wired to your quoting and TMS data, with thresholds agreed before the first reading.

Build the disruption signal panel

Stage 5

Optioned

4% of operators sit here

The operator holds a small number of priced, dated options on the disruptive paths, sized so that being wrong is affordable and being right is exercisable.

Stage 5 is much narrower than "we do disruptive innovation". It is a register of two to four options, each with a stated thesis in one sentence, a cost of holding, a cost of exercising, a review date, and the signal that would trigger exercise. In freight the realistic entries are unglamorous: a productised instant-quote channel for weight breaks the network already handles; a driver-out corridor contingent on a certification path actually opening; a data product built from flows the operator already moves. Options that read like press releases tend not to survive their first pricing.

The discipline that keeps the stage honest is retirement. Options expire, and an option held past its date without a decision has quietly become a programme — one that consumes operating attention while reporting into a forum designed for probes. Reviewing the register on a calendar, and retiring something in most cycles, is what stops the portfolio silting up. The health check is simple: ask what was retired last time. A register that only ever grows is a project list wearing an option's vocabulary.

The binding constraint at stage 5 is not imagination. Every genuinely disruptive path in logistics runs into an asset, a certification or a piece of physics. Driver-out linehaul runs into automated-driving rules and a narrow operational design domain. Drone delivery runs into airspace approvals, payload and energy density. Port and terminal automation runs into capital and labour agreements. Unattended commercial commitment runs into liability and audit evidence. An option that has not written down its bound is not an option, it is a wish — and the register should state the bound and what would have to change for it to move.

In practice

The register that expires

One operator's register holds three rows. Each states the thesis in a sentence, the annual cost of holding, the estimated cost of exercising, the signal that would trigger it, the binding constraint and a review date. At the last review one row was exercised into a funded build, one was extended with a written reason and a new date, and one was retired because its binding constraint — a certification path — had not moved in two years. The retired row cost less over its whole life than a single quarter of the exercised one.

What it looks like

  • A written option register with two to four rows, not fifteen
  • Every row states its thesis, cost of holding, cost of exercising and review date
  • Every row names the asset, certification or physical bound that gates it
  • Something is retired or exercised in most review cycles

Diagnostic signals you can check this week

  • Ask for the option register — three rows with dates is healthier than fifteen without
  • Ask what was retired at the last review; a register that only grows is a programme list
  • Check that every row names its binding constraint: asset, certification, capital or physics
  • Ask what exercising would cost and who would run it — an option nobody could staff is not exercisable

Anti-pattern · Treating an option as a commitment

The characteristic stage-5 failure is escalation of commitment. An option accumulates sunk cost, internal advocates and a narrative, and the review that should retire it turns into a review of whether to double down. The countermeasure is procedural rather than cultural: the register names, in advance, the evidence that would retire the row, and the review opens with that evidence rather than with the advocate's update. Changing the running order of the meeting does more here than any amount of talk about being willing to fail.

What holds you here

Sustaining stage 5 is a retirement discipline — options silt into programmes unless something is killed in most cycles.

Highest-leverage next move

Put a review date and an explicit retirement test on every row, and hold the review even when nothing has changed.

Cost of leaving

Effort
Continuous
Team
A standing portfolio forum — commercial, operations and finance — meeting on a calendar
Risk
Concentrated — few decisions, each consequential, all of them reputational

If this is you, the next step is

We take each row, price the exercise, name the binding constraint and test the retirement evidence.

Stress-test your option register

Where logistics operators actually sit on this ladder

The distribution, and why the jump from rung 2 to rung 3 is a paperwork change that almost nobody makes.

Most logistics operators sit at rung 2 — sustaining by default. They have real AI in production, it is delivering measured savings, and they have no mechanism whatsoever for noticing a disruptive move, because every number they report is computed over freight they accepted. The distribution below is heavily weighted toward that rung, and the drop from rung 2 to rung 3 is the largest single transition loss on the ladder despite being, on paper, the cheapest move: it is a change to two forms and one budget rule.

Distribution of logistics operators across the five rungs

Illustrative distribution. Rung 2 is both the mode and the plateau: competent sustaining delivery with no instrumentation of declined freight. Rung 5 is small by design rather than by failure — an option register is a narrow artefact and most operators do not need one.

Share of operators

  • 22% — 1 · Undifferentiated
  • 38% — 2 · Sustaining by default (the plateau)
  • 24% — 3 · Classified
  • 12% — 4 · Instrumented
  • 4% — 5 · Optioned

Source: Illustrative distribution, synthesised from MHI, Gartner and McKinsey adoption research

22%

at rung 1 — one template, one payback column

Illustrative distribution, sourced above

60%

at rungs 1–2 with no line of sight to declined freight

Illustrative distribution, sourced above

16%

at rungs 4–5 with thresholds or a written option register

Illustrative distribution, sourced above

This is not a logistics-specific failure of nerve. Research houses tracking AI across sectors — Gartner's supply-chain AI programme (opens in a new tab), MHI's annual industry survey (opens in a new tab) and McKinsey's operations research (opens in a new tab) — have consistently found a wide gap between organisations experimenting with AI and organisations reporting material impact. What is logistics-specific is the shape of the gap on the disruptive side: the industry's measurement systems are built entirely around executed shipments, so the market segment where a foothold would form is not merely under-monitored, it is absent from the schema.

The other thing the distribution hides is that the rungs are not evenly hard. Rung 1 to rung 2 is a delivery problem and most operators solve it — it is the subject of our companion page on AI adoption KPIs across the logistics maturity curve. Rung 2 to rung 3 is a governance problem that costs almost nothing and is skipped anyway, because it requires someone senior to say out loud that a favourite initiative is sustaining. Rung 3 to rung 4 is a data problem with a two-week fix. Rung 4 to rung 5 is a discipline problem that never ends.

What AI has actually disrupted in logistics so far

Separating what is running in production today from what published rules and research actually claim — and naming the constraint that holds the rest in place.

Very little of the structure of the freight industry has been disrupted by AI, and a great deal of its cost base has been improved by it. That sentence is unfashionable and it is what the evidence supports. The transaction layer has genuinely changed — instant binding quotes and API booking are real, at scale, and they altered how a large share of dry-van truckload and parcel gets bought. The asset layer has not: capacity is still trucks, drivers, hours of service, buildings and berths, and every AI claim that requires the asset layer to change runs into an asset, a certification or a piece of physics. The table below is the separation, claim by claim.

The claimWhat is actually in productionWhat is still speculation — and whyThe binding constraint
Driverless linehaul removes the driver from freight costDriver-out operation on selected fixed corridors in a small number of jurisdictions, run by named developers under published safety cases with a narrow, documented operational design domainA general driver-out network across mixed weather, unmapped roads and urban pickup and delivery. The ODD is the product, and it is deliberately narrow; every mile outside it is a mile with a driverAutomated-driving rules for commercial vehicles, the safety case and insurance — plus the physical inspection, coupling and yard move a tractor still needs
Instant pricing turns freight brokerage into softwareInstant, binding quotes across a large share of dry-van truckload and parcel, API booking, automated tracking and exception handling at scale by named operatorsThe claim that the asset base becomes irrelevant. Capacity is still trucks, drivers and hours of service, and gross margin still moves with the spot-to-contract spread across the cycleCapacity is physical and cyclical. A software front end changes the transaction, not the truck — which is why the incumbents that survived built the same front end
Warehouse robotics makes labour a solved problemVery large fleets of mobile robots inside purpose-built buildings, plus vision-based induction, sortation, containerised storage and AI-scheduled fleet coordinationRobot-first retrofits of arbitrary legacy buildings at similar economics. Today's largest fleets run in buildings designed around them, and the retrofit case is a different oneBuilding geometry, floor flatness, power and dock configuration; capital per square metre; and the variety of the pick face the robots have to serve
Drones and pavement robots replace the last mileBounded commercial operations at limited scale under specific aviation approvals, with published payload, range and airspace limitsSubstitution for general parcel. Payload, weather, noise and airspace rules bound the addressable share to a specific set of light, urgent items rather than to the parcel streamAviation regulation and airspace integration, payload against energy density, and the ground infrastructure the flights still depend on
Agentic AI will run the supply chain end to endAgents drafting tenders, chasing exceptions, reading and classifying documents, and pre-populating TMS and WMS records with a human approving the commitmentUnattended commercial commitment at scale. The constraint here is not model capability — it is who is liable for a wrong commitment and whether the decision can be reconstructed months laterContractual liability, and audit evidence under security and customs regimes such as ISO 28000 and Authorised Economic Operator programmes
Shared data will make the network interoperableStandardised identification and event capture — GS1 identification keys and EPCIS events — inside consortia, large shipper programmes and regulated lanesA universal network where any parcel routes through any carrier's assets. The blockers are commercial rather than technical: rate structures and customer relationships are the productCommercial incentive. Nobody wants to publish the data that would let a competitor price their customer, and no standard changes that
Six recurring claims about AI disrupting logistics, separated into what is genuinely in production, what remains speculation, and the constraint that decides the timing. Read the last column first: it is the one that can be monitored.

Two of those rows deserve a note, because they are where operators most often get the timing wrong in opposite directions. Driver-out linehaul is bounded by rules for automated driving systems on commercial vehicles — the US federal regulator's remit (opens in a new tab) covers driver qualification, hours of service, inspection and roadside enforcement, and none of those disappear because the cab is empty. The correct posture is not scepticism but specificity: name the corridor, name the operational design domain, name the rule that would have to change, and monitor it. Interoperability is the mirror image. The standards exist and have for years — GS1's EPCIS event standard (opens in a new tab) is mature and widely implemented — and the thing that has not moved is commercial willingness. An operator waiting for a technical breakthrough there is waiting for the wrong event.

The practical consequence of the whole table is the page's central claim: future-readiness in logistics is mostly present-readiness. Every one of the speculative rows, if it arrives, arrives into an operation that needs the same things the sustaining work needs — clean event data with agreed definitions, write-back into the system of record, monitoring, rollback and a reconstructable decision trail. An operator that has built those has bought a cheap option on every row in the table. An operator that has funded a drone trial instead has bought an option on one row, at a higher price, with no residual value if that row does not move.

The disruption test: four questions and a 2×2

A test you can apply in the room, in under a minute, before the funding conversation starts.

The test is four questions, and the first one settles most cases on its own. Ask whether the initiative improves a metric an existing customer already buys on. If it does, it is sustaining — fund it from the operating budget, give it an operations owner and a holdout, and stop discussing it in strategic terms. If it does not, ask the remaining three: who is the customer, and were they buying anything from you before? What is the unit of sale, and is it the one you invoice today? And what asset, licence, certification or channel does it need that you do not currently have? The answers place the initiative on the matrix below.

The disruption test

Plot who buys the improvement against what it needs to work. Three of the four quadrants are legitimate places to invest; only one of them is a disruption in the technical sense, and it is the quadrant your reporting pack cannot see.

Capability bets

  • Driver-out corridor, automated yard tractor, robotics retrofit
  • Same customers, new asset or certification — so a priced option, not a programme
  • Timed by the gate, not by model quality

New-market plays

  • Own carrier network plus own channel; drone or pavement delivery at scale
  • New customer and new asset — rarely an incumbent's best use of capital
  • Usually attempted by funded entrants; watch rather than build

Sustaining — build it now

  • Dwell prediction, load consolidation, carrier scoring, document extraction
  • Operating budget, operations owner, holdout, operational KPI
  • Where nearly all realised logistics AI value has come from

Low-end foothold — the blind spot

  • Instant binding quotes on weight breaks you decline
  • Software-only, cheap for an entrant, invisible in your KPIs
  • The only quadrant that is disruption in the technical sense
What it needs to work — top: An asset, certification or channel you do not have, bottom: Systems and assets you already run
Who buys the improvement — left: Existing customers, on today's metric, right: Freight and shippers you decline today
  • Bottom-left is where the money is, and it deserves to be boring

    Sustaining initiatives in this quadrant have a certain-ish return, a checkable arithmetic and an operations owner whose targets improve. The failure mode is not over-investment; it is that they get delayed while the organisation discusses the top-right quadrant. Run them from the operating budget with a holdout and stop bringing them to the strategy forum.

  • Top-left is an option, and its date is set by a regulator or a capital committee

    Capability bets serve your existing customers with an asset or certification you do not have. That makes them fundable, because the demand is already proven, and it makes them un-schedulable, because the gate is outside your control. Price the option, write the gate down, name who watches it, and set a review date.

  • Bottom-right is the technical definition of disruption and the quadrant you cannot see

    It needs no new asset — only a channel and a decision to serve freight you currently decline. That makes it cheap for an entrant and invisible to you, because your KPIs are computed over accepted volume. The decline register is what makes this quadrant appear on a page for the first time.

  • Top-right is usually somebody else's capital

    New customer plus new asset plus new channel is the hardest combination, and incumbents attempting it typically discover that their advantage — density, relationships, network — does not transfer. Watching this quadrant closely is normally a better use of an incumbent's money than entering it, with the specific exception of operators whose existing density genuinely reaches the new job.

One caution about the test, drawn from watching it used badly. It is a classification instrument, not a ranking. Landing in the sustaining quadrant is not a demotion and landing in the low-end quadrant is not a promotion — the low-end quadrant is where the threat is, which is a different thing from where the return is. A portfolio that is ninety per cent bottom-left, with one instrumented signal panel and two priced options, is a healthy logistics portfolio. A portfolio that has been rebalanced toward the top-right because the test made that quadrant sound important has been misused.

Where AI lands in a logistics network, and which side of the split it is on

Warehouse, yard, linehaul, last-mile, pricing, planning and customs — the decisions worth wiring, the system each lives in, the KPI it moves, and whether it sustains or disrupts.

AI value in logistics concentrates in seven operating domains, and six of them are unambiguously sustaining. That is the most useful fact on this page for anyone building a roadmap: the decision map is dominated by improvements to decisions you already make, in systems you already own, measured in KPIs your budget holders already track. The seventh domain — pricing and quoting — is the exception, and it is the exception precisely because it is where the basis of competition has actually moved. It is also the domain most often owned by a commercial team that does not attend the AI roadmap meeting.

DomainDecisions worth wiringSystem of recordKPI it movesSustaining or disruptive
Warehouse & fulfilmentSlotting and re-slotting, wave release, labour planning, vision-based inductionWMS / labour managementPicks per labour hour, dock-to-stockSustaining
Yard & dockTrailer slot assignment, door scheduling, dwell prediction, gate throughputYMS / WMSDock dwell, detention chargesSustaining
Linehaul & networkLoad consolidation, carrier selection, dynamic routing, backhaul matchingTMSCost per shipment, empty milesSustaining
Last-mileRoute sequencing, in-day re-optimisation, promised ETA, attempt predictionDispatch / route plannerStops per hour, first-attempt deliverySustaining
Pricing & quotingInstant binding quotes, accept/decline scoring, dynamic lane pricing, capacity commitmentQuoting engine / rating tables / TMSTime-to-firm-price, win rate by weight breakDisruptive — this is where the basis of competition moved
Planning & S&OPDemand forecast, capacity and workforce planning, seasonal network shapingAdvanced planning system / ERPForecast bias, OTIFSustaining
Compliance & customsDocument classification, commodity coding, dangerous-goods screening, screening evidenceCustoms / broker platformClearance time, audit findingsSustaining, with a disruptive edge in self-serve clearance for small shippers
The logistics decision landscape, with each domain's decisions, system of record, KPI and classification. Pricing and quoting is the only row where a well-built AI capability changes who can serve the job rather than how cheaply you serve it.

The sequencing advice that falls out of this table is unglamorous. Start in the domains where the system of record is yours — warehouse and yard — because the change-approval path is short and the feedback loop is measured in shifts. Move to linehaul and last-mile next, where the absolute savings are larger but the approvals touch carrier contracts and customer promises. Treat pricing and quoting as a separate track with a separate owner, because it is the only row where the work is not an improvement to an existing decision, and giving it to the team that owns dock scheduling will produce a faster version of your current quoting process rather than a new channel.

The compliance row deserves more attention than it usually gets, and for a reason specific to this page. Logistics operates under ISO 28000 (opens in a new tab) security management, Authorised Economic Operator and equivalent customs programmes, good distribution practice on pharmaceutical lanes and emissions accounting regimes — and every one of them asks for the same artefact: a reconstructable trail of what was decided, by what rule, on what data. That trail is generated as a by-product of doing the sustaining work properly. It is also the precondition for any of the speculative rows in the previous section, because unattended commercial commitment is gated by evidence, not by model quality. Building it is present-readiness that happens to be future-readiness.

The instrument panel: seeing disruption in your own data first

Eight signals, all readable from systems you already run, each with the threshold that should trigger a costed decision rather than a discussion.

Disruption shows up in an operator's own transactional data long before it shows up in a market index or a trade headline. The reason is mechanical: a shipper who starts buying from an instant-quote channel stops sending you requests, or sends them and does not wait for your answer, and both of those events are recorded in your quoting engine and your CRM today. The panel below is eight signals built from those records. None of them requires a new system, six of them require joining data you already store, and the two external series are context rather than instrument.

SignalHow to read itSource systemThe kind of threshold that should trigger a decision
Decline rate by weight breakDeclined or unquoted requests ÷ all inbound requests, split by weight break, lane and requested serviceQuoting engine, CRM, inbound email and phone logAny weight break where you decline more than you accept, for two consecutive quarters
Time to firm priceMedian and 90th-percentile elapsed time from request received to binding price returnedQuoting engine timestamps90th percentile above one working day on non-contract freight
Channel mix of inbound bookingsShare arriving via API or self-serve quote versus phone, email and EDI 204 tenderTMS booking-source fieldSelf-serve share crossing an agreed level in any region or customer segment
Win rate by shipper sizeWon tenders ÷ tenders quoted, banded by annual shipper spendCRM and TMS tender recordsA sustained fall in the smallest band while the largest band holds
Who won the lost tenderNamed counterparty on every loss, coded incumbent, digital entrant or private fleetCRM loss reasons, debrief notesDigital entrants appearing in more than an agreed share of losses
Spot-to-contract spread and tender rejectionMarket context for whether a loss was cyclical or structuralExternal market seriesLosses continuing while the spread narrows — a structural, not cyclical, signal
Price dispersion on your own repeat lanesStandard deviation of quoted price per mile on lanes you quote repeatedlyQuoting engine historyDispersion falling steadily toward a single clearing price — commoditisation
Revenue concentration in your strongest segmentShare of revenue from the segment where your assets are most advantagedFinance and TMSConcentration rising while total inbound requests fall — the upmarket retreat
The disruption instrument panel. Thresholds are examples of the form a threshold should take — a number agreed before the first reading, with a named owner and a pre-costed response. Set your own values against your own baseline.

The two external rows are worth treating differently from the six internal ones. Market series such as the spot-to-contract spread and tender rejection — published in DAT's freight market trendlines (opens in a new tab) among others — answer one question well: was this loss the cycle or was it structural? They are poor primary instruments, because they aggregate the behaviour of thousands of shippers and lag it. The six internal rows lead, because they record individual shipper behaviour at the moment it changes. An operator who buys the index and skips the quote log has bought the comfortable half of the panel.

Instrumentation readiness checklist

Seven conditions. If you cannot tick at least five, your disruption discussion is running on anecdote regardless of how good the anecdotes are. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Nothing ticked — start with the decline register, not the governance

A blank list is common and it is not a crisis. Do not begin with the forms: begin with two weeks of work pulling twelve months of declined and lost quote requests out of your quoting engine and CRM, classified by lane, weight break and channel. That single artefact makes the governance conversation concrete, and it is the one item on this list that changes minds by itself.

Three programmes read against the split

Two sustaining wins at enormous scale and one genuine change to the transaction layer. None is an Atomic Loops engagement; each links to the operator's own published material.

The clearest evidence for this page's thesis is in what large operators actually built and what it actually changed. Two of the three below are sustaining innovations of a scale that dwarfs anything in the speculative column — and neither altered who moves the freight or on what asset base. The third genuinely changed the transaction layer of truckload brokerage, and it is instructive precisely because the asset layer underneath it did not move at all. Outcomes are as reported by the operators themselves; verify figures against the linked source before reusing them.

Two sustaining wins and one transaction-layer change

Read the stage columns as rungs on the disruption-response ladder rather than as a verdict on the operator. Images are illustrative industry scenes from our generated library, not operator photographs.

Scene: parcel network dispatch and loading operationsUPSGlobal parcel network · 500k+ employees23
Challenge
Route sequencing was built from static plans and driver experience, with no systematic way to apply network-level optimisation to the daily dispatch across a very large delivery fleet.
Approach
ORION embedded route optimisation directly into the dispatch and driver-facing systems rather than presenting recommendations separately, and was later extended toward continuous in-day re-optimisation rather than a fixed morning plan.
Reported outcome
UPS has publicly reported ORION delivering annual mileage reductions in the region of 100 million miles, with associated cost savings reported in the hundreds of millions of dollars per year.
What it shows about the curveThis is the largest publicly reported AI win in logistics and it is unambiguously sustaining: same customers, same parcels, same trucks, a decision UPS was already making, made better and delivered into the execution path. Anyone using 'sustaining' as a synonym for 'small' should start here.

UPS newsroom (opens in a new tab)

Scene: automated fulfilment centre with mobile robotic drive units and storage podsAmazonRetail and logistics network · purpose-built fulfilment estate34
Challenge
Scaling throughput and storage density across a very large fulfilment estate, where labour supply, building footprint and pick-face variety all bind at once.
Approach
A very large mobile robot fleet inside buildings designed around it, coordinated by AI — Amazon has described a fleet-coordination foundation model, DeepFleet, alongside the robots themselves.
Reported outcome
Amazon reported in 2025 that it had deployed its one millionth robot, describing itself as the world's largest manufacturer and operator of mobile robotics.
What it shows about the curveThe largest robotics deployment in the industry is a sustaining investment inside a network its owner already controls. It changed unit economics inside the existing value network; it did not create a new class of customer. Note also the constraint the case makes visible: the fleets run in buildings designed for them, which is why the retrofit economics are a different question.

Amazon — one millionth robot and DeepFleet (opens in a new tab)

Scene: truckload marketplace operations with digital load booking on a mobile deviceUber FreightDigital freight marketplace and managed transportation15
Challenge
Truckload capacity was matched through brokers by phone and email, with quoting measured in hours and price discovery opaque to smaller shippers and carriers alike.
Approach
An app-and-API marketplace giving carriers upfront pricing and instant booking, and shippers instant quotes and automated tracking — built as a technology channel over third-party capacity rather than as an owned fleet.
Reported outcome
Uber Freight publishes its marketplace, instant-pricing and managed-transportation offering as a productised service for shippers and carriers, and the instant-quote-and-book pattern is now standard across the segment, including at incumbent brokers who built their own.
What it shows about the curveThis is the one genuine disruption of the three, and it changed the transaction layer, not the asset layer. Capacity remained trucks, drivers and hours of service, and margins remained cyclical with the spot-to-contract spread. The incumbents that responded well did so by building the same channel — which is exactly the 'exercise the option' move rung 5 describes.

Uber Freight (opens in a new tab)

A fourth example is worth naming without a full card, because it shows the pattern from the other direction. Digital forwarding — Flexport (opens in a new tab) being the most visible instance — attacked the visibility and documentation layer of international freight rather than the ships and aircraft, and the incumbent response was instructive: Maersk's digital solutions programme (opens in a new tab) reflects an integrator strategy in which the asset owner builds the software layer itself. In both directions the lesson is the same. The asset layer moves slowly and expensively; the transaction, documentation and visibility layers move quickly and cheaply, and that is where an incumbent's option money is usually best spent.

The two-track operating model, layer by layer

What actually has to exist for an operator to run sustaining delivery and disruptive probes side by side without one eating the other.

Running both tracks needs five layers, and the order matters more than the completeness. The architecture below is deliberately unfashionable: nothing in it is a product, every layer is defined by what it must guarantee, and three of the five are governance artefacts rather than systems. The most common failure is to build the top layer — an option register with impressive rows — on top of a portfolio that has never been classified and a market nobody instruments, at which point the register is a list of enthusiasms with dates attached.

Layers required by rung

Each layer is annotated with the rung that first requires it. An operator building the options layer without the instrumentation layer is holding bets it cannot time.

  1. Classification

    Stage 1+

    • The one-sentence testDoes it improve a metric an existing customer already buys on?
    • Labelled portfolio listEvery live and proposed initiative, with the label and the reason
    • Quarterly review forumCommercial, operations and finance in the same room
  2. Funding

    Stage 3+

    • Operating budgetSustaining work, owned by the operations leader whose KPI moves
    • Capped probe budgetSeparate line, own owner, cannot be borrowed against mid-year
    • No-borrowing ruleWritten, because pressure on the operating number always arrives
  3. Measurement

    Stage 2+

    • Operational KPI plus holdoutFor sustaining work — a lane set, door bank or shift left on the old process
    • Learning milestone plus kill ruleFor probes — a question, the evidence, the date and the number that stops it
    • Decline registerThe market's own view of you, refreshed from the quoting system
  4. Instrumentation

    Stage 4+

    • Quote and tender telemetryTime to firm price, win rate by shipper band, loss counterparty coding
    • Channel-mix reportingAPI and self-serve versus phone, email and EDI 204, by region and segment
    • External market contextSpot-to-contract spread and tender rejection — context, never the instrument
  5. Options and evidence

    Stage 5+

    • Option registerTwo to four rows: thesis, holding cost, exercise cost, trigger, date
    • Binding-constraint statementThe asset, certification, capital or physical limit gating each row
    • Decision logExercise, extend or retire — recorded with the evidence that decided it

Pipeline described

  1. Classification (stage 1+) — The one-sentence test: Does it improve a metric an existing customer already buys on?; Labelled portfolio list: Every live and proposed initiative, with the label and the reason; Quarterly review forum: Commercial, operations and finance in the same room
  2. Funding (stage 3+) — Operating budget: Sustaining work, owned by the operations leader whose KPI moves; Capped probe budget: Separate line, own owner, cannot be borrowed against mid-year; No-borrowing rule: Written, because pressure on the operating number always arrives
  3. Measurement (stage 2+) — Operational KPI plus holdout: For sustaining work — a lane set, door bank or shift left on the old process; Learning milestone plus kill rule: For probes — a question, the evidence, the date and the number that stops it; Decline register: The market's own view of you, refreshed from the quoting system
  4. Instrumentation (stage 4+) — Quote and tender telemetry: Time to firm price, win rate by shipper band, loss counterparty coding; Channel-mix reporting: API and self-serve versus phone, email and EDI 204, by region and segment; External market context: Spot-to-contract spread and tender rejection — context, never the instrument
  5. Options and evidence (stage 5+) — Option register: Two to four rows: thesis, holding cost, exercise cost, trigger, date; Binding-constraint statement: The asset, certification, capital or physical limit gating each row; Decision log: Exercise, extend or retire — recorded with the evidence that decided it
Step-by-step insights
Classification — keep the test to one sentence
The test has to be short enough to apply out loud in a forum without anyone reaching for a document, because its whole value is that it is applied before the funding conversation rather than after it. 'Does it improve a metric an existing customer already buys on?' does almost all the work. The moment the test acquires sub-clauses and exceptions it becomes negotiable, and a negotiable classification test reliably reclassifies whichever initiative has the most senior sponsor. Write the reason next to the label so that next quarter's forum can see whether the same reasoning still holds.
Funding — the no-borrowing rule is the whole layer
Separate budget lines are easy to draw and easy to raid. Halfway through any year with a soft operating number, probe money is the least defended pot in the business, and it disappears without a recorded decision because nobody thinks of a reallocation as a cancellation. Writing the no-borrowing rule down — and naming who has to approve an exception — costs nothing and is the difference between two tracks and one track with an aspiration. Cap the probe pot deliberately low; the constraint that matters is attention, and a small pot forces the portfolio to stay at two or three rows.
Measurement — different success shapes, deliberately
Sustaining work is judged on an operational metric against a holdout, because in a live freight network seasonality and carrier-mix change will otherwise claim the credit or take the blame. Probes are judged on whether they answered a stated question by a stated date, and a probe that answers 'no, the market is not there' at a tenth of the cost of finding out later is a success that should be described as one in the forum. Mixing the two shapes is what produces the familiar pathology of a probe reporting a payback figure that everybody in the room knows is decorative.
Instrumentation — six internal signals before any external series
The internal signals lead because they record individual shipper behaviour at the moment it changes; the external series lag because they aggregate. There is also a political asymmetry worth planning for: the external index is comfortable to present and the decline register is not, because the decline register is a list of business your own team turned away with reasons attached. Expect the first reading to be defensive, socialise it before the forum, and frame it as market information rather than as performance review — otherwise the reason codes quietly degrade to 'no capacity' within two quarters.
Options and evidence — the retirement test is written first
Every row of the register should state, at the moment it is created, the evidence that would retire it. Doing that later never happens, because by then the row has advocates. The decision log matters for the same reason the audit trail matters in the sustaining track: it is what lets a forum eighteen months from now reconstruct why a bet was extended rather than relitigate it from memory. Operators who reach this layer usually find the register shrinks at its second review, which is the point at which it starts being useful.

The layer most often skipped is measurement, and skipping it produces a specific, recognisable failure. Without a holdout on the sustaining side, savings claims become arguable and the operating budget for AI shrinks at the next review. Without a kill rule on the probe side, nothing ever ends, and within two years the probe budget is fully committed to work that has quietly become permanent. Both failures look like a funding problem to the people experiencing them, and both are measurement problems that were cheap to prevent.

A 90-day plan: instrumenting the freight you decline

The rung 2 → rung 3 move made concrete on one problem — the small-shipper LTL your quoting desk turns away — with a priced, dated probe at the end of it.

Moving a rung takes about 90 days when it is scoped to one concrete problem, and several years when it is scoped to a culture. The plan below runs the transition on the most common and most diagnostic problem in this domain: the small-shipment LTL freight a carrier or 3PL declines because pricing it by hand takes most of a working day and the margin does not justify the effort. That freight is simultaneously the classic low-end foothold and something you can measure this month, which makes it the right first target. The quarter contains no model development beyond a simple accept/decline scoring step; it is measurement, governance and one bounded channel experiment.

Rung 2 → rung 3 on declined LTL freight, in one quarter

One weight break, one lane set, one named commercial owner. If any phase overruns its window, narrow the scope — fewer lanes, one weight break — rather than extending the plan.

  1. Days 1–15

    Measure the decline

    Pull twelve months of inbound quote requests and load tenders from the quoting engine, the CRM and the inbound mailbox. Classify every request you declined, priced out of, or answered too late, by lane, weight break, requested service and channel. Add a reason code to each. Name the head of pricing or commercial as owner — this is their number, not the data team's.

    A decline register with reason codes, and a named owner

  2. Days 16–35

    Pick the weight break and write the probe

    Choose the single weight break and lane set with the highest declined volume that your network already physically handles. Write the probe on the second business-case form: the question, the evidence that would answer it, the cost cap, the capacity ceiling and the kill rule — an explicit number of accepted instant quotes by an explicit date.

    One probe, capped, dated and killable

  3. Days 36–70

    Run a bounded instant-quote path

    Stand up an instant binding quote for that weight break behind a hard cap on committed capacity, with an accept/decline score and a fallback to manual quoting on anything outside the bounds. Return a price in seconds, not hours. Instrument win rate, margin per shipment, time to firm price and the counterparty on every loss.

    A live channel with a capacity ceiling and full telemetry

  4. Days 71–90

    Decide against the pre-agreed thresholds

    Compare the result with the kill rule as written, not as remembered. Exercise into a funded build, extend with a written reason and a new date, or retire and record why. Whichever way it goes, add the decline register and channel-mix signals to the standing report with thresholds attached, and log the decision.

    A recorded exercise-or-retire decision and two live signals

The order matters

  1. Measure before you classify

    It is tempting to start with the two forms, because they are cheap and feel like progress. In practice a classification forum with no decline register in front of it re-runs the same argument it had last year with the same seniority weighting. Bring twelve months of declined freight to the first meeting and the argument changes shape in ten minutes.

  2. Cap the capacity before you open the channel

    An instant binding quote is a commitment, and the failure mode is winning freight you cannot cover during a tight capacity week. Set the ceiling in loads per day per lane, enforce it in the quoting path rather than in a procedure, and hold the manual fallback one switch away for anything outside bounds.

  3. Write the kill rule before the first quote goes out

    A kill rule written after the results are visible is a negotiation. Written first, it is the thing that lets the sponsor stop the probe without it reading as a personal failure — which is the actual reason most probes never stop. Put the number, the date and the sponsor's name on the form.

What makes this plan worth running even if the probe fails is the residue. At the end of the quarter the operator owns a decline register that refreshes, a channel-mix report, a second business-case form that has been used in anger, and a recorded decision made against a rule. Those four artefacts are the whole of rung 3 and most of rung 4, and they were produced as by-products of one bounded experiment on one weight break. That is the general shape of the argument on this page: the governance you need for the speculative future is built out of work you can justify on this quarter's numbers.

Failure modes that send the portfolio backwards

Four regressions account for almost all of it, and none is caused by picking the wrong technology.

Progress on this ladder is not monotonic, and operators slide back without noticing because the artefacts survive the discipline that produced them. A labelled portfolio list, a signal panel and an option register can all still exist, on the same template, months after anybody stopped acting on them. Four regressions account for most of the losses, and each has a cheap, procedural prevention.

Likelihood: highImpact: high

The probe budget is quietly reallocated mid-year

Pressure on the operating number arrives every year, and the probe pot is the least defended line in the business. It gets reallocated without anyone recording a cancellation, because a reallocation does not feel like a decision. The following year the probe budget is requested again and refused, on the grounds that last year's produced nothing.

PreventionA written no-borrowing rule naming who must approve an exception, and a recorded decision if one is granted.

Likelihood: highImpact: medium

The decline register degrades to a single reason code

Reason codes are entered by people whose performance is visible in them, so within two quarters most declines are coded 'no capacity' regardless of what happened. The register keeps refreshing, the chart keeps rendering, and the information content goes to zero — which is worse than not having it, because the panel now provides false assurance.

PreventionAudit a sample of declines against the original request each quarter, and frame the register as market information rather than performance review.

Likelihood: mediumImpact: high

An option is held past its date and becomes a programme

Nobody decides to convert an option into a programme; it happens by a review being deferred twice. By the third deferral the row has staff, sunk cost and advocates, and it is consuming operating attention while reporting into a forum designed for probes with a fraction of the scrutiny it now needs.

PreventionHold the review on the calendar date even when nothing has changed, and open it with the pre-written retirement evidence rather than the sponsor's update.

Likelihood: mediumImpact: high

A threshold is crossed and nothing happens

The first crossed threshold is the moment the whole instrumentation layer is tested, and the common outcome is a discussion about whether the data is representative. Once a threshold has been crossed without consequence, every subsequent crossing is discounted and the panel reverts to being a reading exercise with numbers on it.

PreventionCost the response at the same meeting that sets the threshold, so crossing it triggers an approved plan rather than a fresh debate.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Sustaining innovation
An improvement to performance on the dimensions existing customers already value and buy on. In logistics: cost per shipment, OTIF, dock dwell, picks per labour hour, empty miles. Sustaining says nothing about size — the largest publicly reported AI wins in freight are sustaining.
Disruptive innovation
In the technical sense, a process that takes root in applications the incumbent finds unattractive — typically cheaper and more accessible — and then moves upmarket. Not a synonym for 'large change'; the defining feature is where it starts, not how impressive it is.
Low-end foothold
The segment an incumbent is willing to lose because serving it is unprofitable at current cost-to-serve. In freight this is small-shipper LTL, one-pallet loads, sub-scale lanes and anything needing a firm price in seconds.
New-market foothold
Demand created by making a job cheaper or simpler for people who previously did not buy it at all — a shipper who never used a forwarder because the process was opaque, rather than one who left a competitor.
Basis of competition
The dimension on which buyers actually choose. In truckload it has partly moved from price to speed-of-price: two carriers quoting the same rate compete on whether the quote takes four seconds or four hours.
Value network
The web of customers, assets, contracts and cost structures inside which an operator's economics make sense. Sustaining innovation improves performance inside it; disruption arrives from outside it, which is why it is hard to see from inside.
Decline register
A record of every inbound request declined, priced out of or answered too late, classified by lane, weight break, service and channel with a reason code. The one artefact that makes the low-end quadrant visible, because every other KPI is computed over accepted volume.
Channel mix
The split of inbound bookings by how they arrive: API or self-serve quote versus phone, email and EDI 204 tender. A shift here is the earliest visible sign that the basis of competition has moved.
Kill rule
A falsifiable stopping condition written before a probe starts — a number, a date and a sponsor. 'We will stop if it is not working' is not a kill rule; a stated acceptance count by a stated date is.
Option register
A short written list of speculative bets, each with a thesis, cost of holding, cost of exercising, trigger signal, binding constraint and review date. Two to four rows is healthy; fifteen rows without dates is a project list.
Binding constraint
The asset, certification, capital requirement or physical limit that decides when a disruptive path becomes possible — automated-driving rules, airspace approval, payload against energy density, building geometry, contractual liability.
Operational design domain
The specific conditions — road type, geography, weather, speed range — an automated driving system is designed and approved to operate in. For freight autonomy the ODD is effectively the product, and its width is the commercial question.

Frequently asked questions

The questions operators ask most often when they start separating sustaining AI from disruptive AI in a freight business.

What is the difference between disruptive and sustaining AI in logistics?

Sustaining AI improves a metric your existing customers already buy on — cost per shipment, OTIF, dock dwell, empty miles — for the same customers on the same assets. Disruptive AI changes who can serve the job, through what channel and on what asset base, and it typically starts in freight you currently decline. The technology does not determine which one you have. The same price-prediction model is sustaining behind your pricing desk and potentially disruptive exposed as a public API to shippers you have never served.

Is most logistics AI sustaining or disruptive?

Overwhelmingly sustaining, and that is where the realised value has been. The largest publicly reported programmes — UPS's ORION route optimisation, Amazon's robot fleet — improved existing networks for existing customers rather than changing who moves the freight. The genuine disruption in freight so far has been at the transaction layer: instant binding quotes and API booking changed how truckload capacity is bought, while capacity itself remained trucks, drivers and hours of service. Treating a sustaining classification as a disappointment is the most expensive mistake operators make here.

Is autonomous trucking disruptive or sustaining?

For an asset-based carrier it is a capability bet, not a disruption: the same customers buy the same linehaul, on an asset base with a different cost curve. What makes it hard is not the classification but the gate. Driver-out operation is bounded by rules for automated driving systems, by a narrow and documented operational design domain, by insurance and by the physical inspection and yard work a tractor still requires. The right posture is to name the corridor, name the operational design domain, name the rule that must change, and monitor it — rather than to argue about whether it will happen.

Was digital freight brokerage actually disruptive?

Yes, in the technical sense, and it is the clearest freight example. It took root in transactions incumbents found unattractive to serve quickly — smaller, spot, price-sensitive loads — by making pricing instant and booking self-serve, then moved upmarket into managed transportation. What it did not do is change the asset layer: capacity remained physical and cyclical, and margins still move with the spot-to-contract spread. The incumbents that responded well built the same channel themselves, which is exactly the exercise-the-option move rather than a defence.

How do I tell whether an AI initiative is disruptive before funding it?

Ask one question first: does it improve a metric an existing customer already buys on? If yes, it is sustaining — fund it from the operating budget with an operations owner and a holdout. If no, ask three more: who is the customer and were they buying from you before, what is the unit of sale and is it the one you invoice today, and what asset, licence or channel does it need that you do not have. Those four answers place it on the disruption-test matrix and determine which business-case form it belongs on.

Should a 3PL fund disruptive probes at all, or just do sustaining work well?

Do the sustaining work well first, because it pays and because it builds the data foundation every speculative path also needs. Then hold a small number of probes — two or three, capped and dated — rather than a portfolio. The argument for holding any is not upside; it is that the low-end quadrant is invisible in your reporting, so without a probe and an instrument panel you would learn about a channel shift from a lost tender debrief. A capped probe budget is cheap insurance against a blind spot you cannot close by reporting harder.

What signals show that a disruptive entrant is taking our freight?

Six internal ones lead: your decline rate by weight break, your time to return a firm price, the channel mix of inbound bookings, win rate banded by shipper size, the coded identity of whoever won each lost tender, and price dispersion on your repeat lanes. Two external series give context — spot-to-contract spread and tender rejection tell you whether a loss was cyclical or structural. The internal signals lead because they record individual shipper behaviour as it changes; published indices aggregate and lag it by quarters.

How much of the AI budget should go to disruptive probes?

Less than most strategy decks imply, and the constraint that matters is not money. Two or three live probes at any time is a healthy number for a mid-sized operator, capped deliberately low so the portfolio cannot sprawl. The scarce input is senior attention in the forum where classification happens — a large probe pot buys more rows than that forum can genuinely review, and unreviewed rows silt into programmes. Size the pot so that being wrong on every row is affordable within one operating year.

What is a kill rule and how do I write one?

A kill rule is a falsifiable stopping condition agreed before a probe starts: a number, a date and a named sponsor. 'We will stop if it is not working' fails the test because it cannot be false. 'We stop if fewer than forty shippers accept an instant quote on this lane set within eight weeks' passes, because the eighth week arrives whether or not anyone wants it to. Write it on the business-case form, and have the sponsor rather than the finance partner set the number — that is what makes stopping survivable.

Does calling an initiative sustaining mean it is low value?

No, and this is the most common misreading of the whole framework. The classification says where the value comes from, not how much there is. UPS's ORION is sustaining and is the largest publicly reported AI win in the industry; Amazon's robotics fleet is sustaining and is the largest deployment of mobile robots in the world. If your test tells you most of your best opportunities are sustaining, the test is working. The correct response is to fund them faster and stop discussing them in the strategy forum.

How do compliance frameworks like ISO 28000 and AEO interact with disruptive AI?

They set the gate on anything unattended. ISO 28000 security management, Authorised Economic Operator and equivalent customs programmes, and good distribution practice on pharmaceutical lanes all require a reconstructable trail of what was decided, by what rule, on what data. That is exactly what blocks agentic systems from making unattended commercial commitments at scale — the constraint is liability and evidence, not model capability. The trail is also produced as a by-product of doing sustaining work properly, which is why present-readiness and future-readiness are largely the same investment.

Do generative AI and agents change the sustaining versus disruptive calculus?

They change the cost of the sustaining track more than they change the disruptive one. Document classification, tender drafting, exception chasing and record pre-population all got cheaper, which pulls more decisions into the economically-worth-automating set — a sustaining effect, and a large one. The disruptive question is unchanged, because it was never about capability: unattended commercial commitment is gated by who is liable and whether the decision can be reconstructed. Apply the same one-sentence test to an agent proposal that you apply to a forecasting model.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for manufacturing, logistics and energy operators — forecasting, routing, quoting, vision inspection and decision support running against live operational data, integrated into the TMS, WMS and quoting layer rather than delivered as dashboards.

  • · Production deployments across freight, warehousing and last-mile
  • · Innovation-portfolio reviews run jointly with commercial and operations leads
  • · Integration-first delivery: TMS/WMS write-back, monitoring, rollback
  • · 13 cited sources on this page

Sources

  1. Clayton Christensen InstituteDisruptive Innovation Theory (opens in a new tab)
  2. GS1EPCIS event standard (opens in a new tab)
  3. ISOISO 28000 — security and resilience for the supply chain (opens in a new tab)
  4. MHIAnnual Industry Report (opens in a new tab)
  5. DAT Freight & AnalyticsFreight market trendlines — spot and contract rates (opens in a new tab)
  6. Uber FreightDigital freight marketplace and managed transportation (opens in a new tab)
  7. FlexportDigital freight forwarding platform (opens in a new tab)
  8. AmazonAmazon deploys its one millionth robot and launches DeepFleet (opens in a new tab)
  9. Federal Motor Carrier Safety AdministrationRegulation of commercial motor vehicles and automated driving systems (opens in a new tab)
  10. GartnerSupply chain artificial intelligence research (opens in a new tab)
  11. MaerskDigital solutions (opens in a new tab)
  12. McKinsey & CompanyOperations insights (opens in a new tab)
  13. UPSNewsroom — ORION route optimisation (opens in a new tab)

Find out which of your AI initiatives is actually disruptive

We run the classification test against your live portfolio, pull twelve months of declined and lost quotes from your quoting engine and TMS, and leave you with a labelled portfolio, a populated signal panel and a costed 90-day plan for your weakest dimension. You keep all three whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.