Redefining Technology

LogisticsReadiness & Transformation Roadmap

Logistics AI readiness for vendors: evaluating, buying and governing the AI you don't build

Logistics AI vendor readiness is an operator's capability to evaluate, integrate, govern and — when necessary — exit the AI it buys rather than builds. It rests on four disciplines: evidence-led evaluation, integration readiness, commercial and exit terms, and in-production vendor governance. In a market where every TMS and WMS vendor now claims AI, that capability decides what a contract actually delivers.

Logistics control room evaluating AI vendor systems across freight, warehouse and network operations
Logistics · Readiness & Transformation Roadmap

Key takeaways

  1. Vendor readiness is a buyer capability, not a market condition. The same logistics AI vendor produces shelfware at one operator and attributable OTIF gains at another; the difference is almost always the buyer's baseline, integration surface and terms — not the vendor's model.
  2. Gartner expects more than 75% of commercial supply-chain application vendors to ship embedded AI and advanced analytics by 2026 — which means 'has AI' no longer discriminates between vendors. The discriminating questions are integration proof, data terms and in-production evidence.
  3. The demo is a controlled experiment run by the seller. The only evaluation that predicts production is a paid proof of value on your own lanes and live feeds, scored against your own baseline with holdout lanes kept back.
  4. Signature is the last moment of leverage. Data ownership, export formats, exit assistance and SLA metrics negotiated before signing cost almost nothing; the same terms requested at renewal are priced against you.
  5. Vendor continuity is a readiness dimension, not a due-diligence footnote: roughly 18 months separated Convoy's final raise at a $3.8B valuation from its 2023 shutdown. Operators with a drilled exit experienced a migration; the rest experienced an incident.

Abbreviations used on this page

TMS
Transport management system
WMS
Warehouse management system
YMS
Yard management system
EDI
Electronic data interchange (e.g. the 214 shipment status message)
API
Application programming interface
ETA
Estimated time of arrival
OTIF
On-time in-full delivery rate
3PL
Third-party logistics provider
RFP
Request for proposal
PoV
Proof of value — a paid, bounded trial on the buyer's own data
SLA
Service-level agreement
QBR
Quarterly business review — the vendor-run account meeting

Free · 8 questions · ~3 minutes

Score how ready you are to buy AI

Eight questions, one at a time, about three minutes — on how your operation evaluates, integrates, contracts and governs its AI vendors. Answer them and we build your vendor readiness report: your stage on the buying ladder, your score on each of the four disciplines, and the specific gap most likely to cost you at your next selection or renewal. It arrives by email.

0 of 8 answered

Question 1 of 8Evaluation discipline

What evidence decided your most recent logistics AI vendor selection?

The evidence type is the single strongest predictor of buying maturity — demos and references are seller-controlled; trials on your own lanes are not.

How the score maps to a stage
  • 05 — Stage 1, Demo-led. Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
  • 611 — Stage 2, Checklist-led. Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
  • 1216 — Stage 3, Evidence-led. Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
  • 1721 — Stage 4, Production-governed. Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
  • 2224 — Stage 5, Portfolio-orchestrated. The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.

What logistics AI vendor readiness is — and why the demo is the least of it

A definition, the buying curve, and the two paths a vendor claim can travel from demonstration to your production stack.

Logistics AI vendor readiness is the organisational capability to buy AI well: to evaluate vendor claims against your own evidence, to feed and receive vendor systems through integration surfaces you control, to fix data and exit terms while you still have leverage, and to govern vendor performance from your own telemetry for the life of the contract. It is a property of the buyer, not the market — which is why the same vendor produces shelfware at one operator and attributable OTIF gains at another.

The capability matters more every year for a structural reason: AI is ceasing to be a differentiator among logistics vendors and becoming table stakes. Gartner's supply-chain AI research (opens in a new tab) projects that by 2026 more than 75% of commercial supply-chain management application vendors will deliver embedded advanced analytics and AI. When every TMS, WMS and visibility product 'has AI', the phrase stops discriminating between vendors — and the discriminating work moves to the buyer: integration proof, data terms, and evidence from your own lanes. This page is about building that muscle. What it is not about: sensor estates (the IoT readiness page), pipeline and master-data quality (the data readiness page), or which warehouse or freight initiative to sequence first (their own roadmap pages). Those disciplines feed this one; the buying capability is its own.

Value released against buying maturity

The curve is not linear. Through the demo-led and checklist-led stages, contracts accumulate faster than production value — spend rises while decisions stay unchanged. Value inflects when evaluation moves onto the operator's own data and terms are set at signature, and compounds when the stack becomes governable and finally composable.

Production value released per contracted pound by stage

  • Stage 1 · Demo-led — 24% of operators. Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
  • Stage 2 · Checklist-led — 37% of operators. Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
  • Stage 3 · Evidence-led — 26% of operators. Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
  • Stage 4 · Production-governed — 10% of operators. Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
  • Stage 5 · Portfolio-orchestrated — 3% of operators. The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with Gartner's supply-chain AI adoption research.

Two paths from vendor claim to production — and where each leaks

The same vendor claim can travel two paths through a logistics operation. The demo-led path terminates in shelfware renewed by sunk cost; the evidence-led path converts the claim into a governed production capability — and the governance lane is what keeps it converted. Most logistics vendor selection still travels the top lane.

  • AI / model
  • System-of-record action
  • Where value leaks
  • Data & feeds
  • Human in the loop

The process, in words

  • On the demo-led path, evaluation runs on the vendor's curated sample data and scripted pilot, terms are signed with integration unpriced, and the product meets the operator's EDI reality only after signature. The stall that follows is absorbed as sunk cost, and the contract renews by inertia — shelfware with a maintenance fee.
  • On the evidence-led path, the operator baselines its own lanes first, runs a paid PoV for at least two vendors on the same live feeds with holdout lanes kept back, tests integration with one real feed in both directions before signing, and fixes data ownership, exit assistance and SLA metrics while competition is still alive.
  • The production-governance lane is what keeps the evidence honest after go-live: a vendor register with a named owner per vendor, performance computed from the operator's own telemetry rather than the vendor's QBR, renewal gates decided on that evidence, and a drilled exit that keeps the walk-away real.
Step-by-step insights
The demo is a controlled experiment run by the seller
Nothing about a demo is dishonest, and nothing about it is evidence. The data was chosen to flatter the model, the scenario was rehearsed, and the edge cases that dominate a real logistics network — carriers that transmit three of nine EDI 214 event types, item masters with duplicate SKUs, the peak-week data your planners actually live with — are precisely what the demo excludes. Treat demos as theatre with a briefing function: useful for understanding what a product intends to do, structurally incapable of predicting what it will do on your freight.
The baseline is the buyer's only native advantage
The vendor knows its product; the operator knows its network. A computed baseline — your current ETA error by lane, your slotting travel times, your tender acceptance by carrier, from six months of TMS/WMS history — converts that knowledge into the one artefact no vendor can supply or dispute. It is the difference between asking 'is this product good?' (unanswerable) and 'does this beat what we already achieve, on our lanes, by enough to fund the integration?' (a number). Operators who skip the baseline are outsourcing the definition of success to the party being paid to succeed.
Holdout lanes keep everyone honest — including you
A holdout is a set of lanes, doors or facilities the vendor never sees until scoring day. It defends against the two corruptions every trial invites: the vendor tuning to the test set, and the buyer's own champions unconsciously selecting the lanes where the product shines. Choose holdouts to represent the network's hard cases — the carrier with sparse EDI, the seasonal lane, the facility with the messy master data — because production is made of hard cases, and a product that only works on clean lanes will meet the rest of your network eventually.
The integration spike — one real feed beats any architecture slide
Before signature, run one production feed through the vendor's ingestion and one write-back into a test instance of your TMS or WMS — a two-week spike, both directions. This single exercise surfaces what no RFP answer can: the event types the product actually consumes, the master-data assumptions it makes, the latency it adds, the effort your side genuinely requires. Vendors confident in their integration story agree readily; hesitation is itself diligence data. The spike converts the largest unknown cost in the contract into a tested number while you can still walk away.
Terms at signature — the only moment of leverage
Data ownership, export formats, exit assistance, model-improvement rights and the SLA metric all cost approximately nothing to obtain while two vendors are still competing, and are priced as change requests the day after signature. The asymmetry is structural: before signing, the vendor is selling; after, you are retained. The evidence-led path's commercial payoff is exactly here — the trial's real product is not just a selection but a signed contract whose terms were set under competition. A stage-3 trial that ends in a stage-2 contract wasted its leverage.
Own telemetry, the renewal gate and the drilled exit
After go-live, the trial harness keeps running: acceptance rates from your approval logs, ETA error from your TMS events, OTIF from your customer EDI — your numbers, on a file the renewal gate opens. The gate has three honest outcomes: extend on evidence, renegotiate on evidence, or re-compete. But the third outcome exists only if the exit is real — fallback documented, export re-loaded at least once, elapsed time known. That is why the drilled exit sits at the end of the governance lane: it is the difference between a gate and a ceremony.

The five stages of the buying ladder in detail

From demo-led to portfolio-orchestrated: what each stage looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.

Each stage below describes how an operator buys, not what it owns — an operator with a dozen AI contracts can sit at stage 1, and a lean operation with two well-governed vendors can sit at stage 4. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own contracts and telemetry this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Demo-led

24% of operators sit here

Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.

Stage 1 is not gullibility — it is asymmetry. The vendor has run its demonstration hundreds of times against data curated to flatter the product; the operator is seeing it once, without a baseline of its own to compare against. Under that asymmetry, selection quality is capped no matter how experienced the buying team is, because every input to the decision was produced by the seller.

The tell is what happens after signature. Integration effort was a line in the vendor's proposal rather than a tested fact, so the first months of the contract are spent discovering the real shape of the work: the WMS export the product assumed doesn't exist, the EDI 214 feed carries half the events the model needs, the master data has three codings for the same carrier. None of this was hidden — it was simply never tested, because the evaluation never touched the operator's own systems.

The stage is self-perpetuating in a specific way: each disappointing purchase is blamed on the vendor, and the next vendor is selected by exactly the same process. Operators can cycle through three routing or slotting products in five years without once changing the method that chose them — which is the actual defect. The exit from stage 1 is not a better vendor; it is the first internal baseline that makes vendor claims testable.

In practice

The slotting engine bought at a trade show

A regional 3PL signed a warehouse slotting product after a strong conference demo and two reference calls. The demo ran on a model warehouse with clean item master data. The 3PL's own item master, built through two acquisitions, had duplicate SKUs and inconsistent dimensions — the product's first three months were a data-cleansing project nobody had scoped, and the operations team that was promised travel-time savings got a re-keying exercise. The product was quietly shelved before peak; the post-mortem blamed the vendor.

What it looks like

  • Shortlists are assembled from analyst charts, trade shows and demos
  • Pilots run on vendor sample data, in the vendor's own environment
  • Integration effort is estimated by the vendor, after selection
  • No internal baseline exists to test any vendor claim against

Diagnostic signals you can check this week

  • Ask what data the last vendor demo ran on. If the answer is 'theirs', you are here
  • Ask for the internal baseline a current vendor is measured against. Silence is the answer
  • Check whether integration cost was a tested number or the vendor's estimate at signature
  • Count AI products purchased in three years against those in daily production use

Anti-pattern · Fixing it with a bigger shortlist

The instinctive correction is to see more vendors — six demos instead of two, a scoring sheet, a second round. This multiplies the inputs without changing their nature: every one is still a seller-controlled demonstration. Six curated demos aggregate to no more evidence than one. The move that changes the outcome is not widening the funnel but changing what flows through it — one measured baseline of your own, and a requirement that the next vendor beat it on your data.

What holds you here

There is no internal baseline, so every vendor claim is compared against another vendor claim rather than against the operation's own numbers.

Highest-leverage next move

Baseline one decision — ETA accuracy, slotting travel time, tender acceptance — from your own TMS/WMS history, and require the next candidate vendor to beat it on your data.

Cost of leaving

Effort
2–4 months
Team
One operations analyst and one engineer, part-time, to baseline a single decision
Risk
Low — baselining your own lanes commits you to nothing and informs everything
To next stage
2–4 months

If this is you, the next step is

A short engagement: pick the decision, compute the baseline, write the test a vendor must pass.

Baseline one decision before the next demo

Stage 2

Checklist-led

37% of operators sit here

Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.

Stage 2 looks like rigour and often defeats its own purpose. The RFP process is real: requirements are gathered, criteria are weighted, demos are scored by a panel, references are called. But every artefact in the process is still a claim rather than a behaviour. A 400-line requirements matrix scores what the vendor asserts about the product; it cannot score what the product does when it meets the operator's EDI reality, because nothing in the process ever connects the two.

The structural bias of checklist buying is toward feature count, and feature count favours suites and punishes depth. A vendor with forty features that each work at demo quality outscores a vendor with four features that survive production — the matrix cannot see the difference, and the reference calls can't either, because references are curated by the same seller. Meanwhile the two chapters that decide the contract's fate — integration effort and data terms — sit at the back of the RFP as compliance questions ('Do you provide APIs? Yes.') that every vendor passes.

The costly part is timing. At stage 2 the operator discovers integration reality and negotiates data terms after selection, when leverage is gone: the internal announcement has been made, the project has a name, and the vendor knows it. Terms that would have cost nothing to obtain in a competitive process are now priced as change requests. Most of what operators later describe as lock-in was conceded here, in the weeks after the winner was chosen.

In practice

The 400-line RFP that missed the EDI question

A mid-size freight operator ran a disciplined RFP for a predictive-ETA platform: 412 weighted requirements, four vendors, scored demos. The winning vendor scored highest on functionality and analytics. Nobody had asked which EDI 214 status events the models actually consumed, or tested the operator's own feed against them. Post-signature, it emerged the operator's carriers transmitted a fraction of the event types the product was trained on; ETA quality on the operator's real lanes was far below the demo, and the first renewal became a dispute about whose fault that was. The RFP had scored 412 claims and tested none.

What it looks like

  • A formal RFP with weighted criteria governs every AI purchase
  • Demos are scored — but still scripted by the vendor
  • Feature coverage decides selection; integration is a late chapter
  • Data ownership and exit terms surface after the winner is chosen

Diagnostic signals you can check this week

  • Open your last AI RFP: count requirements that were tested versus asserted
  • Check when data-ownership terms were first discussed — before or after selection
  • Ask whether any scored demo ran on your own lanes, doors or item master
  • Look at the criteria weights: if feature coverage outweighs integration proof, you are here

Anti-pattern · Adding more rows to the matrix

When a checklist-led purchase disappoints, the reflex is a longer checklist — the 412 requirements become 600, with a new section on AI capabilities that asks vendors to describe their models. This deepens the defect instead of fixing it: more claims, still no behaviour. The correction is to delete most of the matrix and replace its centre of gravity with one requirement that cannot be answered in prose: run against six months of our history and beat our baseline, on lanes we choose, with holdout lanes you never see until scoring.

What holds you here

The evaluation tests vendor claims, not vendor systems — so selection quality is capped by the honesty of the sales process, and terms are negotiated after leverage is gone.

Highest-leverage next move

Replace the demo-scoring round of your next selection with a paid proof of value on your own lanes and live feeds, scored against your baseline — and put data and exit terms on the table while two vendors are still competing.

Cost of leaving

Effort
3–6 months
Team
Procurement plus one operations owner and one integration engineer with authority over the evaluation design
Risk
Low-to-medium — the change is procedural; the resistance is habit, not cost
To next stage
3–6 months

If this is you, the next step is

We help you design the trial, the baseline and the term sheet before you talk to vendors.

Replace one RFP round with a PoV

Stage 3

Evidence-led

26% of operators sit here

Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.

Stage 3 inverts the asymmetry of stages 1 and 2: the operator now controls the experiment. Candidate vendors run against the same six months of TMS and WMS history, on lanes the operator chose, scored against the operator's own baseline in operational units — ETA error in minutes, OTIF points, dwell minutes — with a set of holdout lanes no vendor sees until scoring day. The demo still happens, but it has been demoted from evidence to theatre, which is what it always was.

Two disciplines make the stage real rather than nominal. First, the trial is paid: a fair fee for a bounded PoV filters for vendors confident enough to be measured, and removes the pretext that a free pilot's failure was under-investment. Free pilots are sales activity, and both sides price them accordingly. Second, at least two vendors run to the end. The moment a single vendor knows it is the only candidate, the operator's negotiating position on data terms, exit assistance and SLA metrics collapses — most of stage 3's commercial value comes from keeping the competition alive until signature.

What stage 3 has not yet solved is the portfolio. Each selection is disciplined, but selections accumulate: three years of good individual decisions produce a visibility platform, two point solutions and a suite module with overlapping capabilities, each renewed on its own cycle by whoever signed it. Nothing in the operation can say what the vendor estate costs per decision, which contracts overlap, or which renewal dates are approaching. The selections are governed; the stack is not.

In practice

The two-vendor ETA bake-off

A retailer's transport team ran two visibility vendors in parallel for eight weeks on the same 40 inbound lanes, live EDI 214 and telematics feeds, with 10 holdout lanes scored blind at the end. Vendor A — the analyst-chart favourite — was two minutes better on mean ETA error; vendor B degraded far less during a carrier-mix shift mid-trial and misclassified fewer late arrivals as on-time. The team chose B, and the trial data did double duty: with a competitor still live, B agreed to explicit data-export terms, a defined exit-assistance package and an SLA metric matching the trial's scoring — terms A's standard contract did not offer.

What it looks like

  • Paid PoV trials on own lanes and live feeds decide selection
  • Holdout lanes are kept back from vendors until scoring
  • An integration spike — one real feed, both directions — is part of diligence
  • Data ownership, export formats and exit assistance are signed at signature

Diagnostic signals you can check this week

  • Ask how the last AI vendor was selected: if the answer names a trial on your data, you are at least here
  • Check whether holdout lanes existed and who chose them
  • Ask whether the PoV was paid, and how many vendors reached the final week
  • Read the signed contract for export formats and exit assistance — present means stage 3, absent means the trial discipline stopped at selection

Anti-pattern · Trialling one vendor

The most common corruption of stage 3 is running a genuine PoV — own data, real baseline, honest scoring — with a single vendor, usually because the trial is framed as a technical validation after a commercial preference has already formed. The evidence is real but the leverage is gone: a vendor that knows it has won concedes nothing at signature, and the operator ends up with stage-3 evidence and stage-2 terms. The trial's cost barely changes with a second vendor; its negotiating value changes completely.

What holds you here

Selections are disciplined one at a time, but the vendor estate as a whole is unmanaged — overlap accumulates, renewals pass by inertia, and nobody can price the stack per decision.

Highest-leverage next move

Stand up a vendor register — every AI vendor, its decision, its owner, its renewal date, its contracted metric — and start computing each vendor's performance from your own telemetry rather than the vendor's QBR.

Cost of leaving

Effort
6–12 months to make evidence-led selection the default across the operation
Team
A repeatable trio: operations owner, integration engineer, commercial lead — plus the discipline to fund paid trials
Risk
Medium — paid two-vendor trials cost real money per selection; the return arrives at signature and at renewal
To next stage
6–12 months

If this is you, the next step is

Trial design, baseline, holdout selection and the term sheet — set up before vendor conversations start.

Design your next PoV

Stage 4

Production-governed

10% of operators sit here

Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.

Stage 4 extends the trial discipline into the life of the contract. The scoring harness built for the PoV — baseline, operational metric, holdout re-checks — keeps running after go-live, so the operator always knows what each vendor's model is doing on its own lanes this quarter, not what the vendor's QBR deck says it is doing. This sounds like bookkeeping and behaves like leverage: renewal conversations change character entirely when the buyer opens with the vendor's own acceptance-rate trend.

The register is the artefact that makes the stage visible. One list: every AI vendor, the decision it serves, the named owner accountable for its outcome, the contracted metric, the renewal date, the annual cost, the last telemetry reading. Most operators assembling it for the first time find things nobody chose: two products computing ETAs for different teams, a suite module licensed but unused since a champion left, a point solution whose owner departed eighteen months ago while the subscription auto-renewed twice. The register does not fix any of this; it makes it undeniable.

What stage 4 cannot yet do is recompose. Each vendor is well-governed in place, but the integrations are point-to-point — the visibility platform wired directly into the TMS, the slotting engine directly into the WMS — so replacing any of them is a project measured in quarters. The operator can measure, negotiate and even decide to switch; it cannot yet switch cheaply. Renewal gates have teeth only when walking away is priced, and at stage 4 it is still priced too high.

In practice

The renewal that got cheaper

A 3PL's vendor register flagged a visibility platform renewal 120 days out. The telemetry file showed exception-alert acceptance by planners had slid from 71% to 44% over a year — alert fatigue from a threshold change the vendor had shipped without notice. Instead of the standard uplift renewal, the 3PL brought the acceptance curve to the QBR, re-ran two holdout lanes as a spot-check, and renewed at a reduced rate with a contractual alert-precision SLA and quarterly threshold reviews. The register turned a rubber-stamp into a negotiation; the telemetry turned the negotiation into a win.

What it looks like

  • A vendor register lists every AI vendor, owner, metric and renewal date
  • Vendor performance is computed from the operator's own telemetry
  • Renewals pass through a gate: extend, renegotiate or re-compete on evidence
  • The QBR reviews the operator's numbers, not the vendor's deck

Diagnostic signals you can check this week

  • Ask for the vendor register. Existence, ownership and freshness are the test
  • Pick one vendor and ask for its current metric from your telemetry, not its QBR
  • Check the last three renewals: how many passed through an evidence gate versus auto-renewing
  • Ask who owns each vendor outcome by name — and whether any named owner has left

Anti-pattern · Governing with the vendor's dashboard

The comfortable version of stage 4 monitors every vendor through the analytics screens the vendor itself provides. It feels like governance and measures the vendor's homework with the vendor's ruler: uptime and prediction volume where decision quality should be, accuracy definitions that quietly exclude the hard cases, baselines that moved when the model did. If the number that decides a renewal is rendered by the party being renewed, it is not governance. The telemetry must come from your own TMS, WMS and approval logs, or the gate is decorative.

What holds you here

Integrations are point-to-point, so exits are unpriced and re-competition is theoretical — the gate can renegotiate but cannot credibly walk away.

Highest-leverage next move

Move vendor integrations onto a layer you own — canonical events, standard identifiers, documented schemas — and rehearse one exit end-to-end, so that at the next gate, switching has a known price.

Cost of leaving

Effort
9–18 months
Team
A small vendor-governance function — often one commercial owner plus the platform engineer who owns telemetry — and named business owners per vendor
Risk
Medium — the telemetry work is modest; the political work of ending auto-renewals is not
To next stage
12–24 months

If this is you, the next step is

We build the register and the telemetry file for your current estate, and run the first renewal gate with you.

Stand up your vendor register

Stage 5

Portfolio-orchestrated

3% of operators sit here

The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.

Stage 5 is not vendor independence — it is vendor optionality. The operator still buys most of its AI, and should: no logistics operator will out-build a visibility network's data advantage or a suite vendor's integration surface across its own product. What changes is the architecture of the relationship. Every vendor consumes and returns data through an integration layer the operator owns — canonical shipment and warehouse events, GS1-style identifiers, documented schemas — so a vendor is a replaceable module in a stack the operator composes, rather than a load-bearing wall the stack was built around.

The discipline that keeps the stage honest is the rehearsed exit. Once a year, one vendor-served decision is deliberately failed over to its documented fallback — the previous rule set, an alternative product, a manual procedure — on a quiet lane set, and the elapsed time and pain are recorded. The first rehearsal is always humbling: the export that had never been re-loaded, the fallback rule set nobody had updated since go-live. But an operator that has rehearsed an exit negotiates differently, buys differently and — as Convoy's customers discovered — absorbs a vendor shutdown as a migration rather than an incident.

Build-versus-buy becomes a live, per-domain decision rather than an identity. Commodity decision domains — visibility, freight audit, standard forecasting — stay bought and are re-competed at renewal. Domains where the operator's data or network shape is genuinely distinctive become candidates to build, because the integration layer that made vendors replaceable is the same substrate an internal model plugs into. Stage 5 operators move specific decisions in-house not on principle but on arithmetic, and sometimes move them back.

In practice

The 30-day migration that wasn't an incident

When a digital freight network in a shipper's carrier mix wound down operations with weeks of notice — as Convoy's market learned can happen in 2023 — one high-volume shipper rerouted its affected tender flow to its documented backup: contract carriers plus a second marketplace, pre-integrated through its own tendering layer and last rehearsed eight months earlier. Spot exposure rose for a month and settled. Its peer, integrated point-to-point with the same network, ran a war room for six weeks re-keying tenders. Same market event; the difference was entirely on the buyer's side.

What it looks like

  • All vendor traffic flows through an owned integration layer with standard events
  • At least one vendor exit has been executed or rehearsed in the last year
  • Build-versus-buy is decided per decision domain, and revisited
  • Renewal gates carry real alternatives — the walk-away is priced and tested

Diagnostic signals you can check this week

  • Ask when a vendor exit was last rehearsed, and for the written result
  • Trace one vendor's data path: through an owned layer, or wired directly into the TMS/WMS?
  • Ask which decision domains are deliberately built versus bought, and who decided
  • Check whether any renewal in the last two years was actually re-competed — not threatened, done

Anti-pattern · Confusing multi-vendor with resilient

Operators sometimes claim stage 5 because they run many vendors — five AI products must surely be a portfolio. Five vendors integrated point-to-point are not a portfolio; they are five hostages, each with its own bespoke integration debt and unpriced exit. The count of vendors is irrelevant. The stage is defined by the layer underneath them: one owned, standard, documented integration surface that makes any single vendor — including the biggest — replaceable at a known cost. Resilience lives in the substrate, not the roster.

What holds you here

Sustaining leverage as the vendor market consolidates — every acquisition of a point solution by a suite, and every network's data advantage, pulls the stack back toward entanglement.

Highest-leverage next move

Keep the annual exit rehearsal and the per-domain build-versus-buy review on the calendar — the stage regresses quietly the year both are skipped.

Cost of leaving

Effort
Continuous
Team
Platform engineering owning the integration layer, a commercial owner running gates and rehearsals, and executive air-cover for re-competition
Risk
Concentrated at market events — consolidation, vendor acquisition and sunset — which is precisely when the stage pays

If this is you, the next step is

We run one exit rehearsal with you and price the walk-away for your next renewal gate.

Pressure-test your portfolio

The logistics AI vendor landscape: five archetypes

Suite modules, visibility networks, point solutions, freight marketplaces and platform tooling — what each actually sells, the lock-in mechanism each carries, and where each fits.

The logistics AI vendor market sorts into five archetypes, and each carries a different lock-in mechanism — which means each demands a different diligence emphasis. Readiness is not choosing the right archetype; every operator of scale ends up buying from most of them. It is knowing which questions each archetype must answer before signature, because a suite module, a visibility network and a freight marketplace fail buyers in entirely different ways.

ArchetypeWhat they actually sellSystems they touchLock-in mechanismFits when
Incumbent suite AI — TMS/WMS/planning vendors adding modulesAI inside the system you already run: forecasting, optimisation and exception features priced into the suiteTheir own TMS / WMS / planning estateThey already own your workflow and your data gravity; AI is bundled into the suite renewalDecisions living wholly inside that suite, where integration cost dominates model quality
Visibility & ETA networksNetwork-effect data products: predictive ETA, exception alerts, dwell and disruption signals trained across many shippers' flowsTMS, carrier EDI, telematicsThe network's data accumulates on their side — your history improves a product you rentMulti-carrier ETA and exception decisions no single operator could train alone
AI-native point solutionsOne decision done deeply: slotting, dock scheduling, demand forecasting, freight audit, carrier vettingWMS / YMS / TMS via APIDeep workflow embed and bespoke integrations that make replacement a projectA decision the suites do shallowly, worth a dedicated product and its integration bill
Digital freight networks & marketplacesMatching, pricing and capacity as a service — the AI is their operation, not your toolTMS tendering and settlementLiquidity and rate history live with the marketplace; leaving means rebuilding relationshipsSpot and backup capacity — benchmarked against independent rate data, never evaluated on quoted savings alone
Hyperscaler & platform toolingModels, pipelines and infrastructure — build-adjacent capability, not logistics productsYour data platformCloud gravity: egress, proprietary services, accumulated pipeline codeOperators with engineering teams that own the decision layer and want vendors only underneath it
The five archetypes of logistics AI vendor. The lock-in column is not an accusation — every mechanism listed is a legitimate business model. It is the thing your terms and architecture must answer before it answers you.

Two archetypes deserve a specific caution each. Marketplace savings claims should never be evaluated on the marketplace's own numbers — independent rate benchmarks such as DAT's freight market data (opens in a new tab) exist precisely so a quoted saving can be tested against what the market actually paid on those lanes in that week. And the visibility networks' honest pitch — that products like Uber Freight's (opens in a new tab) matching or a visibility platform's ETA are trained across flows no single shipper could assemble — is also the lock-in to underwrite: your history is improving an asset you rent. The network effect is real value; the question your contract must answer is what happens to your data, and your operation, when you leave.

A third, quieter shift: your existing suppliers are becoming AI vendors whether you run a selection or not. Carriers ship digital products alongside capacity — Maersk's digital solutions portfolio (opens in a new tab) is a carrier selling software to its own shippers — and every suite upgrade now lands with AI features enabled. Vendor readiness therefore is not only for procurement events; it is the standing discipline of asking, of every product already inside the estate, the same questions a new vendor would face.

Where logistics operators sit on the buying ladder

The distribution across the five stages, and why checklist-led buying is the plateau.

Most logistics operators buy AI at stage 2: the procurement process is formal, but the evidence inside it still belongs to the seller. The distribution below is illustrative — synthesised from published adoption research rather than measured from a single survey — but its shape matches what the industry's own reporting keeps finding: adoption intent and contract volume far ahead of production evidence and governance.

Distribution of logistics operators across the five buying stages

Stage 2 is the mode: formal procurement, seller-owned evidence. The sharpest capability jump on the ladder — and the biggest commercial payoff — is the move to stage 3, when trials shift onto the operator's own lanes and terms move to signature.

Share of operators (illustrative)

  • 24% — 1 · Demo-led
  • 37% — 2 · Checklist-led (the plateau)
  • 26% — 3 · Evidence-led
  • 10% — 4 · Production-governed
  • 3% — 5 · Portfolio-orchestrated

Source: Illustrative distribution, synthesised from MHI, Gartner and McKinsey adoption research

The external evidence for the plateau is consistent. MHI's Annual Industry Report (opens in a new tab) has tracked, year over year, a wide gap between the share of supply-chain organisations planning or piloting AI and the share running it in production — a gap that is precisely the stage-2 signature, since demo-led and checklist-led purchases produce contracts without producing production decisions. McKinsey's operations research (opens in a new tab) reports early AI adopters in supply chains achieving roughly 15% lower logistics costs than slower peers — a prize that accrues to operators whose purchases reach production, which is a buying-capability outcome before it is a modelling one.

One implication is worth making explicit: the move from stage 2 to stage 3 is the highest-return step on the ladder, and it is procedural rather than technical. It requires no new platform, no data-science hires and no transformation programme — it requires the next selection to run as a paid two-vendor trial on your own lanes with the term sheet on the table. Operators repeatedly overestimate the cost of that step and underestimate what stage-2 buying is already costing them in shelfware and conceded terms.

The production-readiness scorecard: what to test before you sign

Six diligence areas, the question each must answer, the evidence to demand, and the walk-away signal — plus the two-axis map that tells you how hard to apply them.

Vendor due diligence in logistics AI comes down to six areas, and the discipline is demanding evidence rather than answers in every one of them. An RFP response is an answer; a spike, a trial score, a named reference running your carrier mix, or a contract clause is evidence. The scorecard below is the working checklist we use when sitting on the buyer's side of a selection — apply it in proportion to the stakes, using the matrix that follows.

Diligence areaThe question that mattersEvidence to demandWalk-away signal
Integration proofWhat does it cost, on both sides, to connect this to our TMS/WMS/YMS estate?A two-week integration spike: one real feed in, one write-back out, effort loggedRefusal to spike; integration priced only after signature
Model evidenceDoes it beat our baseline on our lanes — including the hard ones?Paid PoV on your live feeds, scored on holdout lanes in operational unitsEvaluation offered only on vendor sample data or 'reference customer results'
Data & improvement termsWho owns our operational data, its exports, and the model value trained on it?Contract language: ownership, export format and cadence, exit assistance, improvement rights'Standard terms' that are silent on data — silence always resolves for the vendor
Operational supportWhat happens at 02:00 during peak when the model misbehaves?SLA with response times, a named escalation path, and a tested way to disable the modelSupport scoped to business hours in another timezone; no kill-switch story
Viability & continuityWill this vendor exist, independent and motivated, for the life of the integration?Funding condition, customer concentration, escrow or continuity terms proportionate to their sizeRunway questions deflected; continuity 'never been asked before'
Security & assuranceCan they evidence their claims about controls, and the provenance of their data?Third-party attestation (SOC 2 or equivalent; ISO/IEC 42001 for AI management is emerging), and named data sourcesCertifications 'in progress' for years; training-data provenance they cannot state
The six-area production-readiness scorecard for logistics AI vendors. The walk-away column marks the signals that end a conversation regardless of how the demo went.

Two of these areas have logistics-specific depth worth naming. On security and assurance, ask where the vendor's data actually comes from: carrier-vetting and compliance products, for instance, are largely built on public registries such as FMCSA's safety and registration data (opens in a new tab) — knowing which parts of a product are proprietary intelligence and which are repackaged public data changes what the subscription is worth. On viability, remember that this market's consolidation is not hypothetical: point solutions get acquired by suites, marketplaces exit, and the diligence question is not 'might this happen?' but 'what does our contract and architecture do when it does?'

How hard to apply the scorecard: criticality against switching cost

Plot each vendor — current or candidate — by how core its decision is to your network and how entangled its exit would be. The quadrant sets the governance weight it deserves; the emphasised quadrant is where renewals go wrong.

Strong position

  • Core decision, portable vendor
  • Keep your baseline alive and re-compete at renewal
  • This is where the integration layer earns its keep

The lock-in trap

  • Core decision, entangled exit
  • Full scorecard, continuity terms, annual exit rehearsal
  • Engineer the exit before the renewal, not during it

Commodity

  • Peripheral and portable
  • Buy on price, light-touch governance
  • Do not spend trial budget here

Quiet debt

  • Peripheral but entangled
  • Contain: no scope growth without portability terms
  • The overlooked quadrant where estates silently calcify
Decision criticality — top: Core network decision, bottom: Peripheral decision
Switching cost — left: Portable — standard data and interfaces, right: Entangled — proprietary data and workflow

What the vendor market looks like in public

Three public reference points read against the buying ladder — the build extreme, the continuity shock, and the supplier-turned-vendor. None is an Atomic Loops engagement; each links to the operator's own published material.

The public record teaches the vendor-readiness lesson from three different directions. Amazon marks the far end of build-versus-buy and shows what full integration ownership costs and returns; the Convoy shutdown — and Flexport's acquisition of its technology — is the clearest continuity case study the industry has; and Maersk shows incumbent suppliers becoming AI vendors to their own customers. Read each against the ladder rather than as a template: none of these operators' positions is reachable by imitation.

Three reference points read against the ladder

Outcomes as reported in each operator's own published material. Verify figures against the linked source before reusing them; we have not independently audited them.

Amazon fulfilment centre operations with roboticsAmazonGlobal retail & logistics network · 1.5M+ employees45
Challenge
Operating fulfilment at a scale where no vendor's product roadmap could be allowed to set the ceiling on warehouse decision automation — and where integration depth, not model quality, is the binding constraint.
Approach
Vertical integration of the AI supply chain itself: Amazon builds and deploys its own robotics and AI systems in-house — including the robotic systems and AI-driven inventory-handling technologies covered in its operations newsroom — rather than assembling them from vendor products.
Reported outcome
Amazon has publicly reported deploying robotics at the scale of hundreds of thousands of mobile units across its network, alongside AI systems it states improve inventory handling and fulfilment speed — reported by Amazon's own operations newsroom.
What it shows about the curveThe build pole exists, and it is priced in engineering headcount very few operators have. The transferable lesson is not 'build' — it is that Amazon owns its integration layer and decision data completely, which is exactly the asset a buying operator can own at any scale while still renting the models.

Amazon operations newsroom (opens in a new tab)

Freight forwarding and logistics technology operationsFlexport / ConvoyGlobal freight forwarder · logistics technology platform34
Challenge
Convoy, a digital freight network that had raised $260M at a $3.8B valuation in April 2022, ceased operations in October 2023 — roughly 18 months later — leaving shippers and carriers that depended on its AI-driven matching with weeks to re-route freight.
Approach
Flexport acquired Convoy's technology stack in November 2023, as publicly announced at the time, and later relaunched the platform's capabilities within its own product family — preserving the technology while the vendor entity itself disappeared.
Reported outcome
As publicly reported: the market's freight kept moving, but the transition cost landed on buyers in proportion to their exit readiness — those with documented fallbacks and portable tender flows migrated in weeks; the technology's survival inside Flexport did not spare anyone the migration.
What it shows about the curveVendor continuity is a buyer-side readiness dimension. A well-funded, well-regarded vendor exited the market inside 18 months; nothing in a demo, an RFP score or a reference call would have predicted it. Contract continuity terms and a drilled exit are the only instruments that pay on the day it happens.

Flexport (public announcements) (opens in a new tab)

Maersk container shipping and logistics operationsMaerskGlobal container carrier & logistics integrator34
Challenge
A capacity supplier competing in a digitising market where its shippers increasingly buy visibility, planning and exception-handling capability as software — from third parties, unless the carrier offers it.
Approach
Maersk built a digital-solutions portfolio — visibility, logistics management and data products marketed directly to its shipping customers — positioning the carrier itself as a technology vendor on top of its physical network.
Reported outcome
Maersk publicly markets an expanding portfolio of digital and data products to its customers via its digital-solutions platform — the carrier's own published material is the source for its positioning and capabilities.
What it shows about the curveFor buyers, the supplier-turned-vendor changes the diligence, not the discipline: capability bundled with capacity still deserves the same scorecard — baseline, integration proof, data terms — plus one extra question no standalone vendor raises: what does this data relationship do to our negotiating position on the freight itself?

Maersk — digital solutions (opens in a new tab)

Designing a vendor trial that predicts production

The paid proof of value is the centre of evidence-led buying. Here is the term sheet, the sequence, and the disciplines that keep a trial from becoming a longer demo.

A vendor trial predicts production only when it reproduces production's conditions: your data at its real quality, your lanes including the hard ones, your systems at both ends of the integration, and your definition of success in operational units. Anything less measures the vendor's ability to run trials, which is a skill vendors are uniformly excellent at. The PoV term sheet below is the instrument that holds those conditions in place — agree it before the first vendor conversation, not during them.

TermSet it toWhy it matters
Duration6–10 weeks on live feeds, spanning at least one demand disturbance (promotion, month-end, weather week)Shorter trials measure onboarding polish; a trial that sees no disturbance predicts nothing about peak
DataYour lanes, your live TMS/WMS/telematics feeds, your real master data — flaws includedCleansing the data first tests a network you do not operate
Baseline & holdoutBaseline computed from six months of your history before vendors are briefed; holdout lanes chosen by you, revealed only at scoringThe baseline defines 'better'; the holdout prevents both vendor tuning and champion bias
Success metricOperational units a budget holder already tracks: ETA minutes, OTIF points, dwell minutes, cost per shipmentModel metrics (accuracy, error) cannot fund a contract; operational deltas can
Who paysYou do — a fair, bounded feeA paid trial filters for vendors willing to be measured, and removes 'under-investment' as the excuse for failure
Vendors in parallelTwo, on identical data and scoringThe second vendor costs little and is your entire negotiating position at signature
Exit from trialData returned or deleted with attestation; no auto-conversion to a contract; both outcomes pre-pricedA trial that can only end in signature was a sales process with extra steps
Learnings & IPFindings on your data documented and yours; improvements trained on your data acknowledged in the term sheetThe trial itself generates the first entries of the data-terms conversation
The PoV term sheet for a logistics AI trial. Every row is agreed in writing before the trial starts; the rows are where trials quietly rot when left implicit.

The order matters

  1. Baseline before briefing

    Compute your own numbers before any vendor learns the scope. A baseline built after vendors are engaged inherits their framing of the problem — and their choice of which lanes count. The baseline is also your insurance policy: if the trial fails, you have still built the measurement asset the next selection will use.

  2. Terms on the table before the trial, signature after the scoring

    Circulate the data, exit and SLA term sheet to both vendors at trial start, while each still fears losing. By scoring day you hold evidence and competition simultaneously — the only configuration in which a buyer of logistics AI ever has genuine leverage. Sign only after both are in hand.

  3. Two vendors to the end, even when one is obviously winning

    The temptation at week six is to release the trailing vendor and save the fee. The moment you do, the leading vendor's standard contract reappears. Run both to scoring, negotiate with both results on the table, and treat the second fee as what it is: the price of the terms you are about to obtain.

One benchmark discipline specific to freight: where the product's claim is commercial — a marketplace's savings, a rate-prediction tool's advantage — score it against independent market data for the same lanes and weeks, such as DAT's rate benchmarks (opens in a new tab), not against your historical spend alone. Freight markets move; a vendor evaluated in a softening market will claim as skill what the market gave everyone. The disturbance-spanning trial window and the external benchmark exist for the same reason: to separate the vendor's contribution from the world's.

A 90-day plan: selecting a predictive-ETA vendor for one inbound network

Evidence-led selection made concrete on one common logistics purchase — a visibility/ETA platform for inbound linehaul — from baseline to signed, governed contract in one quarter.

The whole evidence-led method fits in 90 days when it is scoped to one purchase decision. The plan below runs it on the most commonly bought logistics AI product — a predictive-ETA and visibility platform — for one inbound linehaul network: roughly 40 lanes, mixed carrier EDI quality, appointment-driven DCs where a bad ETA costs dwell and detention. Nothing in the plan requires a data platform, a transformation programme, or any modelling capability of your own — the quarter's work is baselining, trial-running and negotiating.

Evidence-led ETA vendor selection, in one quarter

One inbound network, two vendors, one signature. If any phase needs more than its window, narrow the lane set rather than extending the plan — the method survives scope cuts; the calendar rarely survives extensions.

  1. Days 1–15

    Baseline and design the trial

    Compute current ETA performance from six months of TMS events and carrier EDI 214 history: mean absolute error by lane and carrier, late-arrival detection rate, and the dwell and detention cost of ETA misses at your appointment-driven DCs. Choose ~40 trial lanes and 8–10 holdout lanes covering the hard cases — sparse-EDI carriers, seasonal lanes. Write the PoV term sheet and the data/exit/SLA term sheet. Name the selection owner: the inbound transport manager whose dwell numbers are at stake.

    A computed baseline, a trial design, both term sheets drafted

  2. Days 16–55

    Run two vendors on live feeds

    Onboard two shortlisted vendors to identical live data — TMS events, carrier EDI, telematics where present. Each vendor also completes the integration spike in this window: consuming your real feed and writing one predicted-ETA field back into a test instance of your TMS, with both sides' effort logged. Circulate the commercial term sheet to both vendors mid-trial, while the competition is live. Let the trial window span month-end volume.

    Two vendors live on your lanes; integration effort measured on both sides

  3. Days 56–75

    Score on holdouts, negotiate with both results open

    Reveal the holdout lanes and score both vendors against your baseline in operational units: ETA error by lane, late-arrival detection, and modelled dwell/detention impact at the DCs. Take the results into commercial negotiation with both vendors still engaged: data ownership, export format and cadence, exit assistance, an SLA metric that matches the trial's scoring, and continuity terms proportionate to the vendor's size and funding condition.

    A scored comparison and a negotiated contract with terms set under competition

  4. Days 76–90

    Sign, wire the write-back, stand up governance

    Sign the winner. Promote the spike integration to production: predicted ETAs written into the TMS fields and appointment board your planners already use, with the previous ETA source one switch away as fallback. Register the vendor — owner, contracted metric, renewal date, annual cost — and schedule the telemetry file: monthly ETA error and alert-acceptance from your own systems, plus a quarterly holdout re-check. The trial harness becomes the governance harness.

    A governed production capability — not a pilot — at day 90

The quarter's least visible deliverable is the most durable one: the baseline and scoring harness outlive the contract. At the first renewal gate, the same harness re-scores the incumbent; if the relationship ever ends — by your choice or, as the market has demonstrated, the vendor's — the same fallback switch and export terms you wired at day 90 are the difference between a migration and an incident.

Exit readiness: the test of whether you ever owned the capability

Lock-in is not something vendors do to you; it is something architectures and contracts permit. Eight conditions determine whether any vendor in your estate could be replaced at a known cost.

Exit readiness is the ability to replace any AI vendor at a known, bounded cost — and it is the single best proxy for overall vendor readiness, because every discipline on this page compounds into it: the baseline that makes an alternative comparable, the owned integration layer that makes swapping mechanical, the export and assistance terms fixed at signature, the register that sees the renewal coming. An operator that can exit does not need to exit often; the priced walk-away changes every negotiation that precedes it.

The unglamorous foundation is data portability, and logistics is unusually well served here if buyers insist on it: GS1 standards (opens in a new tab) give shipments, locations and logistics units stable, vendor-neutral identifiers, and EPCIS (opens in a new tab) defines a standard event vocabulary for what happened, where and why. A vendor whose exports resolve to standard identifiers and event semantics can be re-loaded into a successor in weeks; a vendor whose exports are proprietary journal dumps keyed to internal IDs cannot — and the difference was decided by a schema question nobody asked at signature. Make export format and cadence a term-sheet row, and re-load one export into a scratch environment once a year to prove the clause is real.

The exit-readiness checklist

Eight conditions, checkable against your estate this week. Score any vendor you depend on — or the estate as a whole. This list works without JavaScript; tick as you verify.

0 of 8 ticked

0 of 8 — every vendor in the estate is load-bearing

No ticks means every exit is unpriced and every renewal is a formality for the vendor. Don't start with the architecture; start with the register and one documented fallback for the vendor whose loss would hurt most this quarter. Both are paperwork, not projects.

Vendor failure modes that send buying readiness backwards

Vendor readiness is not monotonic either. Four failure modes account for most of the regression — and only one of them originates with the vendor.

Buying capability regresses the same way adoption maturity does: quietly, while the artefacts that defined the stage stop being maintained. Registers go stale, baselines stop being computed, rehearsals get skipped in a busy peak — and the estate slides back toward entanglement without a single decision being taken. Four failure modes account for most of it.

Likelihood: mediumImpact: high

The vendor disappears — shutdown, acquisition or pivot

This market removes vendors quickly: Convoy went from a $3.8B valuation to shutdown in roughly 18 months, and point solutions are routinely acquired by suites that sunset them into bundles. The operator's exposure is set entirely before the event — by continuity terms, export re-loads and the drilled fallback — because after the announcement there is no leverage left to negotiate with.

PreventionContinuity terms proportionate to vendor size at signature; renewal-date and funding-condition tracking in the register; one rehearsed exit per year.

Likelihood: highImpact: medium

AI-washing absorbs the renewal

The suite renewal arrives with an AI module bundled in, the QBR deck shows impressive prediction volumes, and nobody re-tests whether the capability beats the operator's own baseline — or whether it is materially different from the rules engine it replaced. The estate's costs ratchet up at each renewal while its measured value stays flat, and the gap is invisible because nothing is being measured.

PreventionEvery renewal passes the same gate a new vendor would: telemetry from your own systems against the contracted metric, and a holdout re-check for material claims.

Likelihood: highImpact: high

The integration hostage

Each vendor was wired point-to-point into the TMS or WMS because it was two weeks faster that way, and five vendors later the estate cannot be recomposed: replacing any of them costs more than tolerating it, which every vendor's pricing quietly reflects. This failure mode originates with the buyer, compounds silently, and is the single largest destroyer of renewal leverage in logistics AI estates.

PreventionRoute every NEW vendor through an owned interface from day one; migrate one incumbent per year, most-expensive-exit first; portability terms before any scope expansion.

Likelihood: mediumImpact: medium

The pilot graveyard poisons the well

Successive demo-led purchases fail in succession, and the organisation learns the wrong lesson: that logistics AI does not work, rather than that the buying method could not tell working from non-working. The next genuinely capable vendor is then evaluated by a cynical organisation with frozen budgets — the anti-readiness position, where the market's real value is inaccessible because the buyer's evidence machinery never existed.

PreventionNo purchase without a baseline and a write-back path; retire the word 'pilot' in favour of PoVs with a pre-agreed production route and pre-agreed kill criteria.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Vendor readiness
The buyer-side capability to evaluate, integrate, govern and exit AI vendors — evidence-led evaluation, integration surfaces the operator controls, terms fixed at signature, and in-production governance from the operator's own telemetry.
Proof of value (PoV)
A paid, bounded trial run on the buyer's own lanes and live feeds, scored against the buyer's baseline in operational units. Distinct from a pilot: it has a pre-agreed production route, kill criteria, and at least two competing vendors.
Baseline
The operator's own measured performance on a decision — ETA error, slotting travel time, tender acceptance — computed from TMS/WMS history before vendors are engaged. The artefact that makes vendor claims testable and alternatives comparable.
Holdout lanes
Lanes, doors or facilities kept back from every vendor during a trial and revealed only at scoring. They defend the evaluation against vendor tuning and against the buyer's own champions choosing flattering scope.
Integration spike
A short pre-signature exercise — one real feed into the vendor's system and one write-back into a test instance of the buyer's TMS or WMS — that converts integration effort from a proposal estimate into a measured number.
AI-washing
Presenting existing rules, heuristics or thin statistical features as AI capability — most consequential at suite renewals, where bundled 'AI modules' ratchet cost without ever being tested against the buyer's baseline.
Vendor register
The single list of every AI vendor in the estate: decision served, named owner, contracted metric, latest telemetry reading, annual cost, renewal date and funding condition. The artefact that turns contracts into a portfolio.
Renewal gate
A scheduled evidence review before each contract renewal with three honest outcomes — extend, renegotiate or re-compete — decided on the operator's own telemetry rather than the vendor's QBR. It has teeth only when the exit is priced.
Exit readiness
The ability to replace any vendor at a known, bounded cost: portable exports, an owned integration layer, documented fallbacks, exit-assistance terms and a rehearsed failover. The priced walk-away that gives every other discipline its leverage.
EPCIS
GS1's standard for sharing supply-chain event data — what happened, where, when and why, in a vendor-neutral vocabulary. Exports that resolve to EPCIS-style events and GS1 identifiers are re-loadable into a successor system; proprietary journal dumps are not.
Continuity terms
Contract provisions that survive a vendor's failure or acquisition: source and model escrow, step-in rights, runbook handover, and data-return obligations — sized in proportion to the vendor's funding condition and customer concentration.
Shelfware
A licensed product no operational decision depends on — the characteristic output of demo-led buying, and invisible in estates with no vendor register because its subscription renews by inertia.

Frequently asked questions

The questions logistics operators ask most often when a vendor selection or renewal is on the table.

How should a logistics operator evaluate an AI vendor?

Run a paid proof of value on your own lanes and live feeds, scored against a baseline you computed before vendors were briefed, with holdout lanes revealed only at scoring — and run two vendors in parallel on identical data. Demos and references are seller-controlled evidence and predict nothing about production. Complete the picture with an integration spike (one real feed, both directions) and the six-area scorecard: integration proof, model evidence, data terms, operational support, viability, and security assurance.

What belongs in a logistics AI RFP that a normal software RFP misses?

Four things. A requirement to beat your computed baseline on your own data, replacing most of the feature matrix. An integration spike with your actual TMS/WMS feeds before signature, with both sides' effort logged. Data terms as selection criteria — ownership, export format, exit assistance — not post-award boilerplate. And continuity questions: funding condition, customer concentration, escrow. The general principle: convert every question a vendor can answer in prose into one it must answer with behaviour.

Should we build or buy AI for logistics operations?

Buy by default, and decide per decision domain rather than as a company identity. No operator will out-build a visibility network's cross-shipper data advantage, so commodity domains — visibility, freight audit, standard forecasting — stay bought and get re-competed at renewal. Domains where your network shape or data is genuinely distinctive are candidates to build once an owned integration layer exists, because the layer that makes vendors replaceable is the same substrate an internal model plugs into. Amazon marks the build extreme; its transferable lesson is owning the integration layer and decision data, not building everything.

How long should a vendor trial run, and who should pay for it?

Six to ten weeks on live feeds, deliberately spanning at least one demand disturbance — month-end, a promotion, a weather week — because a trial that only sees calm weeks predicts nothing about peak. The buyer should pay a fair, bounded fee. Paid trials filter for vendors confident enough to be measured, remove 'under-investment' as the post-hoc excuse for failure, and keep the buyer honest about scoping. Free pilots are sales activity, and both sides treat them accordingly.

What data terms matter most in a logistics AI contract?

Five, all negotiated at signature while competition is alive: ownership of the operational data you feed the vendor; export format and cadence — ideally resolving to standard identifiers and event semantics such as GS1 keys and EPCIS-style events; exit assistance with named duration and cost; rights concerning model improvements trained on your data; and the SLA metric, which should match your trial's scoring so the contract measures what the evaluation measured. Silence on any of these resolves in the vendor's favour.

How do we avoid vendor lock-in with visibility platforms?

Accept the network effect — it is the product's genuine value — and contain the entanglement. Route the platform's feeds through an integration layer you own rather than wiring its portal into planner workflows directly; keep computing your own ETA baseline monthly even while the platform performs; fix export terms at signature and re-load one export annually to prove they work; and keep a documented fallback (the previous ETA source, one switch away). Lock-in is not something vendors do to you; it is something architectures and contracts permit.

What did the Convoy shutdown teach logistics AI buyers?

That vendor continuity is a buyer-side readiness dimension, not a diligence footnote. Convoy raised $260M at a $3.8B valuation in April 2022 and ceased operations in October 2023 — roughly 18 months later — and nothing visible in a demo, an RFP score or a reference call predicted it. Flexport's acquisition of the technology preserved the software but spared no buyer the migration. Operators with documented fallbacks and portable tender flows migrated in weeks; point-to-point integrated peers ran war rooms. The instruments that paid were all pre-purchased: continuity terms, export portability, a rehearsed exit.

Are the AI modules from our incumbent TMS or WMS vendor good enough?

Sometimes — and the honest answer is measurable rather than reputational. Suite modules win on integration cost: they inherit the workflow, the data and the support contract. They lose, when they lose, on depth against AI-native point solutions. The discipline is to subject the bundled module to the same gate a new vendor would face: does it beat your computed baseline on your lanes? Gartner expects more than 75% of supply-chain application vendors to ship embedded AI by 2026, so this question recurs at every suite renewal — treat 'included in the bundle' as a price, not as evidence.

How many AI vendors should a logistics operator run?

The count matters less than the substrate. Five vendors integrated point-to-point are five hostages; five behind an owned integration layer are a portfolio. That said, unmanaged estates accumulate overlap — two products computing ETAs for different teams is the classic register finding — so the practical discipline is one vendor per decision domain unless a deliberate reason says otherwise, a register that makes every overlap visible, and a renewal gate that retires the weaker of any overlapping pair on telemetry rather than on politics.

What certifications and assurances should we ask an AI vendor for?

Third-party security attestation such as SOC 2 (or an equivalent regime) as table stakes, with ISO/IEC 42001 — the AI management system standard — an emerging signal that a vendor governs its models formally. Two logistics-specific checks matter as much as certificates: data provenance (carrier-vetting products, for example, are substantially built on public FMCSA registry data, and you should know which parts of a product are proprietary intelligence versus repackaged public data), and a kill-switch story — a tested way to disable the model's influence on operations during an incident, available at 02:00 in your timezone during peak.

How does vendor readiness relate to data readiness?

They gate each other, and they are different disciplines. Data readiness — pipeline quality, master data, event coverage — determines how fast a vendor trial can start and how well any vendor's model can perform on your freight; a trial that spends six of its eight weeks on data archaeology tests nothing. Vendor readiness determines whether the purchase built on that data serves you — evidence at selection, terms at signature, telemetry at renewal. Operators strong in one and weak in the other fail differently: good data with demo-led buying produces well-fed shelfware; disciplined buying on broken data produces honest trials that fail slowly. This page covers the buying side; the data side has its own.

Do 3PLs and shippers need different vendor readiness?

The disciplines are identical; two constraints differ. A 3PL evaluates vendors against multi-client operations, so trials must test tenant separation and per-contract attribution — a slotting product that cannot report value per client contract fails a 3PL need a shipper never has. And 3PLs face the supplier-turned-vendor question more sharply, since carriers and even customers now ship AI products into the same workflows. Shippers, owning more of their estate, can run integration spikes faster and should exploit that: their route from stage 2 to stage 3 buying is typically one selection cycle shorter.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for manufacturing, logistics and energy operators — forecasting, routing, vision inspection and decision support running against live operational data, integrated into the TMS and WMS layer rather than delivered as dashboards. That vantage point covers both sides of this page: we are evaluated as a vendor, and we sit beside operators evaluating others.

  • · Production deployments across freight, warehousing and last-mile
  • · Vendor evaluations and PoV trials run jointly with operator teams
  • · Integration-first delivery: TMS/WMS write-back, monitoring, rollback
  • · 11 cited sources on this page

Sources

  1. GartnerSupply chain artificial intelligence (topic landing) (opens in a new tab)
  2. MHIAnnual Industry Report (opens in a new tab)
  3. McKinsey & CompanyOperations insights (supply-chain AI research landing) (opens in a new tab)
  4. FlexportFlexport (Convoy technology acquisition, public announcements) (opens in a new tab)
  5. Uber FreightUber Freight (opens in a new tab)
  6. DAT Freight & AnalyticsFreight market benchmark data (opens in a new tab)
  7. AmazonOperations newsroom (robotics and AI) (opens in a new tab)
  8. GS1GS1 standards (opens in a new tab)
  9. GS1EPCIS event data standard (opens in a new tab)
  10. US Department of TransportationFederal Motor Carrier Safety Administration (opens in a new tab)
  11. MaerskDigital solutions (opens in a new tab)

Buy your next AI capability with the leverage on your side

We run the assessment with your operations and procurement leads, benchmark the four buying disciplines against comparable operators, and leave you with a trial design and term sheet for the selection or renewal you are actually facing. You keep all of it whether or not we build with you.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.