LogisticsReadiness & Transformation Roadmap
Logistics AI readiness for vendors: evaluating, buying and governing the AI you don't build
Logistics AI vendor readiness is an operator's capability to evaluate, integrate, govern and — when necessary — exit the AI it buys rather than builds. It rests on four disciplines: evidence-led evaluation, integration readiness, commercial and exit terms, and in-production vendor governance. In a market where every TMS and WMS vendor now claims AI, that capability decides what a contract actually delivers.

Key takeaways
- Vendor readiness is a buyer capability, not a market condition. The same logistics AI vendor produces shelfware at one operator and attributable OTIF gains at another; the difference is almost always the buyer's baseline, integration surface and terms — not the vendor's model.
- Gartner expects more than 75% of commercial supply-chain application vendors to ship embedded AI and advanced analytics by 2026 — which means 'has AI' no longer discriminates between vendors. The discriminating questions are integration proof, data terms and in-production evidence.
- The demo is a controlled experiment run by the seller. The only evaluation that predicts production is a paid proof of value on your own lanes and live feeds, scored against your own baseline with holdout lanes kept back.
- Signature is the last moment of leverage. Data ownership, export formats, exit assistance and SLA metrics negotiated before signing cost almost nothing; the same terms requested at renewal are priced against you.
- Vendor continuity is a readiness dimension, not a due-diligence footnote: roughly 18 months separated Convoy's final raise at a $3.8B valuation from its 2023 shutdown. Operators with a drilled exit experienced a migration; the rest experienced an incident.
Abbreviations used on this page
- TMS
- Transport management system
- WMS
- Warehouse management system
- YMS
- Yard management system
- EDI
- Electronic data interchange (e.g. the 214 shipment status message)
- API
- Application programming interface
- ETA
- Estimated time of arrival
- OTIF
- On-time in-full delivery rate
- 3PL
- Third-party logistics provider
- RFP
- Request for proposal
- PoV
- Proof of value — a paid, bounded trial on the buyer's own data
- SLA
- Service-level agreement
- QBR
- Quarterly business review — the vendor-run account meeting
Free · 8 questions · ~3 minutes
Score how ready you are to buy AI
Eight questions, one at a time, about three minutes — on how your operation evaluates, integrates, contracts and governs its AI vendors. Answer them and we build your vendor readiness report: your stage on the buying ladder, your score on each of the four disciplines, and the specific gap most likely to cost you at your next selection or renewal. It arrives by email.
0 of 8 answered
Pick an option to continue
Report ready
Your vendor readiness report is ready
Tell us where to send it. Your stage appears on screen straight away; the full report — dimension scores, the buying-ladder gap analysis, and the trial and term-sheet checklist for your weakest discipline — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Demo-led
Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
Your next moveBaseline one decision — ETA accuracy, slotting travel time, tender acceptance — from your own TMS/WMS history, and require the next candidate vendor to beat it on your data.
Stage 2 · Checklist-led
Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
Your next moveReplace the demo-scoring round of your next selection with a paid proof of value on your own lanes and live feeds, scored against your baseline — and put data and exit terms on the table while two vendors are still competing.
Stage 3 · Evidence-led
Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
Your next moveStand up a vendor register — every AI vendor, its decision, its owner, its renewal date, its contracted metric — and start computing each vendor's performance from your own telemetry rather than the vendor's QBR.
Stage 4 · Production-governed
Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
Your next moveMove vendor integrations onto a layer you own — canonical events, standard identifiers, documented schemas — and rehearse one exit end-to-end, so that at the next gate, switching has a known price.
Stage 5 · Portfolio-orchestrated
The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.
Your next moveKeep the annual exit rehearsal and the per-domain build-versus-buy review on the calendar — the stage regresses quietly the year both are skipped.
0 / 24
Evaluation discipline
— / 6
Integration readiness
— / 6
Commercial & exit terms
— / 6
Vendor governance
— / 6
Your score maps to a stage on the buying ladder. Read the dimension breakdown before the total: the lowest of the four disciplines is where your next vendor will extract its margin, and it is where the next quarter's effort belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the buying ladder. Read the dimension breakdown before the total: the lowest of the four disciplines is where your next vendor will extract its margin, and it is where the next quarter's effort belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Have a live selection or renewal coming?
We will review your dimension scores against the selection or renewal you are actually facing, and leave you with a trial design, a term sheet and the telemetry file to run the gate with. You keep all three whether or not we work together.
How the score maps to a stage
- 0–5 — Stage 1, Demo-led. Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
- 6–11 — Stage 2, Checklist-led. Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
- 12–16 — Stage 3, Evidence-led. Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
- 17–21 — Stage 4, Production-governed. Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
- 22–24 — Stage 5, Portfolio-orchestrated. The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.
What logistics AI vendor readiness is — and why the demo is the least of it
A definition, the buying curve, and the two paths a vendor claim can travel from demonstration to your production stack.
Logistics AI vendor readiness is the organisational capability to buy AI well: to evaluate vendor claims against your own evidence, to feed and receive vendor systems through integration surfaces you control, to fix data and exit terms while you still have leverage, and to govern vendor performance from your own telemetry for the life of the contract. It is a property of the buyer, not the market — which is why the same vendor produces shelfware at one operator and attributable OTIF gains at another.
The capability matters more every year for a structural reason: AI is ceasing to be a differentiator among logistics vendors and becoming table stakes. Gartner's supply-chain AI research (opens in a new tab) projects that by 2026 more than 75% of commercial supply-chain management application vendors will deliver embedded advanced analytics and AI. When every TMS, WMS and visibility product 'has AI', the phrase stops discriminating between vendors — and the discriminating work moves to the buyer: integration proof, data terms, and evidence from your own lanes. This page is about building that muscle. What it is not about: sensor estates (the IoT readiness page), pipeline and master-data quality (the data readiness page), or which warehouse or freight initiative to sequence first (their own roadmap pages). Those disciplines feed this one; the buying capability is its own.
Value released against buying maturity
The curve is not linear. Through the demo-led and checklist-led stages, contracts accumulate faster than production value — spend rises while decisions stay unchanged. Value inflects when evaluation moves onto the operator's own data and terms are set at signature, and compounds when the stack becomes governable and finally composable.
Production value released per contracted pound by stage
- Stage 1 · Demo-led — 24% of operators. Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
- Stage 2 · Checklist-led — 37% of operators. Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
- Stage 3 · Evidence-led — 26% of operators. Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
- Stage 4 · Production-governed — 10% of operators. Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
- Stage 5 · Portfolio-orchestrated — 3% of operators. The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with Gartner's supply-chain AI adoption research.
Two paths from vendor claim to production — and where each leaks
The same vendor claim can travel two paths through a logistics operation. The demo-led path terminates in shelfware renewed by sunk cost; the evidence-led path converts the claim into a governed production capability — and the governance lane is what keeps it converted. Most logistics vendor selection still travels the top lane.
- AI / model
- System-of-record action
- Where value leaks
- Data & feeds
- Human in the loop
The process, in words
- On the demo-led path, evaluation runs on the vendor's curated sample data and scripted pilot, terms are signed with integration unpriced, and the product meets the operator's EDI reality only after signature. The stall that follows is absorbed as sunk cost, and the contract renews by inertia — shelfware with a maintenance fee.
- On the evidence-led path, the operator baselines its own lanes first, runs a paid PoV for at least two vendors on the same live feeds with holdout lanes kept back, tests integration with one real feed in both directions before signing, and fixes data ownership, exit assistance and SLA metrics while competition is still alive.
- The production-governance lane is what keeps the evidence honest after go-live: a vendor register with a named owner per vendor, performance computed from the operator's own telemetry rather than the vendor's QBR, renewal gates decided on that evidence, and a drilled exit that keeps the walk-away real.
Step-by-step insights
- The demo is a controlled experiment run by the seller
- Nothing about a demo is dishonest, and nothing about it is evidence. The data was chosen to flatter the model, the scenario was rehearsed, and the edge cases that dominate a real logistics network — carriers that transmit three of nine EDI 214 event types, item masters with duplicate SKUs, the peak-week data your planners actually live with — are precisely what the demo excludes. Treat demos as theatre with a briefing function: useful for understanding what a product intends to do, structurally incapable of predicting what it will do on your freight.
- The baseline is the buyer's only native advantage
- The vendor knows its product; the operator knows its network. A computed baseline — your current ETA error by lane, your slotting travel times, your tender acceptance by carrier, from six months of TMS/WMS history — converts that knowledge into the one artefact no vendor can supply or dispute. It is the difference between asking 'is this product good?' (unanswerable) and 'does this beat what we already achieve, on our lanes, by enough to fund the integration?' (a number). Operators who skip the baseline are outsourcing the definition of success to the party being paid to succeed.
- Holdout lanes keep everyone honest — including you
- A holdout is a set of lanes, doors or facilities the vendor never sees until scoring day. It defends against the two corruptions every trial invites: the vendor tuning to the test set, and the buyer's own champions unconsciously selecting the lanes where the product shines. Choose holdouts to represent the network's hard cases — the carrier with sparse EDI, the seasonal lane, the facility with the messy master data — because production is made of hard cases, and a product that only works on clean lanes will meet the rest of your network eventually.
- The integration spike — one real feed beats any architecture slide
- Before signature, run one production feed through the vendor's ingestion and one write-back into a test instance of your TMS or WMS — a two-week spike, both directions. This single exercise surfaces what no RFP answer can: the event types the product actually consumes, the master-data assumptions it makes, the latency it adds, the effort your side genuinely requires. Vendors confident in their integration story agree readily; hesitation is itself diligence data. The spike converts the largest unknown cost in the contract into a tested number while you can still walk away.
- Terms at signature — the only moment of leverage
- Data ownership, export formats, exit assistance, model-improvement rights and the SLA metric all cost approximately nothing to obtain while two vendors are still competing, and are priced as change requests the day after signature. The asymmetry is structural: before signing, the vendor is selling; after, you are retained. The evidence-led path's commercial payoff is exactly here — the trial's real product is not just a selection but a signed contract whose terms were set under competition. A stage-3 trial that ends in a stage-2 contract wasted its leverage.
- Own telemetry, the renewal gate and the drilled exit
- After go-live, the trial harness keeps running: acceptance rates from your approval logs, ETA error from your TMS events, OTIF from your customer EDI — your numbers, on a file the renewal gate opens. The gate has three honest outcomes: extend on evidence, renegotiate on evidence, or re-compete. But the third outcome exists only if the exit is real — fallback documented, export re-loaded at least once, elapsed time known. That is why the drilled exit sits at the end of the governance lane: it is the difference between a gate and a ceremony.
The five stages of the buying ladder in detail
From demo-led to portfolio-orchestrated: what each stage looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps operators there, and what leaving costs.
Each stage below describes how an operator buys, not what it owns — an operator with a dozen AI contracts can sit at stage 1, and a lean operation with two well-governed vendors can sit at stage 4. The hallmarks are observable conditions, the diagnostic signals are checks you can run against your own contracts and telemetry this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Demo-led
24% of operators sit here
Vendors are selected from demonstrations and references; whether the product works on the operator's own freight is discovered after signature.
Stage 1 is not gullibility — it is asymmetry. The vendor has run its demonstration hundreds of times against data curated to flatter the product; the operator is seeing it once, without a baseline of its own to compare against. Under that asymmetry, selection quality is capped no matter how experienced the buying team is, because every input to the decision was produced by the seller.
The tell is what happens after signature. Integration effort was a line in the vendor's proposal rather than a tested fact, so the first months of the contract are spent discovering the real shape of the work: the WMS export the product assumed doesn't exist, the EDI 214 feed carries half the events the model needs, the master data has three codings for the same carrier. None of this was hidden — it was simply never tested, because the evaluation never touched the operator's own systems.
The stage is self-perpetuating in a specific way: each disappointing purchase is blamed on the vendor, and the next vendor is selected by exactly the same process. Operators can cycle through three routing or slotting products in five years without once changing the method that chose them — which is the actual defect. The exit from stage 1 is not a better vendor; it is the first internal baseline that makes vendor claims testable.
In practice
The slotting engine bought at a trade show
A regional 3PL signed a warehouse slotting product after a strong conference demo and two reference calls. The demo ran on a model warehouse with clean item master data. The 3PL's own item master, built through two acquisitions, had duplicate SKUs and inconsistent dimensions — the product's first three months were a data-cleansing project nobody had scoped, and the operations team that was promised travel-time savings got a re-keying exercise. The product was quietly shelved before peak; the post-mortem blamed the vendor.
What it looks like
- Shortlists are assembled from analyst charts, trade shows and demos
- Pilots run on vendor sample data, in the vendor's own environment
- Integration effort is estimated by the vendor, after selection
- No internal baseline exists to test any vendor claim against
Diagnostic signals you can check this week
- Ask what data the last vendor demo ran on. If the answer is 'theirs', you are here
- Ask for the internal baseline a current vendor is measured against. Silence is the answer
- Check whether integration cost was a tested number or the vendor's estimate at signature
- Count AI products purchased in three years against those in daily production use
Anti-pattern · Fixing it with a bigger shortlist
The instinctive correction is to see more vendors — six demos instead of two, a scoring sheet, a second round. This multiplies the inputs without changing their nature: every one is still a seller-controlled demonstration. Six curated demos aggregate to no more evidence than one. The move that changes the outcome is not widening the funnel but changing what flows through it — one measured baseline of your own, and a requirement that the next vendor beat it on your data.
What holds you here
There is no internal baseline, so every vendor claim is compared against another vendor claim rather than against the operation's own numbers.
Highest-leverage next move
Baseline one decision — ETA accuracy, slotting travel time, tender acceptance — from your own TMS/WMS history, and require the next candidate vendor to beat it on your data.
Cost of leaving
- Effort
- 2–4 months
- Team
- One operations analyst and one engineer, part-time, to baseline a single decision
- Risk
- Low — baselining your own lanes commits you to nothing and informs everything
- To next stage
- 2–4 months
If this is you, the next step is
A short engagement: pick the decision, compute the baseline, write the test a vendor must pass.
Stage 2
Checklist-led
37% of operators sit here
Procurement is structured — RFPs, weighted matrices, scored demos — but it still tests what vendors say rather than what their systems do on the operator's freight.
Stage 2 looks like rigour and often defeats its own purpose. The RFP process is real: requirements are gathered, criteria are weighted, demos are scored by a panel, references are called. But every artefact in the process is still a claim rather than a behaviour. A 400-line requirements matrix scores what the vendor asserts about the product; it cannot score what the product does when it meets the operator's EDI reality, because nothing in the process ever connects the two.
The structural bias of checklist buying is toward feature count, and feature count favours suites and punishes depth. A vendor with forty features that each work at demo quality outscores a vendor with four features that survive production — the matrix cannot see the difference, and the reference calls can't either, because references are curated by the same seller. Meanwhile the two chapters that decide the contract's fate — integration effort and data terms — sit at the back of the RFP as compliance questions ('Do you provide APIs? Yes.') that every vendor passes.
The costly part is timing. At stage 2 the operator discovers integration reality and negotiates data terms after selection, when leverage is gone: the internal announcement has been made, the project has a name, and the vendor knows it. Terms that would have cost nothing to obtain in a competitive process are now priced as change requests. Most of what operators later describe as lock-in was conceded here, in the weeks after the winner was chosen.
In practice
The 400-line RFP that missed the EDI question
A mid-size freight operator ran a disciplined RFP for a predictive-ETA platform: 412 weighted requirements, four vendors, scored demos. The winning vendor scored highest on functionality and analytics. Nobody had asked which EDI 214 status events the models actually consumed, or tested the operator's own feed against them. Post-signature, it emerged the operator's carriers transmitted a fraction of the event types the product was trained on; ETA quality on the operator's real lanes was far below the demo, and the first renewal became a dispute about whose fault that was. The RFP had scored 412 claims and tested none.
What it looks like
- A formal RFP with weighted criteria governs every AI purchase
- Demos are scored — but still scripted by the vendor
- Feature coverage decides selection; integration is a late chapter
- Data ownership and exit terms surface after the winner is chosen
Diagnostic signals you can check this week
- Open your last AI RFP: count requirements that were tested versus asserted
- Check when data-ownership terms were first discussed — before or after selection
- Ask whether any scored demo ran on your own lanes, doors or item master
- Look at the criteria weights: if feature coverage outweighs integration proof, you are here
Anti-pattern · Adding more rows to the matrix
When a checklist-led purchase disappoints, the reflex is a longer checklist — the 412 requirements become 600, with a new section on AI capabilities that asks vendors to describe their models. This deepens the defect instead of fixing it: more claims, still no behaviour. The correction is to delete most of the matrix and replace its centre of gravity with one requirement that cannot be answered in prose: run against six months of our history and beat our baseline, on lanes we choose, with holdout lanes you never see until scoring.
What holds you here
The evaluation tests vendor claims, not vendor systems — so selection quality is capped by the honesty of the sales process, and terms are negotiated after leverage is gone.
Highest-leverage next move
Replace the demo-scoring round of your next selection with a paid proof of value on your own lanes and live feeds, scored against your baseline — and put data and exit terms on the table while two vendors are still competing.
Cost of leaving
- Effort
- 3–6 months
- Team
- Procurement plus one operations owner and one integration engineer with authority over the evaluation design
- Risk
- Low-to-medium — the change is procedural; the resistance is habit, not cost
- To next stage
- 3–6 months
If this is you, the next step is
We help you design the trial, the baseline and the term sheet before you talk to vendors.
Stage 3
Evidence-led
26% of operators sit here
Selection is decided by paid trials on the operator's own data with holdout lanes, integration is tested before signature, and data and exit terms are negotiated while vendors still compete.
Stage 3 inverts the asymmetry of stages 1 and 2: the operator now controls the experiment. Candidate vendors run against the same six months of TMS and WMS history, on lanes the operator chose, scored against the operator's own baseline in operational units — ETA error in minutes, OTIF points, dwell minutes — with a set of holdout lanes no vendor sees until scoring day. The demo still happens, but it has been demoted from evidence to theatre, which is what it always was.
Two disciplines make the stage real rather than nominal. First, the trial is paid: a fair fee for a bounded PoV filters for vendors confident enough to be measured, and removes the pretext that a free pilot's failure was under-investment. Free pilots are sales activity, and both sides price them accordingly. Second, at least two vendors run to the end. The moment a single vendor knows it is the only candidate, the operator's negotiating position on data terms, exit assistance and SLA metrics collapses — most of stage 3's commercial value comes from keeping the competition alive until signature.
What stage 3 has not yet solved is the portfolio. Each selection is disciplined, but selections accumulate: three years of good individual decisions produce a visibility platform, two point solutions and a suite module with overlapping capabilities, each renewed on its own cycle by whoever signed it. Nothing in the operation can say what the vendor estate costs per decision, which contracts overlap, or which renewal dates are approaching. The selections are governed; the stack is not.
In practice
The two-vendor ETA bake-off
A retailer's transport team ran two visibility vendors in parallel for eight weeks on the same 40 inbound lanes, live EDI 214 and telematics feeds, with 10 holdout lanes scored blind at the end. Vendor A — the analyst-chart favourite — was two minutes better on mean ETA error; vendor B degraded far less during a carrier-mix shift mid-trial and misclassified fewer late arrivals as on-time. The team chose B, and the trial data did double duty: with a competitor still live, B agreed to explicit data-export terms, a defined exit-assistance package and an SLA metric matching the trial's scoring — terms A's standard contract did not offer.
What it looks like
- Paid PoV trials on own lanes and live feeds decide selection
- Holdout lanes are kept back from vendors until scoring
- An integration spike — one real feed, both directions — is part of diligence
- Data ownership, export formats and exit assistance are signed at signature
Diagnostic signals you can check this week
- Ask how the last AI vendor was selected: if the answer names a trial on your data, you are at least here
- Check whether holdout lanes existed and who chose them
- Ask whether the PoV was paid, and how many vendors reached the final week
- Read the signed contract for export formats and exit assistance — present means stage 3, absent means the trial discipline stopped at selection
Anti-pattern · Trialling one vendor
The most common corruption of stage 3 is running a genuine PoV — own data, real baseline, honest scoring — with a single vendor, usually because the trial is framed as a technical validation after a commercial preference has already formed. The evidence is real but the leverage is gone: a vendor that knows it has won concedes nothing at signature, and the operator ends up with stage-3 evidence and stage-2 terms. The trial's cost barely changes with a second vendor; its negotiating value changes completely.
What holds you here
Selections are disciplined one at a time, but the vendor estate as a whole is unmanaged — overlap accumulates, renewals pass by inertia, and nobody can price the stack per decision.
Highest-leverage next move
Stand up a vendor register — every AI vendor, its decision, its owner, its renewal date, its contracted metric — and start computing each vendor's performance from your own telemetry rather than the vendor's QBR.
Cost of leaving
- Effort
- 6–12 months to make evidence-led selection the default across the operation
- Team
- A repeatable trio: operations owner, integration engineer, commercial lead — plus the discipline to fund paid trials
- Risk
- Medium — paid two-vendor trials cost real money per selection; the return arrives at signature and at renewal
- To next stage
- 6–12 months
If this is you, the next step is
Trial design, baseline, holdout selection and the term sheet — set up before vendor conversations start.
Stage 4
Production-governed
10% of operators sit here
Every vendor-served decision is monitored from the operator's own telemetry against contracted metrics; renewals are decided on evidence at a gate, and a named owner answers for each vendor.
Stage 4 extends the trial discipline into the life of the contract. The scoring harness built for the PoV — baseline, operational metric, holdout re-checks — keeps running after go-live, so the operator always knows what each vendor's model is doing on its own lanes this quarter, not what the vendor's QBR deck says it is doing. This sounds like bookkeeping and behaves like leverage: renewal conversations change character entirely when the buyer opens with the vendor's own acceptance-rate trend.
The register is the artefact that makes the stage visible. One list: every AI vendor, the decision it serves, the named owner accountable for its outcome, the contracted metric, the renewal date, the annual cost, the last telemetry reading. Most operators assembling it for the first time find things nobody chose: two products computing ETAs for different teams, a suite module licensed but unused since a champion left, a point solution whose owner departed eighteen months ago while the subscription auto-renewed twice. The register does not fix any of this; it makes it undeniable.
What stage 4 cannot yet do is recompose. Each vendor is well-governed in place, but the integrations are point-to-point — the visibility platform wired directly into the TMS, the slotting engine directly into the WMS — so replacing any of them is a project measured in quarters. The operator can measure, negotiate and even decide to switch; it cannot yet switch cheaply. Renewal gates have teeth only when walking away is priced, and at stage 4 it is still priced too high.
In practice
The renewal that got cheaper
A 3PL's vendor register flagged a visibility platform renewal 120 days out. The telemetry file showed exception-alert acceptance by planners had slid from 71% to 44% over a year — alert fatigue from a threshold change the vendor had shipped without notice. Instead of the standard uplift renewal, the 3PL brought the acceptance curve to the QBR, re-ran two holdout lanes as a spot-check, and renewed at a reduced rate with a contractual alert-precision SLA and quarterly threshold reviews. The register turned a rubber-stamp into a negotiation; the telemetry turned the negotiation into a win.
What it looks like
- A vendor register lists every AI vendor, owner, metric and renewal date
- Vendor performance is computed from the operator's own telemetry
- Renewals pass through a gate: extend, renegotiate or re-compete on evidence
- The QBR reviews the operator's numbers, not the vendor's deck
Diagnostic signals you can check this week
- Ask for the vendor register. Existence, ownership and freshness are the test
- Pick one vendor and ask for its current metric from your telemetry, not its QBR
- Check the last three renewals: how many passed through an evidence gate versus auto-renewing
- Ask who owns each vendor outcome by name — and whether any named owner has left
Anti-pattern · Governing with the vendor's dashboard
The comfortable version of stage 4 monitors every vendor through the analytics screens the vendor itself provides. It feels like governance and measures the vendor's homework with the vendor's ruler: uptime and prediction volume where decision quality should be, accuracy definitions that quietly exclude the hard cases, baselines that moved when the model did. If the number that decides a renewal is rendered by the party being renewed, it is not governance. The telemetry must come from your own TMS, WMS and approval logs, or the gate is decorative.
What holds you here
Integrations are point-to-point, so exits are unpriced and re-competition is theoretical — the gate can renegotiate but cannot credibly walk away.
Highest-leverage next move
Move vendor integrations onto a layer you own — canonical events, standard identifiers, documented schemas — and rehearse one exit end-to-end, so that at the next gate, switching has a known price.
Cost of leaving
- Effort
- 9–18 months
- Team
- A small vendor-governance function — often one commercial owner plus the platform engineer who owns telemetry — and named business owners per vendor
- Risk
- Medium — the telemetry work is modest; the political work of ending auto-renewals is not
- To next stage
- 12–24 months
If this is you, the next step is
We build the register and the telemetry file for your current estate, and run the first renewal gate with you.
Stage 5
Portfolio-orchestrated
3% of operators sit here
The vendor estate is run as a composable portfolio on an integration layer the operator owns: exits are rehearsed, build-versus-buy is decided per decision domain, and vendors compete at every renewal.
Stage 5 is not vendor independence — it is vendor optionality. The operator still buys most of its AI, and should: no logistics operator will out-build a visibility network's data advantage or a suite vendor's integration surface across its own product. What changes is the architecture of the relationship. Every vendor consumes and returns data through an integration layer the operator owns — canonical shipment and warehouse events, GS1-style identifiers, documented schemas — so a vendor is a replaceable module in a stack the operator composes, rather than a load-bearing wall the stack was built around.
The discipline that keeps the stage honest is the rehearsed exit. Once a year, one vendor-served decision is deliberately failed over to its documented fallback — the previous rule set, an alternative product, a manual procedure — on a quiet lane set, and the elapsed time and pain are recorded. The first rehearsal is always humbling: the export that had never been re-loaded, the fallback rule set nobody had updated since go-live. But an operator that has rehearsed an exit negotiates differently, buys differently and — as Convoy's customers discovered — absorbs a vendor shutdown as a migration rather than an incident.
Build-versus-buy becomes a live, per-domain decision rather than an identity. Commodity decision domains — visibility, freight audit, standard forecasting — stay bought and are re-competed at renewal. Domains where the operator's data or network shape is genuinely distinctive become candidates to build, because the integration layer that made vendors replaceable is the same substrate an internal model plugs into. Stage 5 operators move specific decisions in-house not on principle but on arithmetic, and sometimes move them back.
In practice
The 30-day migration that wasn't an incident
When a digital freight network in a shipper's carrier mix wound down operations with weeks of notice — as Convoy's market learned can happen in 2023 — one high-volume shipper rerouted its affected tender flow to its documented backup: contract carriers plus a second marketplace, pre-integrated through its own tendering layer and last rehearsed eight months earlier. Spot exposure rose for a month and settled. Its peer, integrated point-to-point with the same network, ran a war room for six weeks re-keying tenders. Same market event; the difference was entirely on the buyer's side.
What it looks like
- All vendor traffic flows through an owned integration layer with standard events
- At least one vendor exit has been executed or rehearsed in the last year
- Build-versus-buy is decided per decision domain, and revisited
- Renewal gates carry real alternatives — the walk-away is priced and tested
Diagnostic signals you can check this week
- Ask when a vendor exit was last rehearsed, and for the written result
- Trace one vendor's data path: through an owned layer, or wired directly into the TMS/WMS?
- Ask which decision domains are deliberately built versus bought, and who decided
- Check whether any renewal in the last two years was actually re-competed — not threatened, done
Anti-pattern · Confusing multi-vendor with resilient
Operators sometimes claim stage 5 because they run many vendors — five AI products must surely be a portfolio. Five vendors integrated point-to-point are not a portfolio; they are five hostages, each with its own bespoke integration debt and unpriced exit. The count of vendors is irrelevant. The stage is defined by the layer underneath them: one owned, standard, documented integration surface that makes any single vendor — including the biggest — replaceable at a known cost. Resilience lives in the substrate, not the roster.
What holds you here
Sustaining leverage as the vendor market consolidates — every acquisition of a point solution by a suite, and every network's data advantage, pulls the stack back toward entanglement.
Highest-leverage next move
Keep the annual exit rehearsal and the per-domain build-versus-buy review on the calendar — the stage regresses quietly the year both are skipped.
Cost of leaving
- Effort
- Continuous
- Team
- Platform engineering owning the integration layer, a commercial owner running gates and rehearsals, and executive air-cover for re-competition
- Risk
- Concentrated at market events — consolidation, vendor acquisition and sunset — which is precisely when the stage pays
If this is you, the next step is
We run one exit rehearsal with you and price the walk-away for your next renewal gate.
The logistics AI vendor landscape: five archetypes
Suite modules, visibility networks, point solutions, freight marketplaces and platform tooling — what each actually sells, the lock-in mechanism each carries, and where each fits.
The logistics AI vendor market sorts into five archetypes, and each carries a different lock-in mechanism — which means each demands a different diligence emphasis. Readiness is not choosing the right archetype; every operator of scale ends up buying from most of them. It is knowing which questions each archetype must answer before signature, because a suite module, a visibility network and a freight marketplace fail buyers in entirely different ways.
| Archetype | What they actually sell | Systems they touch | Lock-in mechanism | Fits when |
|---|---|---|---|---|
| Incumbent suite AI — TMS/WMS/planning vendors adding modules | AI inside the system you already run: forecasting, optimisation and exception features priced into the suite | Their own TMS / WMS / planning estate | They already own your workflow and your data gravity; AI is bundled into the suite renewal | Decisions living wholly inside that suite, where integration cost dominates model quality |
| Visibility & ETA networks | Network-effect data products: predictive ETA, exception alerts, dwell and disruption signals trained across many shippers' flows | TMS, carrier EDI, telematics | The network's data accumulates on their side — your history improves a product you rent | Multi-carrier ETA and exception decisions no single operator could train alone |
| AI-native point solutions | One decision done deeply: slotting, dock scheduling, demand forecasting, freight audit, carrier vetting | WMS / YMS / TMS via API | Deep workflow embed and bespoke integrations that make replacement a project | A decision the suites do shallowly, worth a dedicated product and its integration bill |
| Digital freight networks & marketplaces | Matching, pricing and capacity as a service — the AI is their operation, not your tool | TMS tendering and settlement | Liquidity and rate history live with the marketplace; leaving means rebuilding relationships | Spot and backup capacity — benchmarked against independent rate data, never evaluated on quoted savings alone |
| Hyperscaler & platform tooling | Models, pipelines and infrastructure — build-adjacent capability, not logistics products | Your data platform | Cloud gravity: egress, proprietary services, accumulated pipeline code | Operators with engineering teams that own the decision layer and want vendors only underneath it |
Two archetypes deserve a specific caution each. Marketplace savings claims should never be evaluated on the marketplace's own numbers — independent rate benchmarks such as DAT's freight market data (opens in a new tab) exist precisely so a quoted saving can be tested against what the market actually paid on those lanes in that week. And the visibility networks' honest pitch — that products like Uber Freight's (opens in a new tab) matching or a visibility platform's ETA are trained across flows no single shipper could assemble — is also the lock-in to underwrite: your history is improving an asset you rent. The network effect is real value; the question your contract must answer is what happens to your data, and your operation, when you leave.
A third, quieter shift: your existing suppliers are becoming AI vendors whether you run a selection or not. Carriers ship digital products alongside capacity — Maersk's digital solutions portfolio (opens in a new tab) is a carrier selling software to its own shippers — and every suite upgrade now lands with AI features enabled. Vendor readiness therefore is not only for procurement events; it is the standing discipline of asking, of every product already inside the estate, the same questions a new vendor would face.
Where logistics operators sit on the buying ladder
The distribution across the five stages, and why checklist-led buying is the plateau.
Most logistics operators buy AI at stage 2: the procurement process is formal, but the evidence inside it still belongs to the seller. The distribution below is illustrative — synthesised from published adoption research rather than measured from a single survey — but its shape matches what the industry's own reporting keeps finding: adoption intent and contract volume far ahead of production evidence and governance.
Distribution of logistics operators across the five buying stages
Stage 2 is the mode: formal procurement, seller-owned evidence. The sharpest capability jump on the ladder — and the biggest commercial payoff — is the move to stage 3, when trials shift onto the operator's own lanes and terms move to signature.
Share of operators (illustrative)
- 24% — 1 · Demo-led
- 37% — 2 · Checklist-led (the plateau)
- 26% — 3 · Evidence-led
- 10% — 4 · Production-governed
- 3% — 5 · Portfolio-orchestrated
Source: Illustrative distribution, synthesised from MHI, Gartner and McKinsey adoption research
The external evidence for the plateau is consistent. MHI's Annual Industry Report (opens in a new tab) has tracked, year over year, a wide gap between the share of supply-chain organisations planning or piloting AI and the share running it in production — a gap that is precisely the stage-2 signature, since demo-led and checklist-led purchases produce contracts without producing production decisions. McKinsey's operations research (opens in a new tab) reports early AI adopters in supply chains achieving roughly 15% lower logistics costs than slower peers — a prize that accrues to operators whose purchases reach production, which is a buying-capability outcome before it is a modelling one.
One implication is worth making explicit: the move from stage 2 to stage 3 is the highest-return step on the ladder, and it is procedural rather than technical. It requires no new platform, no data-science hires and no transformation programme — it requires the next selection to run as a paid two-vendor trial on your own lanes with the term sheet on the table. Operators repeatedly overestimate the cost of that step and underestimate what stage-2 buying is already costing them in shelfware and conceded terms.
The production-readiness scorecard: what to test before you sign
Six diligence areas, the question each must answer, the evidence to demand, and the walk-away signal — plus the two-axis map that tells you how hard to apply them.
Vendor due diligence in logistics AI comes down to six areas, and the discipline is demanding evidence rather than answers in every one of them. An RFP response is an answer; a spike, a trial score, a named reference running your carrier mix, or a contract clause is evidence. The scorecard below is the working checklist we use when sitting on the buyer's side of a selection — apply it in proportion to the stakes, using the matrix that follows.
| Diligence area | The question that matters | Evidence to demand | Walk-away signal |
|---|---|---|---|
| Integration proof | What does it cost, on both sides, to connect this to our TMS/WMS/YMS estate? | A two-week integration spike: one real feed in, one write-back out, effort logged | Refusal to spike; integration priced only after signature |
| Model evidence | Does it beat our baseline on our lanes — including the hard ones? | Paid PoV on your live feeds, scored on holdout lanes in operational units | Evaluation offered only on vendor sample data or 'reference customer results' |
| Data & improvement terms | Who owns our operational data, its exports, and the model value trained on it? | Contract language: ownership, export format and cadence, exit assistance, improvement rights | 'Standard terms' that are silent on data — silence always resolves for the vendor |
| Operational support | What happens at 02:00 during peak when the model misbehaves? | SLA with response times, a named escalation path, and a tested way to disable the model | Support scoped to business hours in another timezone; no kill-switch story |
| Viability & continuity | Will this vendor exist, independent and motivated, for the life of the integration? | Funding condition, customer concentration, escrow or continuity terms proportionate to their size | Runway questions deflected; continuity 'never been asked before' |
| Security & assurance | Can they evidence their claims about controls, and the provenance of their data? | Third-party attestation (SOC 2 or equivalent; ISO/IEC 42001 for AI management is emerging), and named data sources | Certifications 'in progress' for years; training-data provenance they cannot state |
Two of these areas have logistics-specific depth worth naming. On security and assurance, ask where the vendor's data actually comes from: carrier-vetting and compliance products, for instance, are largely built on public registries such as FMCSA's safety and registration data (opens in a new tab) — knowing which parts of a product are proprietary intelligence and which are repackaged public data changes what the subscription is worth. On viability, remember that this market's consolidation is not hypothetical: point solutions get acquired by suites, marketplaces exit, and the diligence question is not 'might this happen?' but 'what does our contract and architecture do when it does?'
How hard to apply the scorecard: criticality against switching cost
Plot each vendor — current or candidate — by how core its decision is to your network and how entangled its exit would be. The quadrant sets the governance weight it deserves; the emphasised quadrant is where renewals go wrong.
Strong position
- Core decision, portable vendor
- Keep your baseline alive and re-compete at renewal
- This is where the integration layer earns its keep
The lock-in trap
- Core decision, entangled exit
- Full scorecard, continuity terms, annual exit rehearsal
- Engineer the exit before the renewal, not during it
Commodity
- Peripheral and portable
- Buy on price, light-touch governance
- Do not spend trial budget here
Quiet debt
- Peripheral but entangled
- Contain: no scope growth without portability terms
- The overlooked quadrant where estates silently calcify
What the vendor market looks like in public
Three public reference points read against the buying ladder — the build extreme, the continuity shock, and the supplier-turned-vendor. None is an Atomic Loops engagement; each links to the operator's own published material.
The public record teaches the vendor-readiness lesson from three different directions. Amazon marks the far end of build-versus-buy and shows what full integration ownership costs and returns; the Convoy shutdown — and Flexport's acquisition of its technology — is the clearest continuity case study the industry has; and Maersk shows incumbent suppliers becoming AI vendors to their own customers. Read each against the ladder rather than as a template: none of these operators' positions is reachable by imitation.
Three reference points read against the ladder
Outcomes as reported in each operator's own published material. Verify figures against the linked source before reusing them; we have not independently audited them.
AmazonGlobal retail & logistics network · 1.5M+ employees45
- Challenge
- Operating fulfilment at a scale where no vendor's product roadmap could be allowed to set the ceiling on warehouse decision automation — and where integration depth, not model quality, is the binding constraint.
- Approach
- Vertical integration of the AI supply chain itself: Amazon builds and deploys its own robotics and AI systems in-house — including the robotic systems and AI-driven inventory-handling technologies covered in its operations newsroom — rather than assembling them from vendor products.
- Reported outcome
- Amazon has publicly reported deploying robotics at the scale of hundreds of thousands of mobile units across its network, alongside AI systems it states improve inventory handling and fulfilment speed — reported by Amazon's own operations newsroom.
- What it shows about the curveThe build pole exists, and it is priced in engineering headcount very few operators have. The transferable lesson is not 'build' — it is that Amazon owns its integration layer and decision data completely, which is exactly the asset a buying operator can own at any scale while still renting the models.
Flexport / ConvoyGlobal freight forwarder · logistics technology platform34
- Challenge
- Convoy, a digital freight network that had raised $260M at a $3.8B valuation in April 2022, ceased operations in October 2023 — roughly 18 months later — leaving shippers and carriers that depended on its AI-driven matching with weeks to re-route freight.
- Approach
- Flexport acquired Convoy's technology stack in November 2023, as publicly announced at the time, and later relaunched the platform's capabilities within its own product family — preserving the technology while the vendor entity itself disappeared.
- Reported outcome
- As publicly reported: the market's freight kept moving, but the transition cost landed on buyers in proportion to their exit readiness — those with documented fallbacks and portable tender flows migrated in weeks; the technology's survival inside Flexport did not spare anyone the migration.
- What it shows about the curveVendor continuity is a buyer-side readiness dimension. A well-funded, well-regarded vendor exited the market inside 18 months; nothing in a demo, an RFP score or a reference call would have predicted it. Contract continuity terms and a drilled exit are the only instruments that pay on the day it happens.
MaerskGlobal container carrier & logistics integrator34
- Challenge
- A capacity supplier competing in a digitising market where its shippers increasingly buy visibility, planning and exception-handling capability as software — from third parties, unless the carrier offers it.
- Approach
- Maersk built a digital-solutions portfolio — visibility, logistics management and data products marketed directly to its shipping customers — positioning the carrier itself as a technology vendor on top of its physical network.
- Reported outcome
- Maersk publicly markets an expanding portfolio of digital and data products to its customers via its digital-solutions platform — the carrier's own published material is the source for its positioning and capabilities.
- What it shows about the curveFor buyers, the supplier-turned-vendor changes the diligence, not the discipline: capability bundled with capacity still deserves the same scorecard — baseline, integration proof, data terms — plus one extra question no standalone vendor raises: what does this data relationship do to our negotiating position on the freight itself?