Construction & InfrastructureReadiness & Transformation Roadmap
AI vendor selection for infrastructure: how owners and contractors evaluate, contract and govern AI suppliers
Infrastructure AI vendor selection is the procurement discipline of choosing, contracting and governing the suppliers that put AI into a capital programme. It differs from software buying in four places: training-data provenance, ownership of the model your data enriched, a performance floor that can drift, and an exit that returns something you can actually use.

Key takeaways
- Infrastructure AI procurement fails on terms, not on technology. Four clauses decide the outcome — training-data provenance, ownership of the enriched model data, a performance floor with named drift responsibility, and an exit that returns usable artefacts — and a standard software contract contains none of them.
- The enriched model is the asset, not the software. A vendor arrives with a generic detector and leaves with one tuned on your as-built geometry, your defect history and your RFI corpus. If the contract does not name who owns that, you will rent your own project record back at renewal.
- Public owners cannot extend a successful pilot into production on goodwill. The route to production has to be designed into the notice that launched the pilot — an open framework, a dynamic market or a competitive flexible procedure — or the pilot's success triggers a fresh procurement and a year of delay.
- Vendor viability is a technical requirement, not a finance formality. Escrow that covers only source code is worthless for AI: the release package must include model weights, training and inference code, feature definitions and the evaluation harness, and the release must have been tested at least once.
- Portfolio governance starts at vendor four, not vendor ten. By the fourth AI supplier an owner typically has overlapping capability, four copies of the same site imagery leaving the CDE, and four renewal dates nobody is tracking — all of which are cheaper to prevent than to unwind.
Abbreviations used on this page
- BIM
- Building information modelling
- CDE
- Common data environment — the ISO 19650 project information store
- EIR
- Exchange information requirements — the brief a supplier answers under ISO 19650
- AIR
- Asset information requirements — what the operating estate needs handed over
- IFC
- Industry Foundation Classes — the buildingSMART open model exchange format
- ITT
- Invitation to tender
- RFP
- Request for proposal
- CFP
- Competitive flexible procedure — the buyer-designed route under the Procurement Act 2023
- FTS
- Find a Tender Service — the UK public procurement notice platform
- SLA
- Service level agreement
- AIMS
- AI management system — the ISO/IEC 42001 construct
- TCO
- Total cost of ownership
Free · 8 questions · ~3 minutes
Score how you buy AI
Eight questions, one at a time, about three minutes. Answer them and we build your personalised vendor readiness report — where you sit on the buying ladder, your score on each of the four dimensions, and the specific clause or process gap standing between you and the next rung — and send it to your inbox. Your result doubles as the agenda for your next AI tender.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised report is ready
Tell us where to send it. Your rung appears on screen straight away, and the full report — dimension scores, the clause gaps that matter most in your position, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Ad hoc buying
AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
Your next moveWrite one requirement document and one set of standard AI terms, and route every new AI purchase through them — even the £5,000 ones.
Stage 2 · Structured requirements
AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
Your next moveHold back a real package of your own data, score every shortlisted bidder on the same harness on the same day, and sign the model and data terms before the pilot starts.
Stage 3 · Piloted with terms
Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
Your next movePut a priced, time-boxed option to extend into production in the pilot contract, agreed at award, with the unit rates and volume mechanism already in the schedule.
Stage 4 · Contracted for production
AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
Your next moveStand up a vendor and model register with concentration limits, an overlap review and a single renewal calendar across every AI contract in the business.
Stage 5 · Portfolio-managed
AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.
Your next movePut the register, the harness and the renewal calendar on a fixed review cadence with a named owner, and measure how long it takes to score a challenger.
0 / 24
Requirements clarity
— / 6
Data & IP terms
— / 6
Pilot-to-production path
— / 6
Vendor portfolio governance
— / 6
Your score maps to a rung on the buying ladder. The dimension breakdown matters more than the total: the lowest dimension is the one that will decide your next AI contract, and it is almost never the one the team expects. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a rung on the buying ladder. The dimension breakdown matters more than the total: the lowest dimension is the one that will decide your next AI contract, and it is almost never the one the team expects.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want this checked against a live tender?
We will walk your commercial, information-management and delivery leads through the dimension scores, mark up a live or upcoming AI ITT against the clause table on this page, and leave you with the redlines that matter most. No obligation, and you keep the markup either way.
How the score maps to a stage
- 0–4 — Stage 1, Ad hoc buying. AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
- 5–10 — Stage 2, Structured requirements. AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
- 11–15 — Stage 3, Piloted with terms. Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
- 16–20 — Stage 4, Contracted for production. AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
- 21–24 — Stage 5, Portfolio-managed. AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.
What infrastructure AI vendor selection is — and why it is not software buying
A definition, the four places an AI purchase differs from a software purchase, and the path a requirement travels to become a production contract.
Infrastructure AI vendor selection is the procurement discipline of choosing, contracting and governing the suppliers that put AI into a capital programme or an operating estate. It covers the make-or-buy decision, the evaluation method, the data and intellectual property terms, the route to market that public procurement rules allow, the continuity provisions that survive a supplier failing, and the portfolio governance that starts once an owner is running more than three or four AI contracts at once.
It is not software buying with a different noun in the title. Four things differ, and each of them is a clause a standard software contract does not contain. The first is training-data provenance: what the model you are buying was built from, and whether your project information will be used to build the next one. The second is ownership of the enriched model data — the labels, the corrected geometry, the classified elements and the extracted quantities that only exist because your scheme paid for them. The third is that performance degrades on its own, so a system that met its specification at acceptance may not meet it in year three, and somebody has to be contractually responsible for that. The fourth is exit: a software exit returns your records, an AI exit has to return something you can keep operating with.
Value the owner retains, against maturity on the buying ladder
The curve is not linear, and the inflection is commercial rather than technical. Value retained by the owner stays close to flat through ad hoc buying and structured requirements — where most owners are — and rises sharply once terms are signed before the pilot rather than after it, because that is the point at which the enriched model and the data it was built from stop leaving the business by default.
Value retained by the owner by stage
- Stage 1 · Ad hoc buying — 24% of operators. AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
- Stage 2 · Structured requirements — 34% of operators. AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
- Stage 3 · Piloted with terms — 23% of operators. Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
- Stage 4 · Contracted for production — 14% of operators. AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
- Stage 5 · Portfolio-managed — 5% of operators. AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the UK Government's Guidelines for AI procurement.
How an AI requirement becomes a production contract
The commercial path, rung by rung. The rung is determined by where the arrow ends: ad hoc buying terminates in an auto-renewal nobody owns, structured piloting terminates in a gate review, and only a route named at the notice stage reaches a production call-off and a portfolio register. Most owners are in the top lane.
- Where value leaks
- Data & feeds
- AI / model
- System-of-record action
- Human in the loop
The process, in words
- At rungs 1–2, an AI capability enters through a demonstration on the supplier's own data, becomes a free proof of concept running on real project information, and is regularised with a card invoice under the supplier's standard terms. It then auto-renews with no named owner, no re-test and no benchmark. This is where ownership of the enriched model quietly leaves the business.
- At rung 3, the requirement is extracted from the exchange and asset information requirements rather than from a feature list. A held-back package of your own data is withheld from every bidder, all shortlisted suppliers run the same scoring harness on the same day, and the pilot only starts once provenance, ownership, the performance floor and the exit terms are signed. The gate review at the end is a commercial decision, not a technical one.
- At rungs 4–5, the production route was named in the notice that launched the pilot, so exercising a priced option is a purchase order rather than a fresh competition. Viability diligence runs in parallel with the pilot, and every awarded supplier lands on a portfolio register that tracks what data it holds, what floor it must meet and when it is next re-benchmarked.
Step-by-step insights
- The vendor demonstration — why every bidder looks the same
- A demonstration is a supplier's best asset shown on their best data, usually a scheme whose geometry, labelling conventions and lighting the model has already been tuned against. That is not deception; it is what a demonstration is for. The consequence for a buyer is that all credible bidders present indistinguishably strong results, so the evaluation panel is forced onto criteria it can differentiate — price, references, cultural fit — none of which predicts performance on your night-survey footage in February. The single highest-leverage change in infrastructure AI procurement is refusing to be shown anything, and instead handing every bidder the same unfamiliar package.
- The free proof of concept — the most expensive free thing in the estate
- A no-cost trial requires no purchase order, so it requires no terms review, and it typically runs on genuine project data because synthetic data would not prove anything. Two assets transfer during it: your project information goes out, and a model tuned to your conventions comes back — owned by the supplier under an agreement nobody read. When the trial converts, it converts into a negotiation where the supplier already holds the tuned model and the site teams have already changed how they work. Charge for the pilot if it helps; the point is to have a contract, not a price.
- EIR and AIR as the source of the requirement
- Under ISO 19650 the exchange information requirements already state what information the project needs, in what form, at which stage, and to what level of information need. An AI requirement written from that document is testable: named information containers, a defined delivery point, a measurable acceptance metric. An AI requirement written from a supplier's feature list is a wish list of capabilities that every bidder can claim. Writing the requirement from the EIR also has a procedural benefit — it keeps the specification supplier-neutral, which matters a great deal if the competition is ever challenged.
- The held-back package and the one-day bake-off
- The held-back package is a genuine slice of your work — a fortnight of survey imagery, a set of as-built information containers, a tranche of RFI text — deliberately including the awkward cases: poor light, a non-standard asset class, a package modelled by a different designer. Nobody sees it before the day. Every shortlisted bidder runs the same harness, on the same data, in the same window, and the harness was written before any bid arrived. The scores that come back are frequently half what the demonstrations implied, and the ranking is frequently different from the panel's expectation.
- Viability diligence running alongside the pilot
- Technical evaluation and financial diligence usually run in series, which wastes the pilot period. Run them together: filed accounts, funding position, customer concentration, the escrow package's actual contents, change-of-control provisions and where the model would run if the supplier disappeared. This is not a credit check. The question is narrower and more useful — if this company is acquired in eighteen months, what do we hold, and can we run it? An answer of weights, inference code, feature definitions and a tested release is a different risk profile from an answer of a source-code deposit nobody has opened.
- The route named at the notice — the clause that saves a year
- For public owners this is the difference between a pilot that scales and a pilot that expires. If the notice that launched the pilot was a one-off, low-value engagement, a successful result cannot lawfully be extended into a production contract of any size, and the reward for a good pilot is a fresh competition. If instead the pilot was run inside an open framework, a dynamic market or a competitive flexible procedure that anticipated a production phase, the extension is a call-off. The decision is made months before anyone knows whether the technology works, which is precisely why it is so often made wrongly.
The rest of this page is organised around that path. The buying ladder and its five rungs come first, then where owners actually sit, then the build-or-buy decision, then the two documents that do most of the work: a weighted evaluation matrix and a clause-by-clause comparison of what an AI contract must add to a software one. The UK Government's Guidelines for AI procurement (opens in a new tab) and the AI Playbook for the UK Government (opens in a new tab) are the closest public-sector equivalents and are worth reading alongside it.
The five rungs of the buying ladder in detail
For each rung: what it actually looks like in a commercial team, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps owners there, and what leaving costs.
Each rung below is written for a commercial lead and an information manager rather than for a buyer of software. The hallmarks describe observable conditions in your contract file and your CDE access list, the diagnostic signals are checks you can run this week against documents you already hold, and the anti-pattern is the specific mistake most often made trying to leave that rung.
Select a rung
Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Ad hoc buying
24% of operators sit here
AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
Stage 1 is not the absence of AI — most contractors at this stage have more AI in the business than their leadership believes. It is the absence of a buying process, so the capability arrives one project at a time, through whichever supplier ran the most convincing demonstration at whichever site had budget that month. The tools frequently work. What is missing is any record of what was agreed.
The tell is the contract file. At stage 1 the governing document for an AI supplier that reads your drawings, your site imagery and your programme is a click-through agreement nobody in the business has opened, containing a broad licence to use customer content to improve the services. That clause is not unusual and it is not hidden; it is simply never read, because the purchase never passed a desk that reads clauses.
The cost of staying here is not the subscription spend, which is usually trivial. It is that every project's evaluation work amortises nothing — the tenth trial is judged exactly as badly as the first — and that the information leaving your common data environment is uncounted. When someone eventually asks which suppliers hold copies of a scheme's model data, the honest answer takes a fortnight to assemble.
In practice
The site that bought its own computer vision
A project director on a station refurbishment saw a progress-capture demo at a conference, signed up for a three-month trial on a departmental card, and connected the supplier's app to the project's photo library and the shared model. The trial was genuinely useful and the team extended it twice. Eighteen months later the commercial team, preparing a claim, discovered that the definitive weekly progress record for a disputed period sat in a supplier's cloud under terms that gave the project a licence rather than ownership, and no export format better than PDF.
What it looks like
- AI tools enter through a project team's card, not through procurement
- Evaluation is a demo on the vendor's data, judged by impression
- Terms are the vendor's standard online agreement, unread
- Nobody can list which AI suppliers currently touch project information
Diagnostic signals you can check this week
- Ask for a list of AI suppliers that hold project information. If the answer requires an expenses audit, you are here
- Open the terms of the last AI tool a project adopted and search for the words train, improve and aggregate
- Ask who evaluated the tool, and against what written criteria. Usually one person, and none
- Check whether any AI subscription in the business has a named renewal owner
Anti-pattern · Banning the tools instead of buying them properly
The instinctive reaction to discovering uncontrolled AI is a blanket prohibition and a mandatory approval queue. It removes the visible tools and keeps none of the value: teams either revert to the manual process the tool replaced or move the same activity to a personal account, where you now have the same data exposure with none of the logging. The durable fix is a fast, published route to buy — a shortlist, a standard set of terms and a decision inside two weeks — so the compliant path is also the quickest one.
What holds you here
There is no written requirement and no standard term set, so every purchase is judged on a demonstration and governed by the supplier's own paper.
Highest-leverage next move
Write one requirement document and one set of standard AI terms, and route every new AI purchase through them — even the £5,000 ones.
Cost of leaving
- Effort
- 1–3 months
- Team
- One commercial lead and one information manager, part-time
- Risk
- Low — the work is a register, a standard term set and a published route
- To next stage
- 1–3 months
If this is you, the next step is
A two-week exercise: who holds your project data, under what terms, and what to renegotiate first.
Stage 2
Structured requirements
34% of operators sit here
AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
Stage 2 is where most infrastructure owners and tier-one contractors are, and it looks like good practice because it is good practice — for buying software. A requirement is written, a shortlist is scored, a pilot is funded and a business case is prepared. Every step is defensible. The problem is that the artefacts being bought are not software artefacts, and the process never asks the questions that separate one AI supplier from another.
Two omissions do most of the damage. The first is that evaluation happens on the vendor's own data: every bidder demonstrates on a scheme they have already tuned for, so every demonstration looks equally impressive and the score reflects presentation quality. The second is that the pilot is scoped as a proof of value and nothing else, so a successful pilot produces a slide deck, an enthusiastic sponsor and no contractual path to buying the thing at scale.
Time at stage 2 is expensive in a specific way. Each pilot enriches a supplier's model with your project information under terms you did not draft, and each ends without converting, so the supplier accumulates the durable asset and you accumulate the evaluations. Owners who have run four or five pilots in this pattern usually find their negotiating position has quietly weakened, because the incumbent now knows their data better than any challenger can.
In practice
The bake-off that everybody won
A highways client shortlisted four suppliers for automated defect detection and asked each to present. All four demonstrated on their own reference imagery, all four reported recall in the nineties, and the evaluation panel could not separate them on capability, so the award went on price and cultural fit. The winning system, run on the client's own night-survey footage in variable weather, performed far below its demonstration and the difference was not discoverable from anything in the tender.
What it looks like
- A written requirement exists and bidders are scored against it
- Evaluation still happens on vendor-supplied demonstration data
- The contract is the vendor's paper with a few negotiated redlines
- Pilots are funded as one-off proofs of value with no production option
Diagnostic signals you can check this week
- Ask whether any bidder has ever been scored on data they had not seen before
- Look for the words performance floor or acceptance test in the last AI contract signed. Usually absent
- Check whether the last pilot's contract contained a priced option to extend into production
- Ask who owns the labels created during the last pilot. If the answer is a shrug, the vendor owns them
Anti-pattern · Writing a longer requirement instead of a harder test
When the shortlist proves indistinguishable, the reflex is to add requirements — more questions, more compliance schedules, more weighting categories. It makes the document heavier and the decision no better, because every bidder can answer every question affirmatively and no answer is testable. The step that actually separates suppliers is small and unpopular: hold back one real package of your own data, write one scoring harness, and make every shortlisted bidder run on it on the same day.
What holds you here
Bidders are still judged on their own demonstration data and the pilot has no contractual route into production, so a good result cannot be bought at scale.
Highest-leverage next move
Hold back a real package of your own data, score every shortlisted bidder on the same harness on the same day, and sign the model and data terms before the pilot starts.
Cost of leaving
- Effort
- 3–6 months
- Team
- Commercial lead, information manager, one technical assessor, a named sponsor
- Risk
- Medium — the held-back package has to be prepared carefully and lawfully
- To next stage
- 3–6 months
If this is you, the next step is
We prepare the held-back package and the scoring script; you run the bake-off.
Stage 3
Piloted with terms
23% of operators sit here
Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
Stage 3 is the first stage where a pilot is a purchase decision rather than an experiment. The change is not in how the pilot runs but in what has been agreed before it starts: who owns the labels, who owns the tuned model, what the system must achieve in your units, who pays when it stops achieving it, and what you get back if the arrangement ends. Those are one workshop and one schedule, and they are the difference between an option and a sunk cost.
The evaluation discipline is equally concrete. A held-back package — a fortnight of survey imagery, a set of as-built information containers, a slice of the RFI record — is withheld from every bidder, and the scoring harness is written before anyone bids. Scores stop reflecting presentation quality and start reflecting performance on the geometry, the weather and the labelling conventions your business actually produces. Suppliers who object to blind evaluation are giving you information.
What emerges at stage 3 is a different negotiation. Because the terms are drafted before the pilot rather than after it, the leverage sits with the buyer at the moment leverage exists. Owners consistently report that clauses which are impossible to win after a successful pilot — enriched-data ownership, escrow scope, exit pricing — are routine to win before one, because at that point the supplier is competing rather than incumbent.
In practice
The fortnight nobody had seen
A rail infrastructure owner shortlisted five suppliers for automated asset-condition scoring and withheld two weeks of tunnel inspection footage, deliberately including one shift of poor lighting and one of a non-standard asset class. Every bidder ran the same harness on the same Tuesday. Reported recall across the five ranged from acceptable to less than half of what their marketing claimed, and the two suppliers whose numbers held up were not the two the panel had expected from the presentations.
What it looks like
- Shortlisted bidders are scored blind on a held-back package of your data
- Model and data schedules are signed before any pilot begins
- A performance floor is stated in your units, with the measurement method
- A gate review at the end of the pilot has a commercial decision attached to it
Diagnostic signals you can check this week
- Ask to see the scoring harness used on the last AI award. It should exist as a file, not a spreadsheet of opinions
- Check whether the last pilot contract defined enriched model data as customer data
- Ask what the performance floor is for a live AI system, in your units. A number should come back
- Check whether the pilot gate review had a commercial decision attached, or only a technical one
Anti-pattern · Treating the pilot as the procurement
A pilot run well under good terms feels like the hard part is over, so the production purchase is left until the pilot has proved itself. That sequencing hands the supplier the whole negotiation: by the time the pilot has succeeded, the model is tuned on your data, the site teams have changed how they work, and the alternative is a twelve-month restart. Price the production option at award, when three other bidders are still on the table, even if you never exercise it.
What holds you here
The pilot converts on goodwill rather than on a priced, pre-agreed option, so scaling reopens the commercial negotiation from a weaker position.
Highest-leverage next move
Put a priced, time-boxed option to extend into production in the pilot contract, agreed at award, with the unit rates and volume mechanism already in the schedule.
Cost of leaving
- Effort
- 6–12 months
- Team
- Commercial and legal lead, information manager, technical assessor, project sponsor
- Risk
- Medium — the drafting effort is front-loaded and competes with delivery pressure
- To next stage
- 6–12 months
If this is you, the next step is
The four clauses that decide the outcome, written for your CDE and your standard form of contract.
Stage 4
Contracted for production
14% of operators sit here
AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
Stage 4 is where an AI supplier becomes part of the estate rather than part of a project. The system is contracted for a term, with volume mechanics that survive a scheme finishing and another starting, and with the annual re-test that stops a five-year framework quietly becoming a five-year subscription to a 2026 model. The engineering was largely settled at stage 3; what changes here is that the commercial arrangement is built to outlast the people who negotiated it.
Continuity work is what distinguishes stage 4 from a well-run stage 3. AI suppliers in this sector are young, frequently venture-funded, and acquired at a rate that would be unremarkable in software and is alarming when the acquired product holds the definitive condition record for a tunnel. Escrow that covers only source code releases you a repository you cannot run; the release package has to include the weights, the training and inference code, the feature definitions and the evaluation harness, and someone has to have tried releasing it once.
The other stage-4 discipline is unglamorous: the renewal calendar. A production contract creates dates — floor re-test, benchmark, break, exit notice — and those dates only work if a named person owns them. Owners who reach stage 4 without a register find the contract's protections lapse silently, and discover it at the renewal where the price rises and the alternative has not been tested for three years.
In practice
The option that was already priced
An infrastructure client ran a six-month pilot of automated quantity extraction across two packages, with a production option priced at award covering a further four packages at fixed unit rates. The pilot succeeded. Exercising the option took a gate paper and a purchase order — eleven days from decision to live. The comparable client on the same corridor, whose pilot had no option, spent nine months running a new competition and awarded to the same supplier at a higher rate.
What it looks like
- The production option was priced at award and exercised without a new competition
- Escrow covers weights, training and inference code, features and the harness — and has been tested
- The floor is re-tested annually against a fresh held-back package
- Change of control, sub-processing and exit pricing are all fixed at signature
Diagnostic signals you can check this week
- Ask when the escrow release was last tested. If never, you have a document, not a continuity plan
- Check whether the performance floor has been re-tested since award, and on what data
- Look for a change-of-control clause in the largest AI contract in the business
- Ask who owns the renewal date. If it is a calendar reminder on one person's laptop, it is unowned
Anti-pattern · Buying continuity as a document rather than a drill
Escrow, exit plans and transition assistance are easy to agree because they cost nothing at signature and everything at release. The failure mode is uniform: the agreement names a release event, nobody ever exercises it, and when a supplier is acquired the release turns out to omit the training data pipeline, the feature definitions or the environment the weights run in. Run the release once, in a quiet quarter, and treat the gaps you find as contract defects rather than as an inconvenience.
What holds you here
Each supplier is contracted well in isolation, so overlapping capability, duplicated data egress and uncoordinated renewals accumulate across the portfolio.
Highest-leverage next move
Stand up a vendor and model register with concentration limits, an overlap review and a single renewal calendar across every AI contract in the business.
Cost of leaving
- Effort
- 12–18 months
- Team
- Category owner, legal, information manager, an engineering owner for the floor re-test
- Risk
- Higher — continuity and exit obligations have to be real, and testing them costs supplier goodwill
- To next stage
- 12–18 months
If this is you, the next step is
We run the escrow release against a real vendor package and report what is actually missing.
Stage 5
Portfolio-managed
5% of operators sit here
AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.
Stage 5 is narrower than it sounds, and it is a commercial capability rather than a technical one. The register is unremarkable: supplier, model, what data it holds, which schemes it touches, the floor, the renewal date, the exit price. What makes it stage 5 is that the register is consulted before a new purchase, so the sixth AI supplier has to demonstrate that it does something the previous five do not — which is a question stage-4 organisations never get asked.
The concentration question is the one owners reach late. A single supplier holding progress capture, condition scoring and quantity extraction across a whole region is efficient right up to the moment it is acquired, changes its pricing model or suffers an outage during a claims window. Stage-5 owners set an explicit ceiling on how much of the estate one supplier may hold, and accept the small integration cost of keeping a credible second source warm.
Sustaining stage 5 is mostly calendar discipline, and it is the stage most likely to regress. The register goes stale when a scheme closes, the harness rots when the data conventions change, and a year of no re-tests turns a portfolio back into a collection of contracts. The signal to watch is how long it takes to score a challenger: when that number starts rising, the shared harness has stopped being shared.
In practice
The sixth supplier that was refused
A national owner with five AI suppliers on its register was offered a sixth for automated drawing comparison. The register showed two incumbents already extracting the same information containers from the CDE, and the shared harness let the team score the challenger against both in nine days. The challenger was better on one asset class and worse on two, so instead of a sixth contract the owner added the asset class to an existing call-off at the framework's rates.
What it looks like
- One register holds every AI supplier, model, data flow and renewal date
- New capability is tested against the incumbent portfolio before a new supplier is added
- Concentration limits cap how much of the estate any single supplier can hold
- A shared evaluation harness is reused across categories, so a challenger can be scored in days
Diagnostic signals you can check this week
- Ask how long it takes to score a new AI supplier against an incumbent. Under two weeks means the harness is real
- Check whether the register records what data each supplier holds, not just what it costs
- Ask whether any purchase in the last year was refused because an incumbent already covered it
- Check whether a concentration limit exists as a written rule with a named approver for exceptions
Anti-pattern · Consolidating to one supplier because the register is tidy
A good portfolio view makes rationalisation tempting, and consolidating five suppliers into one platform genuinely simplifies integration, invoicing and governance. It also puts the definitive record for the whole estate behind one commercial relationship, one pricing model and one balance sheet, at which point the renewal negotiation has no alternative in it. Rationalise overlap, keep a second source warm in every category that touches the asset record, and treat the integration cost of doing so as insurance rather than waste.
What holds you here
Portfolio discipline decays quietly — registers go stale, harnesses rot, and re-tests slip until the portfolio is a collection of contracts again.
Highest-leverage next move
Put the register, the harness and the renewal calendar on a fixed review cadence with a named owner, and measure how long it takes to score a challenger.
Cost of leaving
- Effort
- Continuous
- Team
- A named AI category owner plus a standing commercial and information-management forum
- Risk
- Concentrated — low frequency, high consequence, and commercial rather than technical in nature
If this is you, the next step is
We map overlap, concentration and renewal exposure across every AI supplier you run.
Where infrastructure owners and contractors actually sit
The distribution across the ladder, and why the drop between structured requirements and piloting with terms is the largest single loss.
Most infrastructure owners and contractors sit at structured requirements — rung two. They write a requirement, score a shortlist and fund a pilot, and they do all of it on the supplier's paper and the supplier's data. A much smaller group signs the model and data terms before the pilot starts, and a very small group has ever exercised a production option that was priced at award.
Illustrative distribution of infrastructure owners and contractors across the buying ladder
Structured requirements is the mode and the plateau. The drop from rung 2 to rung 3 is the largest single transition loss on the ladder, and it is a drafting problem rather than a technology problem. Figures are illustrative, synthesised from the public procurement guidance and construction-productivity research linked beneath the chart, not a measured survey.
Share of owners and contractors
- 24% — 1 · Ad hoc buying
- 34% — 2 · Structured requirements (the plateau)
- 23% — 3 · Piloted with terms
- 14% — 4 · Contracted for production
- 5% — 5 · Portfolio-managed
The plateau has a structural cause. A construction or infrastructure business already knows how to buy software and already knows how to buy design and construction services, and an AI supplier looks like the first while behaving like the second: the deliverable is partly built out of the client's own information, and its quality is contingent on that information. Buying it with a software process produces a contract that is silent on the three things that matter — provenance, the enriched model, and degradation over time. McKinsey's construction-productivity research (opens in a new tab) sets out how little the sector's productivity has moved over two decades, which is the gap every AI supplier is sold against and the reason these purchases get waved through on urgency.
The second cause is that the people who can fix it are rarely in the room together. Provenance and enriched-data ownership are information-management questions, the performance floor is an engineering question, escrow and change of control are legal questions, and the production option is a procurement question. Owners who move up the ladder almost always do it by putting those four people in one workshop before the ITT is issued rather than after the pilot succeeds.
Build, buy or partner: deciding before you go to market
The make-or-buy call turns on two variables — how unusual your data is, and how many credible suppliers exist — and it decides everything downstream about the terms you need.
Build, buy or partner is decided by two questions and not by capability: how unusual is the data the capability depends on, and how many credible suppliers already sell it? Where the data is industry-standard and the market is mature — object detection on site imagery, document search over a contract set — buy, and re-tender often, because the underlying capability improves faster than any owner can track. Where the data is your own accumulated programme record — outturn cost against estimate, historical change and claims, your own defect taxonomy — the asset is the data, no supplier has it, and the correct answer is to build or to partner with the model kept on your side of the line.
| Capability | Default answer | Why | Who must own the data asset |
|---|---|---|---|
| Progress capture from site imagery, drone and laser scan | Buy | Commodity computer vision with many credible suppliers; the differentiator is your as-built model and your labelling conventions, not the detector | You own the imagery, the labels and the classified elements |
| Quantity take-off and model enrichment | Buy, own the output | The extraction service is bought; the enriched information containers are the asset and must land back in the CDE in an open format | You own the enriched containers and the extracted quantities |
| Programme and schedule risk prediction | Partner | Depends on your historical programmes, change history and claims record — data no supplier holds and none can substitute for | You own the trained model and the feature definitions |
| Cost and estimate benchmarking | Build or partner | Your rate history, subcontract packages and outturn costs are the model; sharing them with a supplier who serves competitors is a commercial decision, not a technical one | You own the model and every benchmark derived from it |
| Asset condition and defect detection on the operating estate | Buy, with a floor | Mature market on standard asset classes; contract a recall floor on your classes and your survey conditions rather than a generic accuracy claim | You own the labelled condition dataset and the survey record |
| Document, RFI and contract review assistants | Buy, re-tender often | General language capability moves faster than any owner can build against; assume a two-to-three-year useful life for the supplier choice | You own the retrieval corpus and the full audit log |
| Safety observation and behavioural monitoring | Buy, govern hard | The model is straightforward; the worker-identifiable data is not, and the governance load exceeds the technical load | You own the imagery, the retention schedule and the DPIA |
| Generative design and optioneering | Partner with the designer | Design liability sits with the designer, not with the tool; the capability must live inside an existing design duty rather than beside it | The designer owns the outputs; you own the information deliverables |
The make-or-buy quadrant
Plot the uniqueness of the data the capability depends on against the number of credible suppliers. Three of the four quadrants have an answer that is not build, and the most valuable quadrant is the one owners most often mishandle.
Build or co-develop
- Your data is the capability and nobody sells it
- Cost benchmarking, claims prediction, your own defect taxonomy
- Terms to win: model ownership, feature definitions, no shared training
Buy the engine, own the model
- Mature engines, but tuned on information only you hold
- The highest-value and most mishandled quadrant
- Terms to win: enriched-data ownership, escrowed weights, portability
Wait, or open a dynamic market
- Nothing credible to buy and nothing distinctive to build on
- Common for genuinely novel capability
- Move: a dynamic market or pre-market engagement, not a tender
Buy and re-tender often
- Commodity capability, commodity data, many suppliers
- Document assistants, standard object detection
- Terms to win: short term, low exit cost, open export formats
The quadrant owners handle worst is the top right: a mature engine tuned on information only you hold. It looks like a straightforward purchase because the engine is a product with a price list, and it behaves like a joint development because the version that works on your estate exists nowhere else. Buying it on standard software terms transfers the only genuinely scarce asset in the transaction to the supplier for free, which is why the evaluation matrix in the next section but one puts twenty per cent of the score on performance against your own held-back data and fifteen on who ends up owning the model that produced it.
The vendor evaluation matrix: criteria, weights and disqualifiers
Nine criteria, a weight for each, the evidence to demand at tender stage, and the answer that should end a bid on the spot.
An AI vendor evaluation matrix differs from a software one in where the weight sits: twenty per cent on measured performance against data the bidder has never seen, and thirty per cent across provenance and ownership of what the engagement produces. The matrix below is the version we use with infrastructure owners and tier-one contractors. The weights are a defensible starting point rather than a standard — move them, but move them before the ITT is issued and record why, because a weighting changed after bids arrive is the most common way an award becomes challengeable.
| Criterion | Weight | Evidence to demand at ITT | Automatic disqualifier |
|---|---|---|---|
| Measured performance on your held-back data | 20% | A blind run on a package you withheld — your survey imagery, your as-built containers, your RFI text — scored on one harness you wrote, on one day, with every shortlisted bidder present | Refuses to be evaluated on anything but their own demonstration data |
| Training-data provenance | 15% | A written data lineage statement: what the base model was trained on, on what licence basis, whether any competitor's project information sits in it, and whether your data will train anything shared | Cannot state where the training data came from, or reserves a right to train shared models on your project information |
| Ownership of the enriched model and derived data | 15% | Draft clause text naming who owns the fine-tuned weights, the label set, the embeddings, the corrected geometry and the extracted quantities at the end of the term | Claims ownership of data derived from your CDE, or offers only a licence back to your own information |
| Exit and data portability | 10% | A written exit plan: export formats (IFC, COBie, open label formats, model weights), the delivery window, transition assistance days, and a price fixed at signature | Exit priced at the point of exit, or export available only in a proprietary format |
| Performance floor and drift responsibility | 10% | A floor stated in your units, the measurement method and dataset, who detects drift, who pays for retraining, in what window, and the remedy if the floor is not restored | Performance stated only as accuracy on the supplier's own benchmark, with no floor in your units |
| Integration with the CDE and the ISO 19650 workflow | 10% | A demonstrated read and write against your common data environment with information containers, revision codes and suitability states preserved end to end | Integration by manual export, or a full copy of the CDE held permanently on the supplier's side |
| Vendor viability and continuity | 10% | Filed accounts, funding position, customer concentration, the actual contents of the escrow package, change-of-control provisions and evidence a release has been tested | No escrow, no change-of-control clause, and a refusal to discuss financial runway |
| Security, worker data and site imagery | 5% | DPIA support, a retention schedule for site imagery, worker-identifiability controls, the full sub-processor list and the hosting jurisdiction | Site imagery retained indefinitely, or sub-processors undisclosed at tender stage |
| Five-year total cost of ownership | 5% | Licence, per-seat and per-asset scaling, data egress, retraining, integration maintenance and the cost of the exit, modelled across five years | Year-one price only, with the scaling mechanism left undefined |
Score the evidence, not the answer
Every bidder will answer yes to every question. The score belongs to the artefact attached to the answer: the lineage statement, the draft clause, the exit plan with a price on it. A criterion with no required artefact is a criterion that cannot discriminate, and it should either be given an artefact or removed from the matrix.
Write the harness before the bids arrive
The scoring harness — the metric, the threshold, the data, the script — is written and version-controlled before the ITT is issued. Writing it afterwards invites a specification tuned to whichever bid you liked, and in a regulated procurement it is the sort of thing that turns a challenge into a successful one.
Publish the disqualifiers in the ITT
Disqualifiers are not traps. Stating them up front — we will not accept a right to train shared models on our project information, we will not accept exit priced at exit — saves everybody a bid cycle and quietly changes what suppliers offer. Several will simply agree, because their standard position was never a commercial requirement, only a default.
Keep a criterion for the things that only fail later
Provenance, continuity and exit are all invisible at acceptance and expensive at year three. They carry thirty-five per cent of this matrix precisely because nothing in a pilot will surface them, and because they are the criteria a delivery-pressured panel will drop first if they are not weighted.
This guidance will help inform and empower buyers in the public sector, helping them to evaluate suppliers, then confidently and responsibly procure AI technologies for the benefit of citizens.
What good buying looks like in public
Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the operator's own published material.
The clearest public evidence sits in what large owners and contractors chose to structure rather than in what they chose to buy. In each case below the decisive move was commercial: a contractor that built where its own accumulated project record was the asset, and two public owners that published the route to market before there was anything to buy through it.
Three programmes read against the buying ladder
Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them. Two of the three cards use an industry-scene illustration because no operator image exists in our library — the illustration is not supplied by, or endorsed by, the operator.
SkanskaGlobal contractor and project developer · Nordics, Europe, US24
- Challenge
- Scaling AI capability across many operating units and hundreds of live projects, where every region separately licensing a general-purpose assistant would multiply cost, fragment the data and leave the firm's own accumulated project knowledge outside every tool it bought.
- Approach
- Skanska USA Building established a Digital Transformation and Solutions Team uniting its Data Solutions, Emerging Tech and AI capabilities, and developed the Sidekick suite and Skanska Metriks cost modelling as internal products built on the firm's own project record rather than as licensed general-purpose tools.
- Reported outcome
- Skanska reports that the Sidekick suite has scaled from an initial 2024 pilot to more than 1,000 employee users supporting work across 500-plus projects, with the tools intended to surface safety and operational risks earlier and reduce administrative burden.
- What it shows about the curveWhere the training asset is your own accumulated project record, building keeps the asset and removes the pilot-to-production procurement entirely — the extension from pilot to 500 projects was an internal scaling decision, not a re-tender.
Skanska — press release, Digital Transformation and Solutions Team (opens in a new tab)
National HighwaysPublic owner · England's strategic road network24
- Challenge
- A public owner cannot buy AI capability on a departmental card, and cannot convert a promising trial into an operational contract by agreement. Every route to a supplier has to be a published, competed route, decided long before anyone knows which technology will work.
- Approach
- National Highways sets out a Digital Roads programme covering how the network is designed, built, operated and used with digital data and technology, and publishes its supplier and commercial framework routes so that capability is bought through competed commercial arrangements rather than one-off engagements.
- Reported outcome
- National Highways publishes both the Digital Roads programme and its supplier-facing commercial routes, so the path from an innovation trial to an operational contract runs through published frameworks that suppliers can see and prepare for in advance.
- What it shows about the curveFor a public owner the route to production is itself a procurement artefact, and it has to be published before the pilot. Rung 4 is reached by the notice, not by the pilot result.
Network RailPublic owner · Britain's rail infrastructure34
- Challenge
- AI suppliers appear, merge and disappear faster than a traditional multi-year framework cycle, so a framework let in year one can exclude the best available supplier by year three — while a public owner still has to buy through a competed, published route.
- Approach
- Network Rail states that it revised its commercial and procurement activity for the Procurement Act 2023, which went live on 24 February 2025, reducing its routes to direct award, open framework or the competitive flexible procedure, and using dynamic markets that suppliers can apply to join at any time.
- Reported outcome
- Network Rail publishes that commercial pipelines are visible to suppliers through the Find a Tender Service and its own website, and that open frameworks can be reopened during their lifespan so new suppliers can join.
- What it shows about the curveAn open framework or a dynamic market is the structural answer to a young supplier market: it lets an owner add the AI supplier that did not exist when the framework was let, without abandoning competition.
Read together, the three cases separate the two ways up the ladder. Skanska's route removes the procurement problem by keeping the capability inside the business, which works precisely because the asset — decades of its own project knowledge — is not for sale. The two public owners cannot take that route for most capability, so they solve the same problem from the other end: publish the route to market early enough that a good pilot has somewhere to go. Both are commercial designs made before the technology decision, which is the consistent signature of rung 4.