Redefining Technology

Construction & InfrastructureReadiness & Transformation Roadmap

AI vendor selection for infrastructure: how owners and contractors evaluate, contract and govern AI suppliers

Infrastructure AI vendor selection is the procurement discipline of choosing, contracting and governing the suppliers that put AI into a capital programme. It differs from software buying in four places: training-data provenance, ownership of the model your data enriched, a performance floor that can drift, and an exit that returns something you can actually use.

Infrastructure client team reviewing AI supplier submissions against project model data and framework requirements
Construction & Infrastructure · Readiness & Transformation Roadmap

Key takeaways

  1. Infrastructure AI procurement fails on terms, not on technology. Four clauses decide the outcome — training-data provenance, ownership of the enriched model data, a performance floor with named drift responsibility, and an exit that returns usable artefacts — and a standard software contract contains none of them.
  2. The enriched model is the asset, not the software. A vendor arrives with a generic detector and leaves with one tuned on your as-built geometry, your defect history and your RFI corpus. If the contract does not name who owns that, you will rent your own project record back at renewal.
  3. Public owners cannot extend a successful pilot into production on goodwill. The route to production has to be designed into the notice that launched the pilot — an open framework, a dynamic market or a competitive flexible procedure — or the pilot's success triggers a fresh procurement and a year of delay.
  4. Vendor viability is a technical requirement, not a finance formality. Escrow that covers only source code is worthless for AI: the release package must include model weights, training and inference code, feature definitions and the evaluation harness, and the release must have been tested at least once.
  5. Portfolio governance starts at vendor four, not vendor ten. By the fourth AI supplier an owner typically has overlapping capability, four copies of the same site imagery leaving the CDE, and four renewal dates nobody is tracking — all of which are cheaper to prevent than to unwind.

Abbreviations used on this page

BIM
Building information modelling
CDE
Common data environment — the ISO 19650 project information store
EIR
Exchange information requirements — the brief a supplier answers under ISO 19650
AIR
Asset information requirements — what the operating estate needs handed over
IFC
Industry Foundation Classes — the buildingSMART open model exchange format
ITT
Invitation to tender
RFP
Request for proposal
CFP
Competitive flexible procedure — the buyer-designed route under the Procurement Act 2023
FTS
Find a Tender Service — the UK public procurement notice platform
SLA
Service level agreement
AIMS
AI management system — the ISO/IEC 42001 construct
TCO
Total cost of ownership

Free · 8 questions · ~3 minutes

Score how you buy AI

Eight questions, one at a time, about three minutes. Answer them and we build your personalised vendor readiness report — where you sit on the buying ladder, your score on each of the four dimensions, and the specific clause or process gap standing between you and the next rung — and send it to your inbox. Your result doubles as the agenda for your next AI tender.

0 of 8 answered

Question 1 of 8Requirements clarity

How is an AI supplier's capability actually evaluated before award?

Every bidder looks strong on data they tuned for. Only a blind run on your own package separates them.

How the score maps to a stage
  • 04 — Stage 1, Ad hoc buying. AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
  • 510 — Stage 2, Structured requirements. AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
  • 1115 — Stage 3, Piloted with terms. Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
  • 1620 — Stage 4, Contracted for production. AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
  • 2124 — Stage 5, Portfolio-managed. AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.

What infrastructure AI vendor selection is — and why it is not software buying

A definition, the four places an AI purchase differs from a software purchase, and the path a requirement travels to become a production contract.

Infrastructure AI vendor selection is the procurement discipline of choosing, contracting and governing the suppliers that put AI into a capital programme or an operating estate. It covers the make-or-buy decision, the evaluation method, the data and intellectual property terms, the route to market that public procurement rules allow, the continuity provisions that survive a supplier failing, and the portfolio governance that starts once an owner is running more than three or four AI contracts at once.

It is not software buying with a different noun in the title. Four things differ, and each of them is a clause a standard software contract does not contain. The first is training-data provenance: what the model you are buying was built from, and whether your project information will be used to build the next one. The second is ownership of the enriched model data — the labels, the corrected geometry, the classified elements and the extracted quantities that only exist because your scheme paid for them. The third is that performance degrades on its own, so a system that met its specification at acceptance may not meet it in year three, and somebody has to be contractually responsible for that. The fourth is exit: a software exit returns your records, an AI exit has to return something you can keep operating with.

Value the owner retains, against maturity on the buying ladder

The curve is not linear, and the inflection is commercial rather than technical. Value retained by the owner stays close to flat through ad hoc buying and structured requirements — where most owners are — and rises sharply once terms are signed before the pilot rather than after it, because that is the point at which the enriched model and the data it was built from stop leaving the business by default.

Value retained by the owner by stage

  • Stage 1 · Ad hoc buying — 24% of operators. AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.
  • Stage 2 · Structured requirements — 34% of operators. AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.
  • Stage 3 · Piloted with terms — 23% of operators. Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.
  • Stage 4 · Contracted for production — 14% of operators. AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.
  • Stage 5 · Portfolio-managed — 5% of operators. AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the UK Government's Guidelines for AI procurement.

How an AI requirement becomes a production contract

The commercial path, rung by rung. The rung is determined by where the arrow ends: ad hoc buying terminates in an auto-renewal nobody owns, structured piloting terminates in a gate review, and only a route named at the notice stage reaches a production call-off and a portfolio register. Most owners are in the top lane.

  • Where value leaks
  • Data & feeds
  • AI / model
  • System-of-record action
  • Human in the loop

The process, in words

  • At rungs 1–2, an AI capability enters through a demonstration on the supplier's own data, becomes a free proof of concept running on real project information, and is regularised with a card invoice under the supplier's standard terms. It then auto-renews with no named owner, no re-test and no benchmark. This is where ownership of the enriched model quietly leaves the business.
  • At rung 3, the requirement is extracted from the exchange and asset information requirements rather than from a feature list. A held-back package of your own data is withheld from every bidder, all shortlisted suppliers run the same scoring harness on the same day, and the pilot only starts once provenance, ownership, the performance floor and the exit terms are signed. The gate review at the end is a commercial decision, not a technical one.
  • At rungs 4–5, the production route was named in the notice that launched the pilot, so exercising a priced option is a purchase order rather than a fresh competition. Viability diligence runs in parallel with the pilot, and every awarded supplier lands on a portfolio register that tracks what data it holds, what floor it must meet and when it is next re-benchmarked.
Step-by-step insights
The vendor demonstration — why every bidder looks the same
A demonstration is a supplier's best asset shown on their best data, usually a scheme whose geometry, labelling conventions and lighting the model has already been tuned against. That is not deception; it is what a demonstration is for. The consequence for a buyer is that all credible bidders present indistinguishably strong results, so the evaluation panel is forced onto criteria it can differentiate — price, references, cultural fit — none of which predicts performance on your night-survey footage in February. The single highest-leverage change in infrastructure AI procurement is refusing to be shown anything, and instead handing every bidder the same unfamiliar package.
The free proof of concept — the most expensive free thing in the estate
A no-cost trial requires no purchase order, so it requires no terms review, and it typically runs on genuine project data because synthetic data would not prove anything. Two assets transfer during it: your project information goes out, and a model tuned to your conventions comes back — owned by the supplier under an agreement nobody read. When the trial converts, it converts into a negotiation where the supplier already holds the tuned model and the site teams have already changed how they work. Charge for the pilot if it helps; the point is to have a contract, not a price.
EIR and AIR as the source of the requirement
Under ISO 19650 the exchange information requirements already state what information the project needs, in what form, at which stage, and to what level of information need. An AI requirement written from that document is testable: named information containers, a defined delivery point, a measurable acceptance metric. An AI requirement written from a supplier's feature list is a wish list of capabilities that every bidder can claim. Writing the requirement from the EIR also has a procedural benefit — it keeps the specification supplier-neutral, which matters a great deal if the competition is ever challenged.
The held-back package and the one-day bake-off
The held-back package is a genuine slice of your work — a fortnight of survey imagery, a set of as-built information containers, a tranche of RFI text — deliberately including the awkward cases: poor light, a non-standard asset class, a package modelled by a different designer. Nobody sees it before the day. Every shortlisted bidder runs the same harness, on the same data, in the same window, and the harness was written before any bid arrived. The scores that come back are frequently half what the demonstrations implied, and the ranking is frequently different from the panel's expectation.
Viability diligence running alongside the pilot
Technical evaluation and financial diligence usually run in series, which wastes the pilot period. Run them together: filed accounts, funding position, customer concentration, the escrow package's actual contents, change-of-control provisions and where the model would run if the supplier disappeared. This is not a credit check. The question is narrower and more useful — if this company is acquired in eighteen months, what do we hold, and can we run it? An answer of weights, inference code, feature definitions and a tested release is a different risk profile from an answer of a source-code deposit nobody has opened.
The route named at the notice — the clause that saves a year
For public owners this is the difference between a pilot that scales and a pilot that expires. If the notice that launched the pilot was a one-off, low-value engagement, a successful result cannot lawfully be extended into a production contract of any size, and the reward for a good pilot is a fresh competition. If instead the pilot was run inside an open framework, a dynamic market or a competitive flexible procedure that anticipated a production phase, the extension is a call-off. The decision is made months before anyone knows whether the technology works, which is precisely why it is so often made wrongly.

The rest of this page is organised around that path. The buying ladder and its five rungs come first, then where owners actually sit, then the build-or-buy decision, then the two documents that do most of the work: a weighted evaluation matrix and a clause-by-clause comparison of what an AI contract must add to a software one. The UK Government's Guidelines for AI procurement (opens in a new tab) and the AI Playbook for the UK Government (opens in a new tab) are the closest public-sector equivalents and are worth reading alongside it.

The five rungs of the buying ladder in detail

For each rung: what it actually looks like in a commercial team, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps owners there, and what leaving costs.

Each rung below is written for a commercial lead and an information manager rather than for a buyer of software. The hallmarks describe observable conditions in your contract file and your CDE access list, the diagnostic signals are checks you can run this week against documents you already hold, and the anti-pattern is the specific mistake most often made trying to leave that rung.

Select a rung

Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Ad hoc buying

24% of operators sit here

AI arrives through demos, free trials and expensed subscriptions, with no requirement written down and no terms beyond the vendor's own.

Stage 1 is not the absence of AI — most contractors at this stage have more AI in the business than their leadership believes. It is the absence of a buying process, so the capability arrives one project at a time, through whichever supplier ran the most convincing demonstration at whichever site had budget that month. The tools frequently work. What is missing is any record of what was agreed.

The tell is the contract file. At stage 1 the governing document for an AI supplier that reads your drawings, your site imagery and your programme is a click-through agreement nobody in the business has opened, containing a broad licence to use customer content to improve the services. That clause is not unusual and it is not hidden; it is simply never read, because the purchase never passed a desk that reads clauses.

The cost of staying here is not the subscription spend, which is usually trivial. It is that every project's evaluation work amortises nothing — the tenth trial is judged exactly as badly as the first — and that the information leaving your common data environment is uncounted. When someone eventually asks which suppliers hold copies of a scheme's model data, the honest answer takes a fortnight to assemble.

In practice

The site that bought its own computer vision

A project director on a station refurbishment saw a progress-capture demo at a conference, signed up for a three-month trial on a departmental card, and connected the supplier's app to the project's photo library and the shared model. The trial was genuinely useful and the team extended it twice. Eighteen months later the commercial team, preparing a claim, discovered that the definitive weekly progress record for a disputed period sat in a supplier's cloud under terms that gave the project a licence rather than ownership, and no export format better than PDF.

What it looks like

  • AI tools enter through a project team's card, not through procurement
  • Evaluation is a demo on the vendor's data, judged by impression
  • Terms are the vendor's standard online agreement, unread
  • Nobody can list which AI suppliers currently touch project information

Diagnostic signals you can check this week

  • Ask for a list of AI suppliers that hold project information. If the answer requires an expenses audit, you are here
  • Open the terms of the last AI tool a project adopted and search for the words train, improve and aggregate
  • Ask who evaluated the tool, and against what written criteria. Usually one person, and none
  • Check whether any AI subscription in the business has a named renewal owner

Anti-pattern · Banning the tools instead of buying them properly

The instinctive reaction to discovering uncontrolled AI is a blanket prohibition and a mandatory approval queue. It removes the visible tools and keeps none of the value: teams either revert to the manual process the tool replaced or move the same activity to a personal account, where you now have the same data exposure with none of the logging. The durable fix is a fast, published route to buy — a shortlist, a standard set of terms and a decision inside two weeks — so the compliant path is also the quickest one.

What holds you here

There is no written requirement and no standard term set, so every purchase is judged on a demonstration and governed by the supplier's own paper.

Highest-leverage next move

Write one requirement document and one set of standard AI terms, and route every new AI purchase through them — even the £5,000 ones.

Cost of leaving

Effort
1–3 months
Team
One commercial lead and one information manager, part-time
Risk
Low — the work is a register, a standard term set and a published route
To next stage
1–3 months

If this is you, the next step is

A two-week exercise: who holds your project data, under what terms, and what to renegotiate first.

Get your AI supplier register built

Stage 2

Structured requirements

34% of operators sit here

AI purchases go through a written requirement and a scored evaluation, but the terms are still the vendor's and the pilot has no defined route to production.

Stage 2 is where most infrastructure owners and tier-one contractors are, and it looks like good practice because it is good practice — for buying software. A requirement is written, a shortlist is scored, a pilot is funded and a business case is prepared. Every step is defensible. The problem is that the artefacts being bought are not software artefacts, and the process never asks the questions that separate one AI supplier from another.

Two omissions do most of the damage. The first is that evaluation happens on the vendor's own data: every bidder demonstrates on a scheme they have already tuned for, so every demonstration looks equally impressive and the score reflects presentation quality. The second is that the pilot is scoped as a proof of value and nothing else, so a successful pilot produces a slide deck, an enthusiastic sponsor and no contractual path to buying the thing at scale.

Time at stage 2 is expensive in a specific way. Each pilot enriches a supplier's model with your project information under terms you did not draft, and each ends without converting, so the supplier accumulates the durable asset and you accumulate the evaluations. Owners who have run four or five pilots in this pattern usually find their negotiating position has quietly weakened, because the incumbent now knows their data better than any challenger can.

In practice

The bake-off that everybody won

A highways client shortlisted four suppliers for automated defect detection and asked each to present. All four demonstrated on their own reference imagery, all four reported recall in the nineties, and the evaluation panel could not separate them on capability, so the award went on price and cultural fit. The winning system, run on the client's own night-survey footage in variable weather, performed far below its demonstration and the difference was not discoverable from anything in the tender.

What it looks like

  • A written requirement exists and bidders are scored against it
  • Evaluation still happens on vendor-supplied demonstration data
  • The contract is the vendor's paper with a few negotiated redlines
  • Pilots are funded as one-off proofs of value with no production option

Diagnostic signals you can check this week

  • Ask whether any bidder has ever been scored on data they had not seen before
  • Look for the words performance floor or acceptance test in the last AI contract signed. Usually absent
  • Check whether the last pilot's contract contained a priced option to extend into production
  • Ask who owns the labels created during the last pilot. If the answer is a shrug, the vendor owns them

Anti-pattern · Writing a longer requirement instead of a harder test

When the shortlist proves indistinguishable, the reflex is to add requirements — more questions, more compliance schedules, more weighting categories. It makes the document heavier and the decision no better, because every bidder can answer every question affirmatively and no answer is testable. The step that actually separates suppliers is small and unpopular: hold back one real package of your own data, write one scoring harness, and make every shortlisted bidder run on it on the same day.

What holds you here

Bidders are still judged on their own demonstration data and the pilot has no contractual route into production, so a good result cannot be bought at scale.

Highest-leverage next move

Hold back a real package of your own data, score every shortlisted bidder on the same harness on the same day, and sign the model and data terms before the pilot starts.

Cost of leaving

Effort
3–6 months
Team
Commercial lead, information manager, one technical assessor, a named sponsor
Risk
Medium — the held-back package has to be prepared carefully and lawfully
To next stage
3–6 months

If this is you, the next step is

We prepare the held-back package and the scoring script; you run the bake-off.

Build a blind evaluation harness

Stage 3

Piloted with terms

23% of operators sit here

Pilots run on your own data under negotiated AI terms — provenance, ownership, a performance floor and an exit — with a defined gate into production.

Stage 3 is the first stage where a pilot is a purchase decision rather than an experiment. The change is not in how the pilot runs but in what has been agreed before it starts: who owns the labels, who owns the tuned model, what the system must achieve in your units, who pays when it stops achieving it, and what you get back if the arrangement ends. Those are one workshop and one schedule, and they are the difference between an option and a sunk cost.

The evaluation discipline is equally concrete. A held-back package — a fortnight of survey imagery, a set of as-built information containers, a slice of the RFI record — is withheld from every bidder, and the scoring harness is written before anyone bids. Scores stop reflecting presentation quality and start reflecting performance on the geometry, the weather and the labelling conventions your business actually produces. Suppliers who object to blind evaluation are giving you information.

What emerges at stage 3 is a different negotiation. Because the terms are drafted before the pilot rather than after it, the leverage sits with the buyer at the moment leverage exists. Owners consistently report that clauses which are impossible to win after a successful pilot — enriched-data ownership, escrow scope, exit pricing — are routine to win before one, because at that point the supplier is competing rather than incumbent.

In practice

The fortnight nobody had seen

A rail infrastructure owner shortlisted five suppliers for automated asset-condition scoring and withheld two weeks of tunnel inspection footage, deliberately including one shift of poor lighting and one of a non-standard asset class. Every bidder ran the same harness on the same Tuesday. Reported recall across the five ranged from acceptable to less than half of what their marketing claimed, and the two suppliers whose numbers held up were not the two the panel had expected from the presentations.

What it looks like

  • Shortlisted bidders are scored blind on a held-back package of your data
  • Model and data schedules are signed before any pilot begins
  • A performance floor is stated in your units, with the measurement method
  • A gate review at the end of the pilot has a commercial decision attached to it

Diagnostic signals you can check this week

  • Ask to see the scoring harness used on the last AI award. It should exist as a file, not a spreadsheet of opinions
  • Check whether the last pilot contract defined enriched model data as customer data
  • Ask what the performance floor is for a live AI system, in your units. A number should come back
  • Check whether the pilot gate review had a commercial decision attached, or only a technical one

Anti-pattern · Treating the pilot as the procurement

A pilot run well under good terms feels like the hard part is over, so the production purchase is left until the pilot has proved itself. That sequencing hands the supplier the whole negotiation: by the time the pilot has succeeded, the model is tuned on your data, the site teams have changed how they work, and the alternative is a twelve-month restart. Price the production option at award, when three other bidders are still on the table, even if you never exercise it.

What holds you here

The pilot converts on goodwill rather than on a priced, pre-agreed option, so scaling reopens the commercial negotiation from a weaker position.

Highest-leverage next move

Put a priced, time-boxed option to extend into production in the pilot contract, agreed at award, with the unit rates and volume mechanism already in the schedule.

Cost of leaving

Effort
6–12 months
Team
Commercial and legal lead, information manager, technical assessor, project sponsor
Risk
Medium — the drafting effort is front-loaded and competes with delivery pressure
To next stage
6–12 months

If this is you, the next step is

The four clauses that decide the outcome, written for your CDE and your standard form of contract.

Draft the model and data schedule

Stage 4

Contracted for production

14% of operators sit here

AI is bought on a production contract with a priced extension path, tested continuity terms and a measurable floor that survives the pilot's champion.

Stage 4 is where an AI supplier becomes part of the estate rather than part of a project. The system is contracted for a term, with volume mechanics that survive a scheme finishing and another starting, and with the annual re-test that stops a five-year framework quietly becoming a five-year subscription to a 2026 model. The engineering was largely settled at stage 3; what changes here is that the commercial arrangement is built to outlast the people who negotiated it.

Continuity work is what distinguishes stage 4 from a well-run stage 3. AI suppliers in this sector are young, frequently venture-funded, and acquired at a rate that would be unremarkable in software and is alarming when the acquired product holds the definitive condition record for a tunnel. Escrow that covers only source code releases you a repository you cannot run; the release package has to include the weights, the training and inference code, the feature definitions and the evaluation harness, and someone has to have tried releasing it once.

The other stage-4 discipline is unglamorous: the renewal calendar. A production contract creates dates — floor re-test, benchmark, break, exit notice — and those dates only work if a named person owns them. Owners who reach stage 4 without a register find the contract's protections lapse silently, and discover it at the renewal where the price rises and the alternative has not been tested for three years.

In practice

The option that was already priced

An infrastructure client ran a six-month pilot of automated quantity extraction across two packages, with a production option priced at award covering a further four packages at fixed unit rates. The pilot succeeded. Exercising the option took a gate paper and a purchase order — eleven days from decision to live. The comparable client on the same corridor, whose pilot had no option, spent nine months running a new competition and awarded to the same supplier at a higher rate.

What it looks like

  • The production option was priced at award and exercised without a new competition
  • Escrow covers weights, training and inference code, features and the harness — and has been tested
  • The floor is re-tested annually against a fresh held-back package
  • Change of control, sub-processing and exit pricing are all fixed at signature

Diagnostic signals you can check this week

  • Ask when the escrow release was last tested. If never, you have a document, not a continuity plan
  • Check whether the performance floor has been re-tested since award, and on what data
  • Look for a change-of-control clause in the largest AI contract in the business
  • Ask who owns the renewal date. If it is a calendar reminder on one person's laptop, it is unowned

Anti-pattern · Buying continuity as a document rather than a drill

Escrow, exit plans and transition assistance are easy to agree because they cost nothing at signature and everything at release. The failure mode is uniform: the agreement names a release event, nobody ever exercises it, and when a supplier is acquired the release turns out to omit the training data pipeline, the feature definitions or the environment the weights run in. Run the release once, in a quiet quarter, and treat the gaps you find as contract defects rather than as an inconvenience.

What holds you here

Each supplier is contracted well in isolation, so overlapping capability, duplicated data egress and uncoordinated renewals accumulate across the portfolio.

Highest-leverage next move

Stand up a vendor and model register with concentration limits, an overlap review and a single renewal calendar across every AI contract in the business.

Cost of leaving

Effort
12–18 months
Team
Category owner, legal, information manager, an engineering owner for the floor re-test
Risk
Higher — continuity and exit obligations have to be real, and testing them costs supplier goodwill
To next stage
12–18 months

If this is you, the next step is

We run the escrow release against a real vendor package and report what is actually missing.

Stress-test a continuity plan

Stage 5

Portfolio-managed

5% of operators sit here

AI suppliers are governed as a portfolio — one register, deliberate overlap, concentration limits, a shared evaluation harness and a coordinated renewal cycle.

Stage 5 is narrower than it sounds, and it is a commercial capability rather than a technical one. The register is unremarkable: supplier, model, what data it holds, which schemes it touches, the floor, the renewal date, the exit price. What makes it stage 5 is that the register is consulted before a new purchase, so the sixth AI supplier has to demonstrate that it does something the previous five do not — which is a question stage-4 organisations never get asked.

The concentration question is the one owners reach late. A single supplier holding progress capture, condition scoring and quantity extraction across a whole region is efficient right up to the moment it is acquired, changes its pricing model or suffers an outage during a claims window. Stage-5 owners set an explicit ceiling on how much of the estate one supplier may hold, and accept the small integration cost of keeping a credible second source warm.

Sustaining stage 5 is mostly calendar discipline, and it is the stage most likely to regress. The register goes stale when a scheme closes, the harness rots when the data conventions change, and a year of no re-tests turns a portfolio back into a collection of contracts. The signal to watch is how long it takes to score a challenger: when that number starts rising, the shared harness has stopped being shared.

In practice

The sixth supplier that was refused

A national owner with five AI suppliers on its register was offered a sixth for automated drawing comparison. The register showed two incumbents already extracting the same information containers from the CDE, and the shared harness let the team score the challenger against both in nine days. The challenger was better on one asset class and worse on two, so instead of a sixth contract the owner added the asset class to an existing call-off at the framework's rates.

What it looks like

  • One register holds every AI supplier, model, data flow and renewal date
  • New capability is tested against the incumbent portfolio before a new supplier is added
  • Concentration limits cap how much of the estate any single supplier can hold
  • A shared evaluation harness is reused across categories, so a challenger can be scored in days

Diagnostic signals you can check this week

  • Ask how long it takes to score a new AI supplier against an incumbent. Under two weeks means the harness is real
  • Check whether the register records what data each supplier holds, not just what it costs
  • Ask whether any purchase in the last year was refused because an incumbent already covered it
  • Check whether a concentration limit exists as a written rule with a named approver for exceptions

Anti-pattern · Consolidating to one supplier because the register is tidy

A good portfolio view makes rationalisation tempting, and consolidating five suppliers into one platform genuinely simplifies integration, invoicing and governance. It also puts the definitive record for the whole estate behind one commercial relationship, one pricing model and one balance sheet, at which point the renewal negotiation has no alternative in it. Rationalise overlap, keep a second source warm in every category that touches the asset record, and treat the integration cost of doing so as insurance rather than waste.

What holds you here

Portfolio discipline decays quietly — registers go stale, harnesses rot, and re-tests slip until the portfolio is a collection of contracts again.

Highest-leverage next move

Put the register, the harness and the renewal calendar on a fixed review cadence with a named owner, and measure how long it takes to score a challenger.

Cost of leaving

Effort
Continuous
Team
A named AI category owner plus a standing commercial and information-management forum
Risk
Concentrated — low frequency, high consequence, and commercial rather than technical in nature

If this is you, the next step is

We map overlap, concentration and renewal exposure across every AI supplier you run.

Review your AI vendor portfolio

Where infrastructure owners and contractors actually sit

The distribution across the ladder, and why the drop between structured requirements and piloting with terms is the largest single loss.

Most infrastructure owners and contractors sit at structured requirements — rung two. They write a requirement, score a shortlist and fund a pilot, and they do all of it on the supplier's paper and the supplier's data. A much smaller group signs the model and data terms before the pilot starts, and a very small group has ever exercised a production option that was priced at award.

Illustrative distribution of infrastructure owners and contractors across the buying ladder

Structured requirements is the mode and the plateau. The drop from rung 2 to rung 3 is the largest single transition loss on the ladder, and it is a drafting problem rather than a technology problem. Figures are illustrative, synthesised from the public procurement guidance and construction-productivity research linked beneath the chart, not a measured survey.

Share of owners and contractors

  • 24% — 1 · Ad hoc buying
  • 34% — 2 · Structured requirements (the plateau)
  • 23% — 3 · Piloted with terms
  • 14% — 4 · Contracted for production
  • 5% — 5 · Portfolio-managed

Source: Illustrative distribution, synthesised from UK Government AI procurement guidance and McKinsey construction-productivity research

The plateau has a structural cause. A construction or infrastructure business already knows how to buy software and already knows how to buy design and construction services, and an AI supplier looks like the first while behaving like the second: the deliverable is partly built out of the client's own information, and its quality is contingent on that information. Buying it with a software process produces a contract that is silent on the three things that matter — provenance, the enriched model, and degradation over time. McKinsey's construction-productivity research (opens in a new tab) sets out how little the sector's productivity has moved over two decades, which is the gap every AI supplier is sold against and the reason these purchases get waved through on urgency.

The second cause is that the people who can fix it are rarely in the room together. Provenance and enriched-data ownership are information-management questions, the performance floor is an engineering question, escrow and change of control are legal questions, and the production option is a procurement question. Owners who move up the ladder almost always do it by putting those four people in one workshop before the ITT is issued rather than after the pilot succeeds.

Build, buy or partner: deciding before you go to market

The make-or-buy call turns on two variables — how unusual your data is, and how many credible suppliers exist — and it decides everything downstream about the terms you need.

Build, buy or partner is decided by two questions and not by capability: how unusual is the data the capability depends on, and how many credible suppliers already sell it? Where the data is industry-standard and the market is mature — object detection on site imagery, document search over a contract set — buy, and re-tender often, because the underlying capability improves faster than any owner can track. Where the data is your own accumulated programme record — outturn cost against estimate, historical change and claims, your own defect taxonomy — the asset is the data, no supplier has it, and the correct answer is to build or to partner with the model kept on your side of the line.

CapabilityDefault answerWhyWho must own the data asset
Progress capture from site imagery, drone and laser scanBuyCommodity computer vision with many credible suppliers; the differentiator is your as-built model and your labelling conventions, not the detectorYou own the imagery, the labels and the classified elements
Quantity take-off and model enrichmentBuy, own the outputThe extraction service is bought; the enriched information containers are the asset and must land back in the CDE in an open formatYou own the enriched containers and the extracted quantities
Programme and schedule risk predictionPartnerDepends on your historical programmes, change history and claims record — data no supplier holds and none can substitute forYou own the trained model and the feature definitions
Cost and estimate benchmarkingBuild or partnerYour rate history, subcontract packages and outturn costs are the model; sharing them with a supplier who serves competitors is a commercial decision, not a technical oneYou own the model and every benchmark derived from it
Asset condition and defect detection on the operating estateBuy, with a floorMature market on standard asset classes; contract a recall floor on your classes and your survey conditions rather than a generic accuracy claimYou own the labelled condition dataset and the survey record
Document, RFI and contract review assistantsBuy, re-tender oftenGeneral language capability moves faster than any owner can build against; assume a two-to-three-year useful life for the supplier choiceYou own the retrieval corpus and the full audit log
Safety observation and behavioural monitoringBuy, govern hardThe model is straightforward; the worker-identifiable data is not, and the governance load exceeds the technical loadYou own the imagery, the retention schedule and the DPIA
Generative design and optioneeringPartner with the designerDesign liability sits with the designer, not with the tool; the capability must live inside an existing design duty rather than beside itThe designer owns the outputs; you own the information deliverables
The make-or-buy map for construction and infrastructure AI. The default answer is a starting point, not a rule; the last column is the part that must survive whichever answer you choose.

The make-or-buy quadrant

Plot the uniqueness of the data the capability depends on against the number of credible suppliers. Three of the four quadrants have an answer that is not build, and the most valuable quadrant is the one owners most often mishandle.

Build or co-develop

  • Your data is the capability and nobody sells it
  • Cost benchmarking, claims prediction, your own defect taxonomy
  • Terms to win: model ownership, feature definitions, no shared training

Buy the engine, own the model

  • Mature engines, but tuned on information only you hold
  • The highest-value and most mishandled quadrant
  • Terms to win: enriched-data ownership, escrowed weights, portability

Wait, or open a dynamic market

  • Nothing credible to buy and nothing distinctive to build on
  • Common for genuinely novel capability
  • Move: a dynamic market or pre-market engagement, not a tender

Buy and re-tender often

  • Commodity capability, commodity data, many suppliers
  • Document assistants, standard object detection
  • Terms to win: short term, low exit cost, open export formats
Uniqueness of your data — top: Your own programme record, bottom: Industry-standard data
Supplier market maturity — left: Few credible suppliers, right: Many credible suppliers

The quadrant owners handle worst is the top right: a mature engine tuned on information only you hold. It looks like a straightforward purchase because the engine is a product with a price list, and it behaves like a joint development because the version that works on your estate exists nowhere else. Buying it on standard software terms transfers the only genuinely scarce asset in the transaction to the supplier for free, which is why the evaluation matrix in the next section but one puts twenty per cent of the score on performance against your own held-back data and fifteen on who ends up owning the model that produced it.

The vendor evaluation matrix: criteria, weights and disqualifiers

Nine criteria, a weight for each, the evidence to demand at tender stage, and the answer that should end a bid on the spot.

An AI vendor evaluation matrix differs from a software one in where the weight sits: twenty per cent on measured performance against data the bidder has never seen, and thirty per cent across provenance and ownership of what the engagement produces. The matrix below is the version we use with infrastructure owners and tier-one contractors. The weights are a defensible starting point rather than a standard — move them, but move them before the ITT is issued and record why, because a weighting changed after bids arrive is the most common way an award becomes challengeable.

CriterionWeightEvidence to demand at ITTAutomatic disqualifier
Measured performance on your held-back data20%A blind run on a package you withheld — your survey imagery, your as-built containers, your RFI text — scored on one harness you wrote, on one day, with every shortlisted bidder presentRefuses to be evaluated on anything but their own demonstration data
Training-data provenance15%A written data lineage statement: what the base model was trained on, on what licence basis, whether any competitor's project information sits in it, and whether your data will train anything sharedCannot state where the training data came from, or reserves a right to train shared models on your project information
Ownership of the enriched model and derived data15%Draft clause text naming who owns the fine-tuned weights, the label set, the embeddings, the corrected geometry and the extracted quantities at the end of the termClaims ownership of data derived from your CDE, or offers only a licence back to your own information
Exit and data portability10%A written exit plan: export formats (IFC, COBie, open label formats, model weights), the delivery window, transition assistance days, and a price fixed at signatureExit priced at the point of exit, or export available only in a proprietary format
Performance floor and drift responsibility10%A floor stated in your units, the measurement method and dataset, who detects drift, who pays for retraining, in what window, and the remedy if the floor is not restoredPerformance stated only as accuracy on the supplier's own benchmark, with no floor in your units
Integration with the CDE and the ISO 19650 workflow10%A demonstrated read and write against your common data environment with information containers, revision codes and suitability states preserved end to endIntegration by manual export, or a full copy of the CDE held permanently on the supplier's side
Vendor viability and continuity10%Filed accounts, funding position, customer concentration, the actual contents of the escrow package, change-of-control provisions and evidence a release has been testedNo escrow, no change-of-control clause, and a refusal to discuss financial runway
Security, worker data and site imagery5%DPIA support, a retention schedule for site imagery, worker-identifiability controls, the full sub-processor list and the hosting jurisdictionSite imagery retained indefinitely, or sub-processors undisclosed at tender stage
Five-year total cost of ownership5%Licence, per-seat and per-asset scaling, data egress, retraining, integration maintenance and the cost of the exit, modelled across five yearsYear-one price only, with the scaling mechanism left undefined
Weighted evaluation matrix for an infrastructure AI supplier. The evidence column is what goes into the ITT as a required submission; the disqualifier column is what ends the bid regardless of the rest of the score.
  • Score the evidence, not the answer

    Every bidder will answer yes to every question. The score belongs to the artefact attached to the answer: the lineage statement, the draft clause, the exit plan with a price on it. A criterion with no required artefact is a criterion that cannot discriminate, and it should either be given an artefact or removed from the matrix.

  • Write the harness before the bids arrive

    The scoring harness — the metric, the threshold, the data, the script — is written and version-controlled before the ITT is issued. Writing it afterwards invites a specification tuned to whichever bid you liked, and in a regulated procurement it is the sort of thing that turns a challenge into a successful one.

  • Publish the disqualifiers in the ITT

    Disqualifiers are not traps. Stating them up front — we will not accept a right to train shared models on our project information, we will not accept exit priced at exit — saves everybody a bid cycle and quietly changes what suppliers offer. Several will simply agree, because their standard position was never a commercial requirement, only a default.

  • Keep a criterion for the things that only fail later

    Provenance, continuity and exit are all invisible at acceptance and expensive at year three. They carry thirty-five per cent of this matrix precisely because nothing in a pilot will surface them, and because they are the criteria a delivery-pressured panel will drop first if they are not weighted.

This guidance will help inform and empower buyers in the public sector, helping them to evaluate suppliers, then confidently and responsibly procure AI technologies for the benefit of citizens.

What good buying looks like in public

Three publicly reported programmes, read against the ladder. None is an Atomic Loops engagement — each links to the operator's own published material.

The clearest public evidence sits in what large owners and contractors chose to structure rather than in what they chose to buy. In each case below the decisive move was commercial: a contractor that built where its own accumulated project record was the asset, and two public owners that published the route to market before there was anything to buy through it.

Three programmes read against the buying ladder

Outcomes as reported by the operators themselves. Verify figures against the linked source before reusing them; we have not independently audited them. Two of the three cards use an industry-scene illustration because no operator image exists in our library — the illustration is not supplied by, or endorsed by, the operator.

Contractor project team reviewing digital construction data on a large infrastructure siteSkanskaGlobal contractor and project developer · Nordics, Europe, US24
Challenge
Scaling AI capability across many operating units and hundreds of live projects, where every region separately licensing a general-purpose assistant would multiply cost, fragment the data and leave the firm's own accumulated project knowledge outside every tool it bought.
Approach
Skanska USA Building established a Digital Transformation and Solutions Team uniting its Data Solutions, Emerging Tech and AI capabilities, and developed the Sidekick suite and Skanska Metriks cost modelling as internal products built on the firm's own project record rather than as licensed general-purpose tools.
Reported outcome
Skanska reports that the Sidekick suite has scaled from an initial 2024 pilot to more than 1,000 employee users supporting work across 500-plus projects, with the tools intended to surface safety and operational risks earlier and reduce administrative burden.
What it shows about the curveWhere the training asset is your own accumulated project record, building keeps the asset and removes the pilot-to-production procurement entirely — the extension from pilot to 500 projects was an internal scaling decision, not a re-tender.

Skanska — press release, Digital Transformation and Solutions Team (opens in a new tab)

Illustrative scene: strategic road network operations team reviewing digital asset and supplier performance dataNational HighwaysPublic owner · England's strategic road network24
Challenge
A public owner cannot buy AI capability on a departmental card, and cannot convert a promising trial into an operational contract by agreement. Every route to a supplier has to be a published, competed route, decided long before anyone knows which technology will work.
Approach
National Highways sets out a Digital Roads programme covering how the network is designed, built, operated and used with digital data and technology, and publishes its supplier and commercial framework routes so that capability is bought through competed commercial arrangements rather than one-off engagements.
Reported outcome
National Highways publishes both the Digital Roads programme and its supplier-facing commercial routes, so the path from an innovation trial to an operational contract runs through published frameworks that suppliers can see and prepare for in advance.
What it shows about the curveFor a public owner the route to production is itself a procurement artefact, and it has to be published before the pilot. Rung 4 is reached by the notice, not by the pilot result.

National Highways — Digital Roads (opens in a new tab)

Illustrative scene: infrastructure owner and supplier teams working through a technology partnership agreementNetwork RailPublic owner · Britain's rail infrastructure34
Challenge
AI suppliers appear, merge and disappear faster than a traditional multi-year framework cycle, so a framework let in year one can exclude the best available supplier by year three — while a public owner still has to buy through a competed, published route.
Approach
Network Rail states that it revised its commercial and procurement activity for the Procurement Act 2023, which went live on 24 February 2025, reducing its routes to direct award, open framework or the competitive flexible procedure, and using dynamic markets that suppliers can apply to join at any time.
Reported outcome
Network Rail publishes that commercial pipelines are visible to suppliers through the Find a Tender Service and its own website, and that open frameworks can be reopened during their lifespan so new suppliers can join.
What it shows about the curveAn open framework or a dynamic market is the structural answer to a young supplier market: it lets an owner add the AI supplier that did not exist when the framework was let, without abandoning competition.

Network Rail — Procurement Act 2023 (opens in a new tab)

Read together, the three cases separate the two ways up the ladder. Skanska's route removes the procurement problem by keeping the capability inside the business, which works precisely because the asset — decades of its own project knowledge — is not for sale. The two public owners cannot take that route for most capability, so they solve the same problem from the other end: publish the route to market early enough that a good pilot has somewhere to go. Both are commercial designs made before the technology decision, which is the consistent signature of rung 4.

The model data problem: who owns what your project enriched

Seven assets change hands during an AI engagement. A default vendor contract assigns most of them to the vendor, and ISO 19650 has a name for only some of them.

Ownership in an AI engagement is not one question but seven, and a default vendor contract answers most of them in the vendor's favour by omission rather than by argument. The raw capture is usually acknowledged as yours. Everything downstream of it — the labels created from it, the model tuned on it, the corrected geometry, the extracted quantities, the evaluation harness and the decision log — is either unmentioned or defined as output of the service, which is a licence rather than a transfer. The asymmetry is not adversarial; it is the shape of a software contract applied to something that is not software.

The reason this matters more in infrastructure than elsewhere is that the information has a statutory and contractual afterlife. Under ISO 19650 (opens in a new tab) and the UK BIM Framework (opens in a new tab), project information is produced against exchange information requirements, held in a common data environment with revision codes and suitability states, and handed over as an asset information model that the operator will rely on for decades. An AI supplier that enriches those containers is producing project information whether or not the contract calls it that. If the enriched containers cannot be exported in IFC or another open format (opens in a new tab), the handover is defective in a way nobody notices until the operating estate needs it.

AssetWhat a default vendor contract saysWhat your clause should sayWhere it sits in ISO 19650 terms
Raw capture — photographs, video, point clouds, drone surveyCustomer data, but hosted and retained on the supplier's platform for the term and often beyondCustomer-owned, retention capped and evidenced on deletion, with no training use unless separately and specifically agreedProject information in the CDE, typically in the shared or published state
Labels and annotations created during the engagementSupplier intellectual property, on the basis that the supplier's staff or tooling produced themCustomer-owned. The label set is the reusable asset and the single thing that makes your next tender competitiveAn information container with its own revision code and suitability state
The fine-tuned model and its weightsSupplier product; the customer licenses access for the term and receives nothing on expiryCustomer-owned, or jointly owned, with an escrowed copy, the feature definitions and a tested portability routeNot an ISO 19650 container at all — which is precisely why it goes missing at handover
Enriched model data — classified elements, corrected geometry, extracted quantitiesOutput of the service, licensed to the customer for the duration of the subscriptionCustomer data by definition; the supplier holds a limited processing licence for the term and nothing after itInformation containers delivered against the EIR, exportable in IFC or the native authoring format
The evaluation harness and the held-back test setNot mentioned anywhere in the agreementCustomer-owned. It is your only means of re-testing the floor next year or scoring a replacement supplier in days rather than monthsHeld alongside the AIR as an acceptance and verification artefact
Aggregated benchmarks derived from your dataSupplier may use aggregated and anonymised customer data freely, including in market-facing productsPermitted only for named purposes, with no re-identification and no competitor-facing benchmark containing your outturn costs or ratesOutside the CDE entirely — govern it in the contract, not in the information protocol
The decision and audit logRetained by the supplier for support purposes, with access during the termCustomer-owned, exportable in a readable format, retained for the same period as the rest of the project recordPart of the project information model handed over at completion
The seven assets that change hands in an infrastructure AI engagement, what a default vendor contract typically says about each, and the position to hold. The final column shows where each asset sits — or conspicuously fails to sit — in the ISO 19650 information-management structure.
  • Define enriched model data in the definitions, not in a schedule

    The single most effective drafting change is putting a definition of enriched model data — labels, embeddings, corrected geometry, classified elements, extracted quantities — into the definitions clause and folding it into customer data. Everything downstream then inherits the right answer without further negotiation, including confidentiality, retention, deletion and the exit obligation.

  • Restrict training rights narrowly and name the exception

    A blanket prohibition on training is often refused and sometimes unreasonable — a supplier may genuinely need to tune on your data to serve you. Prohibit training of shared or other-customer models on your project information, permit training of the model instance dedicated to you, and name any exception explicitly with a stated purpose and duration.

  • Make the export format a specification, not a promise

    An exit clause that says data will be returned in a mutually agreed format is unenforceable at the moment you need it. Name the formats: IFC or COBie for model and asset data, an open label format for annotations, a stated serialisation for weights, and readable exports for logs. The National Institute of Building Sciences (opens in a new tab) and buildingSMART both publish the open standards to point at.

  • Treat worker-identifiable imagery as a separate category

    Site imagery containing identifiable workers carries obligations that model data does not. The ICO's guidance on AI and data protection (opens in a new tab) sets the expectations; the contractual consequence is a distinct retention schedule, a named lawful basis, sub-processor transparency and a deletion obligation you can evidence.

One practical test settles most of these arguments quickly. Ask the supplier what you would hold on the day after termination, and ask for it as a list of files rather than as a sentence. A supplier at the top of this market will answer with the enriched containers in IFC, the label set, a serialised model, the feature definitions and the logs. A supplier at the bottom will answer with a PDF export and an offer to discuss it at the time. That single question separates the field faster than any section of the ITT.

Public procurement: which routes let you pilot AI at all

For a public infrastructure owner the procurement route is decided before the technology, and it determines whether a successful pilot can ever become a contract.

Public procurement rules do not stop an infrastructure owner piloting AI, but they decide which pilots can become contracts — and that decision is made at the notice stage, months before anyone knows whether the technology works. A trial run as an isolated, below-threshold engagement is lawful and useful, and its success creates no route to a production purchase of any size. A trial run inside an open framework, a dynamic market or a competitive flexible procedure that anticipated production creates one automatically. Same technology, same result, entirely different outcome.

In the UK the Procurement Act 2023 (opens in a new tab) went live on 24 February 2025, and Network Rail describes the practical effect for a large infrastructure buyer: routes reduced to direct award, open framework or the competitive flexible procedure, frameworks that can be reopened during their lifespan, dynamic markets suppliers may join at any time, and pipeline visibility through the Find a Tender Service (opens in a new tab). In the EU regime the public procurement directives (opens in a new tab) — 2014/24/EU for public contracts and 2014/25/EU for utilities — provide the competitive dialogue and innovation partnership procedures for the same purpose. The vocabulary differs; the design question does not.

RouteWhat it isWhat it gives an AI buyerWhere it fails for AI
Pre-market engagementPublished supplier days, requests for information and prior information notices issued before any competition beginsThe only defensible way to learn what the AI market can actually do before writing a requirement no supplier can meet — or one only the incumbent canEngagement that names a preferred supplier, or shares information unevenly, taints the competition that follows and invites challenge
Direct awardAward without competition on the narrow grounds set out in the regimeSpeed where a genuine single-supplier justification exists — for example continuing an existing system that nothing else can interoperate withAlmost never available for capability several suppliers could provide, and a challenged direct award costs a programme more time than the competition would have
Open frameworkA framework that can be reopened during its life so new suppliers can join laterThe structural answer to a young supplier market: you can add the AI supplier that did not exist when the framework was let, without abandoning competitionOnly works if the lots were drafted broadly enough to contain AI capability at all — narrow lot definitions are the usual reason a framework cannot be used
Dynamic marketA register suppliers may apply to join at any time, used to run quick competitions among qualified membersShort, repeatable competitions against a live supplier list — the fastest compliant route from a scoped requirement to a running pilotThe register is only as good as its qualification questions; shallow qualification produces a long shortlist and moves all the work back into evaluation
Competitive flexible procedureA procedure the buyer designs, which may include negotiation, staged shortlisting and multiple roundsOne lawful process containing the whole path: shortlist, paid bake-off on your held-back data, then award with the production option already pricedThe design effort sits with the buyer, and a badly designed bespoke procedure is harder to defend than a standard one run properly
Competitive dialogue (EU regimes)Structured dialogue with shortlisted bidders to develop a solution before final tenders are submittedRoom to co-develop a specification for a capability you genuinely cannot specify up front, with the requirement converging during the processLong and expensive for bidders; small AI suppliers frequently withdraw, leaving you with the systems integrators rather than the specialists
Innovation partnership (EU regimes)Research and development plus purchase of the resulting solution inside a single procedureA lawful path from development through to production purchase without a second competitionHeavy machinery for a small pilot; proportionate only where the capability is genuinely novel and the eventual volume is large
Below-threshold purchaseBuying beneath the value threshold under lighter-touch rulesA legitimate way to fund a short, bounded proof of value quickly when the outcome is genuinely uncertainSplitting a production requirement into a series of below-threshold pilots to avoid competition is unlawful, and it is the single most common mistake in this space
Routes to market for an infrastructure AI purchase, and what each one does for a buyer who cannot yet specify the capability precisely. Names follow the UK Procurement Act 2023 regime, with the EU equivalents noted where they differ.

Designing a compliant AI pilot that can scale

  1. Decide the production route before the pilot scope

    Write down, in one paragraph, how a successful pilot becomes a production contract — which framework, which lot, which option, which threshold. If that paragraph cannot be written, the pilot is an experiment rather than a procurement step, and it should be funded and scoped accordingly.

  2. Run pre-market engagement in public and symmetrically

    Publish the engagement, invite the market rather than a shortlist, share the same material with everyone and record what was shared and when. Procurement policy notes (opens in a new tab) and the Construction Playbook (opens in a new tab) both set expectations for early market engagement in this sector.

  3. Specify outputs and evidence, never a named technique

    Specify the information you need, the format it must arrive in and the measured floor it must meet. Specifying an architecture, a model family or a named product narrows the field artificially, dates the specification within a year, and is the most challengeable part of an AI ITT.

  4. Pay for the bake-off

    A blind run on a held-back package is real work. Paying a modest, equal fee to each shortlisted bidder keeps small specialists in the process, makes the obligation to run properly contractual, and removes the argument that unpaid evaluation effort favours large suppliers with bid teams.

  5. Price the production option at award

    Put the option, its window, its unit rates and its volume mechanism in the contract signed at the end of the competition — when alternatives still exist. Exercising it later becomes a gate paper and a purchase order rather than a nine-month restart.

Two other regimes touch the same purchase. The EU AI Act's regulatory framework (opens in a new tab) classifies systems by risk and imposes obligations that flow to whoever places a system on the market or puts it into service, which for an infrastructure owner means the classification question belongs in the ITT rather than in a later compliance review. And where a capability is bought through a central purchasing body such as Crown Commercial Service (opens in a new tab), the framework's own terms may already answer — or already fix — several rows of the clause table in the next section, so read them before drafting a schedule you cannot use.

The clause table: what an AI contract must add to a software contract

Twelve clauses, side by side. The left column is what your current template says; the middle column is what has to be added before an AI supplier signs it.

An AI contract must add twelve things to a software contract, and each addition exists because a specific failure has a specific price. The table below is the working document: run your current template down the left column, and anywhere the middle column has content your template does not, you have found a clause that will be argued about at renewal, at exit, or in the week your supplier is acquired. Nothing in the middle column is exotic. Most of it is standard in other regulated procurement and simply absent from software paper.

ClauseWhat a software contract usually saysWhat an infrastructure AI contract must addCost of leaving it out
Scope of licenceA licence to use the software for internal business purposesSeparate three assets explicitly: the base model, the model fine-tuned on your data, and the data derived from your models. Three assets, three ownership positions, three exit obligationsThe tuned model is treated as the supplier's product. You funded it, and you cannot take it anywhere
Training rights over your informationSilent, or a broad right to use customer content to improve the servicesAn explicit prohibition on training shared or other-customer models on your project information, with any exception named, purpose-limited and time-limitedYour as-built geometry, defect history and rate data improve a model a competitor rents next year
Ownership of enriched model dataThe customer owns customer data, which is defined as data the customer uploadsDefine enriched model data — labels, embeddings, extracted quantities, classified elements, corrected geometry — as customer data, with the supplier holding a processing licence for the term onlyYou pay to license your own quantities and classifications back at every renewal
Performance floorAn availability SLA covering uptime and response timeA quality floor in your units — detection recall on your asset classes, quantity variance against measured, false-alarm rate per shift — plus the agreed measurement method and the dataset it is measured onThe system is available and wrong, and there is no contractual event to trigger
Drift and retraining responsibilityMaintenance, support and updates during the termWho detects drift and on what signal, who pays for retraining, within what window it must be restored, and the remedy if the floor is not recoveredSilent degradation across a five-year framework with no remedy and no early warning
Human-decision boundaryNot addressed at allWhich decisions the system may make, which it may only recommend, what record is kept for each, and how that maps to your AI management system and to the EU AI Act risk class where it appliesA safety-relevant or design-relevant decision made by a system nobody ever classified
Evaluation and acceptanceAcceptance on delivery of the specified featuresAcceptance on a blind run against a held-back package, using the harness annexed to the contract, with an annual right to re-test on fresh dataAcceptance degrades to it installed, and you have no later basis to challenge quality
Pilot-to-production optionA separate order form for each additional purchaseA priced, time-boxed option to extend into production, agreed at award, with the unit rates, the volume mechanism and the mobilisation window already in the scheduleA successful pilot triggers a fresh procurement and a nine-to-twelve-month gap in capability
Continuity and escrowSource-code escrow, if anything at allEscrow of model weights, training and inference code, feature definitions and the evaluation harness, with a defined release event and evidence that a release has actually been testedThe supplier is acquired, maintenance stops, and you hold a repository you cannot run
Exit and transitionData returned in a standard format on requestNamed export formats for each asset, a delivery window in days, a stated number of transition assistance days, and a price fixed at signature rather than quoted at exitThe exit is priced at the exact moment you have the least leverage in the relationship
Sub-processing, hosting and worker dataSub-processors listed on a web page the supplier may change at willNamed sub-processors, stated jurisdiction, notice and objection rights for changes, plus a retention limit and deletion evidence for site imagery and worker-identifiable dataFootage of your workforce moves to a jurisdiction your data protection impact assessment never covered
Benchmark and re-tender rightsAutomatic renewal unless cancelled within a notice windowAn annual benchmark against the market, a re-test of the performance floor on fresh data, and an express right to run a mini-competition inside the frameworkYou renew a 2026 model at 2029 prices, having never tested whether anything better exists
Contract clause comparison for an infrastructure AI purchase. The right-hand column is the cost of leaving the clause out — in every case a cost paid later, by someone who did not negotiate the contract.

Two rows in that table deserve a note. The human-decision boundary row should reference the AI management system your organisation is building — ISO/IEC 42001 (opens in a new tab) is the emerging standard, and the NIST AI Risk Management Framework (opens in a new tab) is the most widely used free equivalent — because a contractual decision boundary that does not map to an internal control is a sentence rather than a control. And the escrow row is the one most often agreed and least often tested: an escrow package that omits the feature definitions releases you a model you cannot feed.

Structuring a pilot so success is not a re-procurement

  1. One contract, two phases, one signature

    Sign a single agreement covering the pilot and the production phase, with the production phase drafted as an option the buyer may exercise. This is the structural change that removes the second procurement, and it is far easier to agree while three other bidders are still live.

  2. Put the exercise conditions in writing, and make them measurable

    The option should be exercisable on stated conditions — the floor met on the held-back package, integration demonstrated against the CDE, the exit plan tested. Conditions phrased as satisfaction of the buyer are unenforceable in one direction and unfundable in the other.

  3. Price the production phase now, index it later

    Fix unit rates for the option period and index them to a published measure rather than leaving them to negotiation. A supplier who will not price production before the pilot is telling you what the renewal conversation will look like.

  4. Time-box the option, and let it lapse

    Give the option a window — typically six to twelve months from the gate — after which it expires. A perpetual option is a perpetual obligation on the supplier's pricing and will be priced accordingly, and a lapsed option is a legitimate reason to go back to the market.

  5. Make the gate review commercial as well as technical

    The gate paper should carry the harness scores, the integration evidence, the continuity diligence and the exit test in one place, signed by the commercial lead and the information manager together. Gate reviews attended only by the delivery team approve every pilot they see.

Vendor viability, continuity and running a portfolio

AI suppliers in this sector are young. The commercial stack that survives one of them failing, and the governance that starts at supplier four.

Vendor viability is an engineering requirement in this market, not a finance formality, because the supplier base is genuinely young and consolidating. The question to answer is narrow: if this company is acquired, pivots or ceases trading in eighteen months, what do we hold, and can we run it? Answering it properly changes what you ask for at tender — the contents of the escrow package, the change-of-control provision, the portability route — and it changes them before the relationship makes the conversation awkward.

The commercial stack, layer by layer

The commercial equivalent of a reference architecture. Each layer is annotated with the rung of the buying ladder that first requires it — an owner attempting rung 4 without the model and data schedule is running a rung-2 purchase with a longer term attached.

  1. Category strategy

    Stage 2+

    • Named AI category ownerOne person accountable across every AI contract
    • Capability roadmapWhat is being bought, in what order, over 24 months
    • Make-or-buy policyWhich data assets never leave the business
  2. Route to market

    Stage 2+

    • Pre-market engagementPublished, symmetrical, recorded
    • Procedure and noticeFramework, dynamic market or CFP — chosen first
    • Evaluation model and harnessWeights and scoring script fixed before bids
  3. Master agreement

    Stage 3+

    • Framework or master services agreementThe terms every call-off inherits
    • Data processing agreementLawful basis, retention, deletion evidence
    • Security and sub-processor scheduleNamed parties, jurisdiction, change notice
  4. Model and data schedule

    Stage 3+

    • Training-data provenance statementWhat the base model was built from
    • Performance floor and measurementIn your units, on a named dataset
    • Enriched-data ownershipDefined in the definitions clause, not a schedule
    • Drift and retraining dutySignal, owner, window, remedy
  5. Production call-off and continuity

    Stage 4+

    • Priced option to extendRates, volumes and window agreed at award
    • Escrow and change of controlWeights, code, features, harness — release tested
    • Exit plan and transition daysFormats named, price fixed at signature
  6. Portfolio governance

    Stage 5+

    • Vendor and model registerSupplier, model, data held, floor, renewal date
    • Overlap and concentration reviewA written ceiling with a named approver
    • Renewal and re-benchmark calendarOwned dates, not calendar reminders

Pipeline described

  1. Category strategy (stage 2+) — Named AI category owner: One person accountable across every AI contract; Capability roadmap: What is being bought, in what order, over 24 months; Make-or-buy policy: Which data assets never leave the business
  2. Route to market (stage 2+) — Pre-market engagement: Published, symmetrical, recorded; Procedure and notice: Framework, dynamic market or CFP — chosen first; Evaluation model and harness: Weights and scoring script fixed before bids
  3. Master agreement (stage 3+) — Framework or master services agreement: The terms every call-off inherits; Data processing agreement: Lawful basis, retention, deletion evidence; Security and sub-processor schedule: Named parties, jurisdiction, change notice
  4. Model and data schedule (stage 3+) — Training-data provenance statement: What the base model was built from; Performance floor and measurement: In your units, on a named dataset; Enriched-data ownership: Defined in the definitions clause, not a schedule; Drift and retraining duty: Signal, owner, window, remedy
  5. Production call-off and continuity (stage 4+) — Priced option to extend: Rates, volumes and window agreed at award; Escrow and change of control: Weights, code, features, harness — release tested; Exit plan and transition days: Formats named, price fixed at signature
  6. Portfolio governance (stage 5+) — Vendor and model register: Supplier, model, data held, floor, renewal date; Overlap and concentration review: A written ceiling with a named approver; Renewal and re-benchmark calendar: Owned dates, not calendar reminders
Step-by-step insights
Category strategy — one owner is the whole intervention
Almost every dysfunction on this page traces back to AI being bought by whichever function needed it, under whichever contract was to hand. Naming one category owner — usually in commercial, with a standing relationship to information management — does more than any policy document, because it creates a person who sees the fourth contract and recognises it as the third supplier doing the same thing. The role is not a gatekeeper. Its useful output is a published two-week route to buy, which is what stops teams routing around it.
Route to market — the layer chosen earliest and reconsidered least
The procedure is usually inherited: whatever framework the last technology purchase used. That inheritance is where most pilot-to-production failures originate, because a framework let for IT services frequently has no lot that contains an AI capability, and a lot that does not contain it cannot be called off against. Read the lot definitions before the requirement is written. If none fits, a dynamic market or a competitive flexible procedure is a smaller amount of work than discovering the gap after a successful pilot.
Model and data schedule — the layer that is not in your template
The master agreement, the data processing agreement and the security schedule almost certainly exist in your business already, in a form your legal team trusts. The model and data schedule does not, and it is the layer carrying provenance, the performance floor, enriched-data ownership and drift responsibility — four of the five things that decide whether this purchase ages well. Write it once, properly, and reuse it. It is roughly eight pages and it is the single highest-return document in AI procurement.
Escrow that actually releases something runnable
Software escrow releases source code, and source code is not what runs an AI system. The release package needs the model weights in a stated serialisation, the inference code, the training pipeline, the feature definitions, and the evaluation harness that proves the released version still meets the floor. It also needs an environment specification, because weights that only run inside the supplier's platform are a file rather than a capability. Exercise the release once, in a quiet quarter, and treat every gap as a contract defect to be fixed rather than a surprise to be noted.
Concentration limits and the warm second source
Consolidation is genuinely efficient and genuinely dangerous. An explicit ceiling — no single supplier holds more than a stated share of the asset record, or covers more than a stated number of categories — converts an emotional argument at renewal into a policy decision made in advance. Keeping a credible second source warm costs integration effort nobody enjoys funding, and it is the only thing that gives a renewal negotiation an alternative in it. Price it as insurance, not as duplication.
The renewal calendar as the portfolio's heartbeat
Every protection in the layers above is dated: the floor re-test, the benchmark, the break clause, the exit notice, the escrow verification. Dates without an owner lapse silently, and the first evidence of a lapse is usually a renewal quote. Put every date in one calendar with one named owner and a standing quarterly review, and measure one number — how long it takes to score a challenger against an incumbent. When that number rises above two weeks, the shared harness has stopped being maintained and the portfolio is drifting back to a collection of contracts.

Portfolio governance becomes necessary earlier than most owners expect — at supplier four rather than supplier ten. By the fourth AI contract a typical infrastructure owner has two suppliers extracting overlapping information from the same common data environment, four separate egress paths for the same site imagery, four sets of terms with four different definitions of customer data, and four renewal dates in four calendars. None of that is a crisis on any single day, and all of it is far cheaper to prevent than to unwind.

Likelihood: highImpact: high

The pilot that cannot be extended

A pilot succeeds, the sponsor is delighted, and the commercial team discovers there is no lawful or contractual route to buy it at scale. The programme either restarts the procurement — nine to twelve months, during which the site team keeps using the trial — or quietly abandons the capability. This is the single most common failure in public infrastructure AI buying.

PreventionWrite the route to production, in one paragraph, before the pilot is scoped. If it cannot be written, fund the pilot as research and say so.

Likelihood: highImpact: high

The enriched model leaves with the supplier

The engagement ends and the labels, corrected geometry and tuned weights go with it, because they were defined as service output rather than customer data. The next tender is then uncompetitive by construction: only the incumbent has a model that works on your conventions, and every challenger bids to build from scratch.

PreventionDefine enriched model data as customer data in the definitions clause, and specify the export formats and the delivery window.

Likelihood: mediumImpact: high

The supplier is acquired mid-framework

A specialist is acquired by a larger platform, the product is folded into a suite, the pricing model changes and the roadmap you bought disappears. The escrow agreement exists, has never been exercised, and turns out to omit the training pipeline and the feature definitions — so the released package cannot be run.

PreventionChange-of-control clause, escrow covering weights, code, features and harness, and one tested release before it is needed.

Likelihood: mediumImpact: medium

Six suppliers, six copies of the CDE

Each AI contract was sensible in isolation, and collectively they mean six external parties hold overlapping copies of the same project information under six different retention regimes. The exposure surfaces during a data protection review, a cyber incident or a handover, and unwinding it requires renegotiating six live contracts at once.

PreventionA vendor and model register consulted before every new purchase, with a written concentration and overlap rule.

Likelihood: highImpact: medium

The performance floor nobody can test

A floor was agreed in the contract, but the held-back package was never refreshed, the harness was not retained, and the person who wrote it has moved on. The floor is now unenforceable in practice, so degradation is argued about qualitatively and resolved by whoever has more patience.

PreventionKeep the harness and the test set as owned, version-controlled artefacts, and refresh the held-back package annually.

A 90-day plan: buying progress capture for one infrastructure package

The rung 2 → 3 move made concrete on one purchase — automated progress capture for a single motorway or station package, from requirement to a signed pilot with a priced production option.

Moving one rung takes about 90 days when it is scoped to a single purchase, and multiple years when it is scoped to a policy. To make it concrete, the plan below runs the move on a specific and very common infrastructure problem: automated progress capture from site imagery and scans for one package — a motorway section, a station box, a substation — where the weekly progress record is currently assembled by hand and disputed monthly. There is no model development in this quarter at all. Every day of it is requirement, evaluation and contract.

Rung 2 → rung 3 on one progress-capture purchase, in one quarter

One package, one requirement, one harness, one gate. If any phase runs long, narrow the scope — fewer asset classes, one week of held-back data instead of two — rather than extending the plan. The critical path is legal review of the model and data schedule, not evaluation.

  1. Days 1–20

    Write the requirement from the EIR, not from a demo

    Extract the requirement from the package's exchange information requirements: which information containers, at which stage, to what level of information need, delivered where in the CDE. Define the acceptance metric in your units — element-level completion accuracy against a manual survey on named asset classes. In parallel, name the production route in one paragraph: which framework, which lot, which option, which threshold.

    A testable requirement and a named route to production

  2. Days 21–45

    Prepare the held-back package and the scoring harness

    Withhold two weeks of real capture from the package, deliberately including the awkward cases: poor light, a temporary works arrangement, one asset class modelled by a different designer. Write the scoring harness as a script with a fixed metric and threshold, version it, and have the engineering owner and the information manager both sign it off before any bid arrives.

    A blind test set and a version-controlled harness

  3. Days 46–70

    Run pre-market engagement and the one-day bake-off

    Publish the engagement, shortlist five to eight suppliers, and pay each a modest equal fee to run the harness on the held-back package on the same day. Run viability diligence in parallel: accounts, runway, customer concentration, the actual contents of the escrow package. Score the evidence artefacts, not the answers.

    Scores on your data, and a viability read on each bidder

  4. Days 71–90

    Award with the model schedule signed and the option priced

    Award on the weighted matrix. Sign the model and data schedule — provenance, enriched-data ownership, the floor and its measurement, drift responsibility, the exit specification — before the pilot starts, and include a priced, time-boxed production option with unit rates and a volume mechanism. Book the gate review date on the day of award.

    A signed pilot with a production route already priced

The order matters

  1. Terms before pilot, never after

    Every clause on this page is routine to win before a pilot and hard to win after one. Before, you are one of several options to a supplier who is competing; after, you are an account with a tuned model, a changed site process and no alternative inside a year.

  2. Your data, your harness, one day

    The bake-off is worth more than every other evaluation activity combined, and it only works if the data is genuinely unseen, the harness is genuinely fixed and every bidder runs in the same window. Loosen any of the three and you are back to comparing presentations.

  3. Buy the option, not the promise

    A supplier's assurance that scaling will be straightforward is not a commercial instrument. A priced, time-boxed, conditional option is. It costs almost nothing at award, and it is the difference between an eleven-day extension and a nine-month re-procurement.

AI ITT readiness checklist

If you cannot tick all eight, the tender is a software tender with the word AI in the title. Tick as you go — this list works without JavaScript.

0 of 8 ticked

Tick honestly — the blank list is data too

Most owners can genuinely tick one or two of these, not zero. If none apply yet, don't start with a policy: take the next AI purchase already in your pipeline and run the 90-day plan above on that one. Every item on this list falls out of doing it once, and the second tender inherits the artefacts.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Enriched model data
The information created when an AI system processes your project data — labels, embeddings, corrected geometry, classified elements, extracted quantities. It exists only because your scheme paid for it, and a default vendor contract usually defines it as service output rather than customer data.
Held-back package
A genuine slice of your own project data — survey imagery, as-built information containers, RFI text — deliberately withheld from every bidder so that capability can be measured on unfamiliar work rather than demonstrated on tuned reference data.
Bake-off
A blind evaluation in which every shortlisted supplier runs the same scoring harness on the same held-back package in the same window. It is the single highest-signal activity in an AI procurement and the one most often skipped for time.
Performance floor
A contractual minimum quality expressed in the buyer's own units — detection recall on named asset classes, variance against measured quantity, false alarms per shift — with the measurement method and dataset stated. Distinct from an availability SLA, which measures only whether the system is running.
Drift responsibility
The contractual allocation of who detects performance degradation, on what signal, who pays to restore it, within what window, and what remedy applies if the floor is not recovered. Absent from almost every standard software agreement.
Model escrow
A deposit arrangement whose release package includes model weights, inference code, the training pipeline, feature definitions and the evaluation harness — not only source code. Escrow that omits the feature definitions releases a model that cannot be fed.
Exit and data portability
The obligation to return your assets in named, usable formats — IFC or COBie for model and asset data, open formats for labels, a stated serialisation for weights — within a defined window, at a price fixed at signature rather than quoted at exit.
Competitive flexible procedure
The buyer-designed procurement route introduced by the UK Procurement Act 2023, which may include negotiation and staged rounds. It is the route that most readily contains a shortlist, a paid bake-off and an award with a priced production option inside one lawful process.
Open framework
A framework agreement that can be reopened during its life so suppliers who were not present at the original competition can join. The structural answer to a supplier market in which the best AI vendor for a category may not exist when the framework is let.
Dynamic market
A register of qualified suppliers that applicants may join at any time, used to run short competitions among members. Useful for AI because it keeps the qualified list current without re-running a full framework competition.
Option to extend
A priced, time-boxed, conditional right to move a pilot into production, agreed at award with unit rates and a volume mechanism already in the schedule. Its absence is why most successful public-sector AI pilots are followed by a fresh procurement.
Vendor concentration risk
The exposure created when a single AI supplier holds too large a share of an owner's asset record or covers too many capability categories, so that an acquisition, a pricing change or an outage has no credible alternative behind it.

Frequently asked questions

The questions infrastructure owners and contractors ask most often when buying AI.

What is different about an AI RFP compared with a software RFP?

Four things a software RFP never asks for. Training-data provenance: what the model was built from and whether your project information will train anything shared. Ownership of the enriched model data — the labels, corrected geometry and extracted quantities the engagement produces. A performance floor in your units with named responsibility for drift, because AI degrades where software simply runs or fails. And an exit that returns something operable rather than a data dump. A software RFP that adds the word AI to its title contains none of these, and every one of them is argued about later at a worse moment.

Who owns the model a vendor trains on our project data?

Under most standard vendor terms, the vendor does. The base model was theirs, the tuning was performed by their tooling, and the resulting weights are treated as their product — so you hold a licence for the term and nothing afterwards. That position is negotiable and routinely conceded before a pilot, rarely after one. The workable middle ground is joint ownership or vendor ownership with an escrowed copy, the feature definitions and a tested portability route, so a change of supplier does not mean rebuilding from raw capture.

Can we extend a successful AI pilot into production without re-tendering?

Only if the route was designed in before the pilot. For a public owner that means the pilot ran inside an open framework, a dynamic market or a competitive flexible procedure that anticipated a production phase, or the contract carried a priced option agreed at award. If the pilot was an isolated below-threshold engagement, its success creates no lawful route to a production purchase of any size, and splitting the requirement into further small pilots to avoid competition is unlawful. Write the route down in one paragraph before scoping the pilot.

Should we build or buy AI for infrastructure delivery?

It depends on two variables, not on capability. Where the data is industry-standard and several credible suppliers exist — object detection on site imagery, document search over a contract set — buy, and re-tender often, because the underlying capability improves faster than any owner can track. Where the capability depends on your own accumulated programme record, such as outturn cost against estimate or your own defect taxonomy, no supplier holds that data and the correct answer is to build or to partner with the model kept on your side of the line.

How do we evaluate AI vendors fairly when every demo looks the same?

Stop being shown anything. Withhold a genuine package of your own data — two weeks of capture including the awkward cases — write a scoring harness before the ITT is issued, and require every shortlisted supplier to run that harness on that data on the same day. Pay a modest equal fee so small specialists stay in the process. Reported performance frequently halves against unfamiliar data, and the ranking is frequently different from what the presentations implied. Nothing else in an AI evaluation carries comparable signal.

What should an AI performance floor look like in a construction contract?

It should read like a specification, not an SLA. Name the metric in your units — recall on stated asset classes, variance against a manual survey, false alarms per shift — state the threshold, name the dataset it is measured against, and name the measurement method. Then allocate drift: who detects it, on what signal, who pays to restore it, in what window, and the remedy if it is not restored. An availability SLA tells you the system is running; only a floor tells you it is still right.

What happens if our AI vendor is acquired or fails?

That depends entirely on what your escrow package contains and whether anyone has tested it. Source-code escrow alone releases a repository you cannot run. A usable package includes the model weights in a stated serialisation, the inference code, the training pipeline, the feature definitions, the evaluation harness and an environment specification. Pair it with a change-of-control clause and exercise the release once in a quiet quarter. Suppliers in this market are young and consolidating, so treat continuity as a design requirement rather than a remote contingency.

How does ISO 19650 affect AI vendor contracts?

It gives you the language to write the requirement and the structure the outputs must land in. The exchange information requirements already state what information the project needs, in what form and at which stage, so an AI requirement drawn from them is testable and supplier-neutral. It also exposes the gap: enriched information containers are project information and belong in the common data environment with revision codes and suitability states, while the trained model is not an ISO 19650 container at all — which is exactly why it disappears at handover unless the contract names it.

How many AI vendors should an infrastructure owner run?

Fewer than arrive by default, and more than one in any category touching the asset record. Portfolio problems begin at supplier four, not supplier ten: by then most owners have overlapping extraction from the same CDE, several egress paths for the same imagery, and renewal dates in different calendars. Set a written concentration ceiling, consult a register before every new purchase, and keep one credible second source warm in each category that touches the asset record. The integration cost of that second source is insurance, not duplication.

What does a pilot-to-production contract structure look like?

One agreement, two phases, one signature. The pilot phase runs under the full model and data schedule. The production phase is drafted as a buyer's option, exercisable on stated measurable conditions — the floor met on the held-back package, integration demonstrated against the CDE, the exit test passed — with unit rates, a volume mechanism and a mobilisation window already agreed. Time-box the option to six to twelve months so it lapses rather than persisting. Exercising it becomes a gate paper and a purchase order instead of a new competition.

Does the EU AI Act change how we buy AI for infrastructure?

It changes what the tender has to establish. The framework classifies systems by risk and places obligations on providers and on organisations putting a system into service, so the classification question belongs in the ITT rather than in a later compliance exercise. Practically that means asking each bidder to state the intended purpose, the risk classification they assert and the evidence behind it, and requiring that the classification be maintained through changes. It sits naturally alongside an AI management system built to ISO/IEC 42001 or the NIST AI Risk Management Framework.

How long should an AI procurement take?

About 90 days for a single scoped purchase, if the route to market already exists. Roughly twenty days to write the requirement from the exchange information requirements, twenty-five to prepare the held-back package and the scoring harness, twenty-five for pre-market engagement and the bake-off, and twenty to award with the model schedule signed and the production option priced. The critical path is legal review of the model and data schedule rather than the evaluation. Where the route does not exist, add the time to establish a framework or dynamic market first.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for contractors and infrastructure owners — computer-vision progress and condition capture, quantity and schedule intelligence, and decision support wired into project controls and the common data environment. We are frequently on the other side of the table from the buyer, which is why this page is written from the evaluation harness outwards rather than from the pitch deck inwards.

  • · Production AI deployments with contractors and public infrastructure owners
  • · Evaluation harnesses and blind bake-offs run against clients' own held-back project data
  • · Delivery includes the commercial artefacts: model schedules, exit plans, escrow scope, portfolio registers
  • · 22 cited sources on this page

Sources

  1. ISOISO 19650-1 — organisation and digitisation of information about buildings and civil engineering works (opens in a new tab)
  2. ISOISO/IEC 42001 — artificial intelligence management system (opens in a new tab)
  3. UK BIM FrameworkUK BIM Framework (opens in a new tab)
  4. buildingSMART InternationalOpen standards for the built asset industry (opens in a new tab)
  5. NIBSNational Institute of Building Sciences (opens in a new tab)
  6. NISTAI Risk Management Framework (opens in a new tab)
  7. UK Government (DSIT / Office for AI)Guidelines for AI procurement (opens in a new tab)
  8. UK GovernmentAI Playbook for the UK Government (opens in a new tab)
  9. UK Government (Cabinet Office)Procurement Act 2023 guidance documents (opens in a new tab)
  10. UK Government (Cabinet Office)Procurement policy notes (opens in a new tab)
  11. UK Government (Cabinet Office)The Construction Playbook (opens in a new tab)
  12. UK GovernmentFind a Tender Service (opens in a new tab)
  13. Crown Commercial ServiceCrown Commercial Service (opens in a new tab)
  14. European CommissionPublic procurement in the single market (opens in a new tab)
  15. European CommissionRegulatory framework for AI (opens in a new tab)
  16. Information Commissioner's OfficeArtificial intelligence and data protection (opens in a new tab)
  17. National HighwaysDigital Roads (opens in a new tab)
  18. National HighwaysSuppliers and commercial frameworks (opens in a new tab)
  19. Network RailProcurement Act 2023 — what it means for suppliers (opens in a new tab)
  20. SkanskaSkanska USA Building establishes Digital Transformation and Solutions Team (opens in a new tab)
  21. SkanskaCurbing equipment emissions through artificial intelligence (opens in a new tab)
  22. McKinsey & CompanyReinventing construction through a productivity revolution (opens in a new tab)

Buy AI the way you buy everything else on a major project

We run the assessment with your commercial, information-management and delivery leads, mark up a live or upcoming AI tender against the clause table on this page, and leave you with a costed 90-day plan for your weakest dimension. You keep the markup and the plan whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.