Construction & InfrastructureAI Adoption & Maturity Curve
AI adoption KPIs in construction and infrastructure: the denominator problem, and how to fix it
AI adoption KPIs are the metrics that prove an AI programme is changing how a construction business builds, not merely that software has been bought. In construction they fail for one structural reason: every project is a one-off, so a number with no agreed denominator, baseline or comparison cohort cannot be defended.

Key takeaways
- A construction AI adoption KPI is only as good as its denominator. Absolute claims — hours saved, documents processed, queries closed — cannot be compared between a £3m fit-out and a £300m viaduct, so they die at the first commercial challenge. Normalise per £m of certified value, per 1,000 programme activities, per 100,000 worked hours or per inspection, and say which.
- Construction already owns one properly built KPI family: safety. RIDDOR-reportable rates are expressed per 100,000 worked hours, over a defined period, for a defined population. Copy that grammar — numerator, denominator, population, period, exclusions — for every AI adoption KPI you intend to report twice.
- The baseline must be taken before the tool is switched on, from records that already exist in the CDE, the cost ledger and the programme. A baseline reconstructed afterwards inherits the pilot's own effects, and on a project with a re-baselined programme it cannot be reconstructed at all.
- Attribution needs a counterfactual, and in construction that is a matched cohort rather than a holdout: comparable packages on the same framework, same phase, same designer, left on the previous process. Without one, weather, design maturity and subcontractor mix will take the credit or the blame.
- Five measurement stages — Anecdotal, Counted, Baselined, Normalised, Assured — decide which KPIs you may honestly report. Most firms sit at Counted, reporting licence utilisation and documents processed, which is activity data rather than adoption data and never survives a renewal conversation.
Abbreviations used on this page
- CDE
- Common data environment (the ISO 19650 information container)
- EIR
- Exchange information requirements — what the client asks for, contractually
- BEP
- BIM execution plan — how the supply chain says it will deliver the EIR
- AIR
- Asset information requirements — what the operator needs at handover
- RFI
- Request for information (US/global usage)
- TQ
- Technical query — the UK civils equivalent of an RFI
- ITP
- Inspection and test plan
- NCR
- Non-conformance report
- CVR
- Cost value reconciliation — the monthly commercial truth on a project
- EVM
- Earned value management, and its indices SPI (schedule) and CPI (cost)
- NRM
- RICS New Rules of Measurement — the industry's standard measurement basis
- RIDDOR
- Reporting of Injuries, Diseases and Dangerous Occurrences Regulations
Free · 8 questions · ~3 minutes
Score how well you can actually measure AI adoption
Eight questions, one at a time, about three minutes. They score the measurement system rather than the technology: whether your AI numbers have a denominator, a baseline, a counterfactual and an owner. Answer them and we build your personalised report — your stage on the measurement ladder, your score on each of the four dimensions, and the specific fix that unlocks the next stage — and send it to your inbox.
0 of 8 answered
Pick an option to continue
Report ready
Your personalised measurement report is ready
Tell us where to send it. Your stage appears on screen straight away, and the full report — dimension scores, the normalisation bases that suit your portfolio mix, and a 90-day plan for your weakest dimension — arrives in your inbox.
Your result
Your full report is on its way to your inbox.
Stage 1 · Anecdotal
AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
Your next moveWrite one metric definition sheet — numerator, denominator, exclusions, owner — for one decision on one live package, and take the before-figure from the CDE.
Stage 2 · Counted
Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
Your next moveChoose one decision, define the outcome KPI in the project's own units, and capture the pre-AI baseline before the next project mobilises.
Stage 3 · Baselined
One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
Your next moveAgree the normalisation base and the classification the KPI is expressed in, then reproduce it on a second, structurally different project.
Stage 4 · Normalised
The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
Your next moveMake the KPI recomputable by someone outside the team: version the definition, retain the inputs, and build the evidence pack.
Stage 5 · Assured
The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.
Your next movePut the KPI set on the management system's review cycle: annual definition review, a change log, and one live recomputation test a year.
0 / 24
Denominator discipline
— / 6
Baseline & counterfactual
— / 6
Instrumentation & recomputability
— / 6
Assurance & decision rights
— / 6
Your score maps to a stage on the measurement ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the credibility of every number you report, and it is where the next fortnight of work belongs. Your lowest-scoring dimension is —, and that is where the next investment belongs.
Your score maps to a stage on the measurement ladder. The dimension breakdown matters more than the total: the lowest dimension is what actually caps the credibility of every number you report, and it is where the next fortnight of work belongs.Your four dimensions score evenly, so there is no single weak link to attack — follow the stage’s next move above rather than picking a dimension.
Want your reported figures independently recomputed?
We take one AI benefit figure you have already published internally, ask for the definition and the records behind it, and attempt to reproduce it. You get the recomputation, the gaps we hit, and a definition sheet you can use on the next project — whether or not we work together.
How the score maps to a stage
- 0–4 — Stage 1, Anecdotal. AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
- 5–9 — Stage 2, Counted. Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
- 10–14 — Stage 3, Baselined. One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
- 15–19 — Stage 4, Normalised. The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
- 20–24 — Stage 5, Assured. The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.
What AI adoption KPIs are in construction — and why they break
A definition, the two KPI families, and the path a number travels from a site event to a board pack — including the shortcut that makes most construction AI numbers unusable.
AI adoption KPIs in construction are the metrics that prove an AI capability has changed how a project is delivered, rather than proving that software has been bought and logged into. They come in two families that only mean something together: adoption KPIs, which measure whether the capability has entered the work — share of technical queries triaged, share of inspections captured through the tool, time from site event to a usable answer — and value KPIs, which measure what changed in the units the project already runs on: days on the critical path, non-conformances per 1,000 inspections, cost of rework, disputed variation value, RIDDOR-reportable incidents per 100,000 worked hours.
The reason these KPIs fail in construction and not, say, in a factory is structural. A production line offers a stable denominator — units, shifts, machine hours — that persists for years. A construction business is a portfolio of temporary organisations, each with its own package breakdown, its own designer, its own subcontract mix, its own weather, and a programme that will be re-baselined at least once. A number computed on one job therefore has no natural counterpart on the next, and the honest answer to 'is our AI adoption improving?' is unavailable until somebody chooses a base, a comparison group and a set of exclusions and writes them down.
Which KPIs you can honestly report depends on where you sit on the measurement ladder set out below. A firm at Counted can report usage and nothing else; a firm at Baselined can report a delta on one project with caveats; only a firm at Normalised can put a portfolio figure in a bid without inviting a challenge it cannot answer. Reporting above your stage is not ambition, it is the mechanism by which a programme loses credibility — and it is much more common than under-reporting.
Credible measurement against time on the ladder
What is plotted here is not value delivered but value that can be defended. It stays close to flat through Anecdotal and Counted — where most construction firms are — and inflects at Baselined, the first stage at which a number survives a commercial challenge. This is why firms with genuinely useful AI tools often cannot show a single defensible figure.
Benefit that survives a commercial challenge by stage
- Stage 1 · Anecdotal — 34% of operators. AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
- Stage 2 · Counted — 38% of operators. Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
- Stage 3 · Baselined — 19% of operators. One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
- Stage 4 · Normalised — 7% of operators. The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
- Stage 5 · Assured — 2% of operators. The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.
Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with the UK construction industry KPI programme.
How an AI adoption KPI is manufactured on a construction project
The same site event can produce two completely different numbers. The top lane is the path most reported construction AI figures actually take — through the vendor's dashboard and into a slide, arriving at the board pack with its definition stripped off. The middle and bottom lanes are what it costs to make the number checkable and then portable.
- Data & feeds
- AI / model
- Where value leaks
- Human in the loop
- System-of-record action
The process, in words
- In the site and package lane, an event — a technical query raised, a pour signed off, a snag logged — is picked up by the AI tool, which produces a flag, a score or a drafted response. The tool counts its own work on its own dashboard, using a numerator and a base the vendor chose, and somebody screenshots the result into the monthly pack. The number arrives at the board with its definition stripped off, which is why it cannot be recomputed and does not survive a challenge.
- In the project-controls lane, the same site event is written to records that already exist — the CDE, the cost ledger, the programme. A metric definition sheet fixes the numerator, the base, the exclusions and the owner; a scheduled job recomputes the figure from retained records; and the result lands in the monthly review and the CVR, a forum with the authority to change scope, spend or sequence.
- In the portfolio and assurance lane, the definition is expressed in a normalisation base that holds across projects and compared against a matched cohort of packages left on the previous process. What comes out is an attributed delta rather than a raw improvement, and the evidence pack behind it — definition version, retained inputs, recomputation route — is what makes the figure usable in a tender or a framework performance review.
- The dashed red arrow is the shortcut nearly every construction business takes at least once: the vendor's dashboard number promoted straight into the assured position without a base, a baseline or a trail. It is fast, it looks identical in a slide, and it is the single reason most construction AI benefit claims cannot be repeated a year later.
Step-by-step insights
- The site event — one event, two registers, two numbers
- The most under-appreciated fact about construction measurement is that the same event is usually recorded twice: once in the tool and once in the contractual register. A technical query exists in the AI assistant's log and in the TQ register in the CDE, with different timestamps, different closure semantics and sometimes different identities, because the tool created a thread that the register treats as one query with two responses. Any KPI you intend to defend has to be computed from the contractual register, with the tool's log used only to establish which records it touched. The moment the tool's log becomes the numerator, the number has left the world the contract operates in.
- The vendor dashboard — a marketing instrument doing a management job
- Vendor dashboards are not dishonest; they are built to demonstrate that the product is working, which is a different question from whether the project is better off. Their base is almost always the tool's own activity — items processed, items flagged, users active — and their period is the licence period rather than the reporting period. They are perfectly good as an operational health check and completely unusable as a KPI, because the numbers cannot be reproduced from records you retain, and the definition can change with a product release. Treat them the same way you would treat a subcontractor's own productivity claim: useful signal, not evidence.
- The definition sheet — where the arguments belong
- The definition sheet is one page and it is where the difficult conversations happen. Does a query withdrawn on the day it was raised count. Is a query bounced back for missing information the same query or a new one. Does closure mean the designer's formal response or site acknowledgement of it. Do weekends count. These look like pedantry and they are the whole of the measurement: two honest people computing from the same register with different exclusion rules will produce numbers ten or fifteen per cent apart, and every subsequent argument about the AI tool will actually be an argument about the exclusion rules. Settle them once, in writing, with the commercial team in the room.
- Scheduled recomputation — the difference between a metric and a memory
- The reporting job should be rerunnable by someone who has never met the project, over retained extracts, producing the same answer. That has two consequences people underestimate. First, it forces retention: if the CDE export is overwritten each month, last quarter's figure can never be checked. Second, it forces the definition into code, where a change is visible in a diff rather than in someone's habit. Construction firms that have done this once usually find the first recomputation disagrees with the originally reported number, and that discovery — early, internal, harmless — is exactly the point of building it.
- The normalisation base — chosen per domain, never universally
- The base has to reflect what actually drives the numerator. Information-intensive work — design queries, technical submittals, drawing revisions — scales with design packages and model elements, not with contract value. Commercial workload scales with certified value and the number of variations. Production scales with programme activities. Safety already has its base: incidents per 100,000 worked hours, which is the grammar used in RIDDOR-based industry reporting and the reason safety is the one construction KPI family that survives comparison between firms. Copy that grammar per domain and write down why each base was chosen; a base with no stated rationale gets changed by the next person who finds it inconvenient.
- The matched cohort — construction's substitute for a holdout
- You cannot randomise a viaduct, so attribution in construction is done by matching rather than by holding out. Match on the variables that drive the metric — scheme type, contract form, delivery phase, designer, whether the design was novated, approximate value band — and keep a written register of what you could not match on, typically weather, ground conditions, design maturity at contract award and subcontractor mix. Read the register out with the number. A delta presented with its confounders is believed; the same delta presented as a clean result invites the room to invent the confounders themselves, and they will find bigger ones than you did.
The five stages of the measurement ladder
Anecdotal, Counted, Baselined, Normalised, Assured — what each looks like on a live project, the signals a reviewer can check in an afternoon, and the anti-pattern that traps firms there.
The ladder below measures the measurement system, not the technology — a firm running sophisticated computer vision on every site can sit at Counted, and a firm with one language model and a disciplined quantity surveyor can sit at Baselined. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own registers this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Anecdotal
34% of operators sit here
AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
Anecdotal is the default state, and it is not a sign that nothing is happening. On most sites something genuinely is: an engineer is using a language model to draft technical query responses, a planner is testing an assistant against the four-week look-ahead, a quantity surveyor is pulling clauses out of a subcontract faster than they used to. The work is real. What does not exist is any artefact a second person could use to verify it.
The failure mode is specific and it is commercial rather than technical. A remembered number has no numerator anybody agreed, no denominator at all, and no before-figure. When the commercial manager asks the obvious question — over how many queries, compared with what, on which package — there is no answer, and the claim is quietly withdrawn. That withdrawal is what teaches a project board that AI benefits are soft, and it is much harder to undo than to avoid.
This stage is cheap to leave. It does not require a platform, a data strategy or a business case; it requires one page of writing. Pick a single decision on a single live package, write down what you are counting, what you are dividing it by, what you are excluding and who owns the number, and take the before-figure from records that already exist. That page is the whole difference between Anecdotal and the stage above it.
In practice
The three-hour claim at the monthly review
On a highways widening scheme, a site engineer tells the monthly review that the AI document assistant saves him about three hours a week finding clauses and specification references. The project director likes it and repeats it to the client. The commercial manager asks two questions: three hours against what, and across how many technical queries? Nobody has the TQ register open, and nobody knows what last quarter's median closure time was. By the following month the claim has been dropped from the pack — not because it was untrue, but because it could not be checked.
What it looks like
- Benefit is quoted in remembered time — 'about half a day a week'
- No written definition of what is being counted, or of what it is divided by
- The only record is a slide in the digital lead's deck
- The claim does not survive project closeout or a change of project director
Diagnostic signals you can check this week
- Ask for the definition of any AI benefit number quoted in the last board pack. If it exists only in a slide, you are here
- Ask what the figure is divided by. Silence, or 'per project', means there is no denominator
- Ask what the before-figure was and where it came from. 'We think it was about…' is the tell
- Check whether any AI claim survived a project handover into the next job's bid or lessons-learned pack
Anti-pattern · Commissioning a benefits study
The instinct at this stage is to buy a business-wide benefits assessment — a consultant, a survey of project teams, a modelled figure for the group. It produces a number with a decimal point and no provenance, and the first framework client or auditor who asks how it was derived collapses it. Worse, it teaches the organisation that AI value is estimated rather than measured. Write one metric definition sheet for one decision instead; it costs a morning and it is the only artefact that compounds.
What holds you here
Nothing is written down, so no claim can be checked by a second person — and an unchecked claim is worth nothing commercially.
Highest-leverage next move
Write one metric definition sheet — numerator, denominator, exclusions, owner — for one decision on one live package, and take the before-figure from the CDE.
Cost of leaving
- Effort
- 2–3 months
- Team
- Digital lead plus half a day a week from a quantity surveyor or planner
- Risk
- Low — the work is definitional and nothing operational depends on it yet
- To next stage
- 2–3 months
If this is you, the next step is
A half-day workshop with your commercial and digital leads on one live package.
Stage 2
Counted
38% of operators sit here
Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
Counted is where most construction businesses actually sit, and it is a genuine improvement on Anecdotal: the numbers exist, they are produced on a schedule, and they come out of a system rather than a memory. Licence utilisation, active users per project, documents processed, images analysed, drafts generated — all of it is real telemetry. It is simply telemetry about the tool rather than about the business.
The structural problem is that a count has no denominator and no outcome, which makes it unusable in exactly the two conversations that matter. It cannot be compared between projects, because a 400-unit residential job will always process more documents than a substation upgrade; and it cannot be defended at renewal, because nothing in the count tells the commercial director whether a date moved or a cost changed. Counts also invite the worst possible intervention: an adoption campaign that raises usage without raising value, after which the count is not merely uninformative but actively misleading.
Leaving this stage is not a matter of counting better. It is a matter of choosing one decision the tool is supposed to change, naming the outcome that decision moves in the units the project already runs on — days, pounds, non-conformances, worked hours — and taking the before-figure before the next project starts. One decision measured properly outranks forty projects counted.
In practice
The 78% licence utilisation slide
A tier-one contractor reports 78% seat utilisation for its AI document platform across forty live projects, up eleven points on the quarter, with a chart of documents processed per month. At the renewal meeting the commercial director asks what the eleven points bought. The digital team can show that usage rose fastest on the three projects with the most rework, which is either evidence that the tool goes where the pain is, or evidence that the tool creates work — and nothing in the reporting distinguishes the two. The renewal goes through on relationship rather than on evidence, which is the point at which the programme becomes politically fragile.
What it looks like
- Monthly digital report leads with seat utilisation and active-user counts
- Volume metrics — drawings scanned, queries drafted, images processed
- Counts are absolute, so a big project always looks like a success
- No commercial or delivery metric appears next to any of it
Diagnostic signals you can check this week
- Open the last AI report. If every metric is a count of activity rather than a rate against a base, you are here
- Ask whether any AI metric appears next to CPI, SPI, NCR counts or accident frequency rate in the same pack
- Check whether the number would go up if the project simply got bigger. If it would, it is a count, not a KPI
- Ask what decision changed as a result of last quarter's report. If the answer is a renewal, the reporting is serving the vendor
Anti-pattern · Running an adoption campaign to make the count go up
When usage plateaus, the standard response is a push — training, mandates, league tables between projects. Usage duly rises, and the relationship between usage and outcome, which was never established, becomes unrecoverable: you can no longer tell whether high-usage projects use the tool because they are well run or because they are struggling. Any incentive attached to a count will be met, and the data it produces afterwards is worse than the data you had before. Define the outcome metric first, then let usage be whatever it needs to be.
What holds you here
The count has no denominator and no outcome, so it cannot be compared between projects or defended when the licence comes up for renewal.
Highest-leverage next move
Choose one decision, define the outcome KPI in the project's own units, and capture the pre-AI baseline before the next project mobilises.
Cost of leaving
- Effort
- 3–6 months
- Team
- Digital lead, a commercial manager, a planner, and one project willing to be first
- Risk
- Medium — renewal and rollout decisions are being taken on the wrong number in the meantime
- To next stage
- 3–6 months
If this is you, the next step is
We pick the decision, define the metric and find the before-figure in your existing records.
Stage 3
Baselined
19% of operators sit here
One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
Baselined is the first stage at which a number can survive a challenge. The definition is written, the before-figure came from a system rather than a recollection, and the current figure is produced by rerunning a query over records that are retained. If the project director is asked in six months where the figure came from, the answer is a script and a record set, not a person's diary.
The character of the work here is unglamorous and mostly commercial. The arguments are about exclusions — does a technical query raised and withdrawn on the same day count; is a query bounced back for missing information a new query or the same one; does a Saturday count as a working day for closure time. Those arguments feel like pedantry and they are the entire value of the stage: an unstated exclusion rule is the most common reason two people compute different numbers from the same register and both are honest.
The constraint that ends this stage is comparability. The number is true for this project, and it says nothing about the next one, because activity identifiers, cost codes, query classifications and package breakdowns differ between jobs — often between phases of the same job. A firm that stops here has one well-measured project and no way to aggregate, which means the KPI cannot inform an investment decision or a bid.
In practice
The technical-query baseline on a viaduct package
Before switching on AI-assisted TQ triage on the viaduct substructure package, the project-controls lead exports eighteen months of the TQ register from the CDE, computes the median and 90th-percentile closure time in working hours, and writes down the rules: withdrawn queries excluded, queries reopened within five days treated as the same query, closure measured to the designer's formal response rather than the site acknowledgement. The current figure is recomputed weekly by a scheduled job. When the number improves, the argument at the review is about whether the design was simply more mature this quarter — a real question, and one the project cannot yet answer.
What it looks like
- A metric definition sheet exists: numerator, denominator, population, period, exclusions, owner
- The baseline was taken from the CDE, the cost ledger or the programme before go-live
- The KPI is recomputed on a schedule, from raw records, by a job rather than a person
- It appears in the monthly project review alongside the commercial numbers
Diagnostic signals you can check this week
- Ask to see the metric definition sheet. If the exclusions section is blank, the definition is not finished
- Ask who could rerun the calculation if the digital lead were on leave, and whether the inputs are retained
- Check whether the baseline predates the tool's go-live date, from evidence rather than assertion
- Ask whether the same number could be produced for a second project without new hand-mapping
Anti-pattern · Mandating it across the portfolio before it has been keyed twice
One clean measurement produces immediate pressure to make it a corporate KPI. It gets mandated across thirty projects that classify technical queries differently, use different work-breakdown levels and run three different CDEs, and within a quarter the group figure is an average of incompatible things. Prove the definition on a second, deliberately different project first — different sector, different client, different designer — and fix what breaks. The second project is where you learn which parts of the definition were actually project-specific assumptions.
What holds you here
The number is true for one project and comparable with nothing, because identifiers, classifications and package breakdowns differ across jobs.
Highest-leverage next move
Agree the normalisation base and the classification the KPI is expressed in, then reproduce it on a second, structurally different project.
Cost of leaving
- Effort
- 4–8 months
- Team
- Project-controls lead, a data engineer part-time, commercial sign-off on the definition
- Risk
- Medium — the definitional arguments are with the commercial team and cannot be skipped
- To next stage
- 4–8 months
If this is you, the next step is
We take your existing definition to a deliberately different job and report what breaks.
Stage 4
Normalised
7% of operators sit here
The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
Normalised is the stage at which a construction business can finally answer the question it has been asked since stage one: is this working, across our portfolio, and by how much. That becomes possible only when two things exist together — a base that survives differences in project size and type, and something to compare against that is not simply the last job.
The base is domain-specific and choosing it is real work. Design-side metrics normalise well per design package or per thousand model elements; commercial metrics per £m of certified value; production metrics per thousand programme activities; safety metrics per 100,000 worked hours, which is the base the industry already uses under RIDDOR reporting. One base applied to everything is the classic error: divide a fit-out and a viaduct by contract value and the viaduct will look efficient at everything, because value per unit of information is completely different in the two.
The counterfactual in construction is a matched cohort rather than a holdout, because you cannot randomise a bridge. Matching is on the variables that actually drive the metric — scheme type, contract form, phase, designer, client, and whether the design was novated — and the confounders you cannot match on go in a register that is read out with the number. Weather, design maturity and subcontractor mix will otherwise take the credit or the blame, and everyone in the room will know it.
In practice
The cohort that made the framework case
A civils contractor on a five-year water framework runs AI-assisted technical-query triage on six schemes and leaves six comparable schemes on the previous process — matched on AMP delivery phase, designer, contract form and approximate value band. The KPI is reported as median TQ closure hours and TQs reopened per 100 raised, both per scheme, with a confounder note on two schemes affected by a wet winter. The reported delta is smaller than the pilot claimed and it is believed, which is what allows the client to write the measure into the next framework's information requirements.
What it looks like
- Every KPI carries a normalisation base — per £m certified, per 1,000 activities, per 100,000 worked hours
- A matched cohort exists: comparable packages or schemes left on the previous process
- A re-baseline register records every programme and scope change that breaks a series
- The KPI sits in the portfolio pack next to CPI, SPI and accident frequency rate
Diagnostic signals you can check this week
- Ask what base each KPI is divided by, and whether the same base is used across sectors. One base everywhere is a red flag
- Ask to see the cohort definition and the matching variables, in writing
- Ask what happens to the KPI series when a programme is re-baselined. If nothing happens, the series is silently broken
- Check whether the confounder register is read out with the number or filed separately
Anti-pattern · Normalising everything by contract value
Contract value is available for every project, which is exactly why it gets adopted as the universal base — and it is wrong for most metrics. Information intensity does not scale with value: a £5m hospital fit-out can generate more technical queries than a £150m earthworks package. Using value alone makes complex low-value work look like the worst performer in the portfolio and rewards teams for being on big, information-light jobs. Choose the base per domain, write down why, and review it when the portfolio mix changes.
What holds you here
The numbers are credible internally but carry no assurance trail, so they cannot be used in a bid, an insurance discussion or a client audit.
Highest-leverage next move
Make the KPI recomputable by someone outside the team: version the definition, retain the inputs, and build the evidence pack.
Cost of leaving
- Effort
- 9–15 months
- Team
- Project controls, commercial analytics, a data engineer, and a portfolio owner with the authority to standardise
- Risk
- Higher — standardising classifications across live projects is change management, not engineering
- To next stage
- 9–15 months
If this is you, the next step is
Two workshops: one on bases per domain, one on cohort matching against your live portfolio.
Stage 5
Assured
2% of operators sit here
The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.
Assured is narrow, and it should be. It does not mean every AI metric in the business is audited; it means a small, named set of KPIs has been promoted to the same status as certified value or reportable-incident rates. Those are the numbers that appear in a prequalification questionnaire, in a framework performance review, or in an insurer's questions about how design-review decisions are being made — and each one carries a definition, a version, retained inputs and a recomputation route.
The engineering at this stage is largely done. The work is management-system work: an owner, a change log, a review cycle, a retention rule, an evidence pack that a person who has never met the project can follow. Firms already running an ISO 19650 information management function and, increasingly, an AI management system in the shape of ISO/IEC 42001 have the machinery for this; the KPI set simply becomes another controlled artefact inside it.
The characteristic failure of Assured is not collapse but slow decay. Definitions age while the operation moves, the portfolio mix shifts and the base stops fitting, personnel change and the recomputation script stops being run. Assurance is a cycle rather than a state: an annual definition review, a change log that says why each version changed, and one live recomputation test a year against a real project's records.
In practice
The recomputation on a framework audit
A client's assurance team picks one reported figure — non-conformances raised per 1,000 inspections on the AI-supported quality process — and asks the contractor to reproduce it. The contractor supplies the definition at the version in force that quarter, the extract of the inspection and NCR registers with record identifiers, and the script. The recomputed figure differs by under two per cent, explained by three records reclassified after the report date, and the difference itself is documented. That exchange takes four days and is the reason the measure is accepted in the next tender.
What it looks like
- Definitions are versioned and referenced in the EIR and the BEP, not held in a spreadsheet
- Inputs are retained with lineage for the full contractual retention period
- Internal audit recomputes a sample and the result is reproducible within tolerance
- The figures appear in tender submissions and withstand client and insurer questions
Diagnostic signals you can check this week
- Ask whether the KPI definitions are version-controlled with a change log, and who approves a change
- Ask whether the inputs behind last year's reported figure are still retrievable today
- Ask when the recomputation test was last run, and by whom outside the reporting team
- Check whether the KPI set is referenced anywhere contractual — the EIR, the framework performance schedule, a prequalification response
Anti-pattern · Freezing the definition to protect the trend
Once a KPI has two years of history, changing it breaks the series, and the reflex is to leave it alone. The operation moves anyway: the portfolio shifts toward frameworks, the classification changes, a new CDE arrives. The metric slowly stops describing the business while continuing to look consistent, which is worse than an obvious break. Version the definition, restate the affected periods where you can, and record the discontinuity — a documented break is defensible, a silent drift is not.
What holds you here
Assurance decays quietly — definitions age, the portfolio changes shape, and last year's evidence pack no longer satisfies this year's audit.
Highest-leverage next move
Put the KPI set on the management system's review cycle: annual definition review, a change log, and one live recomputation test a year.
Cost of leaving
- Effort
- Continuous
- Team
- Information management, internal audit, commercial, with a standing review forum
- Risk
- Concentrated — low frequency, high consequence, contractual and reputational in nature
If this is you, the next step is
We attempt to recompute a KPI you have already published, and report exactly where it breaks.
Two things are worth noticing about the shape of this ladder. The first is that the biggest single gap is between Counted and Baselined, and nothing in it is technical: it is a page of definitions, an extract taken before go-live, and a commercial manager willing to argue about exclusions. The second is that the top two stages are governed by documents rather than by systems. Normalised depends on a base and a cohort definition; Assured depends on version control, retention and a recomputation route. A firm can climb this ladder without buying anything.
Where construction and infrastructure firms actually sit
The distribution across the ladder, why Counted is the mode, and the external evidence that the industry's measurement problem predates AI entirely.
Most construction and infrastructure firms sit at Counted. They can tell you how many licences are active, how many documents the model has processed and how usage compares between regions, and they cannot tell you what any of it did to a date, a cost or a non-conformance. A minority have one project with a real baseline, and a very small number can put a normalised, cohort-compared figure in front of a client.
Distribution of construction firms across the measurement ladder
Counted is the mode and the plateau. The drop from Counted to Baselined is the largest transition loss on the ladder, and it is a documentation gap rather than a technology gap.
Share of firms
- 34% — 1 · Anecdotal
- 38% — 2 · Counted (the plateau)
- 19% — 3 · Baselined
- 7% — 4 · Normalised
- 2% — 5 · Assured
The distribution above is a model-derived illustration rather than a survey, and it is anchored to three external bodies of evidence about how construction measures itself. The UK industry has published normalised headline construction KPIs through Constructing Excellence (opens in a new tab) since the late 1990s, which is why safety, client satisfaction and predictability have a shared grammar and almost nothing else does. The Infrastructure and Projects Authority (opens in a new tab) reports annually on the delivery confidence of the government's major projects, and the National Audit Office (opens in a new tab) has repeatedly found that infrastructure performance claims are hard to compare because baselines move. Every one of those findings applies to AI benefit claims without modification.
5%
Of project cost is directly measurable avoidable error, before indirect and latent costs, in GIRI's published research
Get It Right Initiative
per 100k
Worked hours — the base the industry already uses for RIDDOR-reportable incident rates, and the model for every AI adoption KPI
HSE
1999
The UK construction industry has published normalised headline KPIs annually since the late 1990s
Constructing Excellence
It is worth being precise about what this distribution does and does not say. It is not a claim that two thirds of the industry gets no value from AI; plenty of firms in the Anecdotal and Counted bands have tools that genuinely help engineers and quantity surveyors every day. It is a claim about evidence: that the value is invisible to the commercial function, absent from the cost and output statistics (opens in a new tab) the sector reports, and therefore unavailable in the two conversations where it would matter most — the investment case for scaling, and the tender in which a client asks what your digital capability is actually worth to them.
The KPI specification: which metrics, in which units, from which record
The honest KPI set for each stage, the decision map that says where each metric comes from, and the eight tests a number must pass before it enters a board pack.
The right AI adoption KPI is the one your stage can actually support, computed from a register that already exists and expressed in a base that survives the next project. That gives three requirements, in order: pick metrics your evidence can carry, source each from a named system of record rather than from the tool, and divide each by something that does not change when the job gets bigger. The tables below are the working specification — the first says what each stage may honestly claim, the second says where every metric comes from in a construction estate.
| Stage | Adoption KPIs — has it entered the work? | Value KPIs — did anything change? | The claim this stage cannot support |
|---|---|---|---|
| 1 · Anecdotal | None that are checkable. At best, a written list of which decisions the tool touches | None. The only honest output is a pre-AI baseline extracted from the CDE, the cost ledger and the programme | Any benefit figure at all — including a conservative one |
| 2 · Counted | Licence and active-user counts, items processed, coverage of a register (share of TQs the tool touched) | None attributable. Value talk at this stage is projection, and should be labelled as such | That usage growth means adoption, or that high-usage projects perform better |
| 3 · Baselined | Coverage against a defined population, median time from event to usable answer, share of outputs accepted without rework by the responsible engineer | One named delta against a pre-AI baseline on one project — TQ closure hours, NCRs per 1,000 inspections, days to assess a variation | That the project delta generalises to the portfolio |
| 4 · Normalised | Coverage and acceptance normalised per domain base, reported per scheme and per cohort | Attributed deltas against a matched cohort — rework cost per £m certified, SPI at package level, disputed variation value per £m | That the figure is audit-ready, or usable in a contractual performance schedule |
| 5 · Assured | The same measures under version control, with retention and a recomputation test on record | Attributed deltas with confounders registered, restated across definition changes, referenced in tender and framework reporting | That assurance holds without an annual definition review — it decays quietly |
Two disciplines make the whole specification trustworthy, and both are unpopular. The first is that adoption and value KPIs are always reported as a pair: a coverage figure with no delta is theatre, and a delta with no coverage figure is unexplainable, because nobody can tell whether it came from the tool or from the two packages where the tool was never used. The second is that every value KPI names its comparison in the same sentence as the number. 'TQ closure time down 22%' is a claim; 'median TQ closure down from 61 to 47 working hours on the six AI-supported schemes, against 58 to 55 on the six matched schemes' is a measurement.
| Operating domain | Decision AI touches | System of record | KPI it moves | Normalisation base | Honest from |
|---|---|---|---|---|---|
| Design & engineering | Technical query triage, drafting responses, clash and change review | CDE (ISO 19650 information containers), TQ/RFI register | Median TQ closure hours; queries reopened per 100 raised | Per design package, or per 1,000 model elements | Stage 3 |
| Pre-construction & estimating | Take-off checking, subcontract comparison, risk-item extraction from tender documents | Estimating system and tender document set | Tender coverage achieved; variance between estimate and first CVR | Per £m of tender value, by work section | Stage 3 |
| Site production | Progress capture, short-term programme validation, plant and cycle analysis | Programme (P6 / Asta) and the field app | SPI at package level; days of float consumed per month | Per 1,000 programme activities | Stage 4 |
| Quality & rework | Defect and non-conformance detection, ITP evidence completeness | Field quality app, ITP and NCR registers | NCRs raised per 1,000 inspections; cost of rework | Per 1,000 inspections, and per £m certified | Stage 3 |
| Health & safety | Observation triage, permit and RAMS checking, high-risk activity flagging | H&S management system and observation register | Reportable incidents and near-miss closure rate | Per 100,000 worked hours (the RIDDOR base) | Stage 4 |
| Commercial | Variation and compensation-event assessment, subcontract payment checking | Cost ledger, CVR, application and payment records | Days to assess a variation; disputed value carried at period end | Per £m of certified value | Stage 3 |
| Handover & asset data | O&M and asset-data completeness checking against the AIR | CDE and the asset information model | Share of assets with complete data at handover; post-handover data queries | Per asset record, by classification | Stage 4 |
Design and commercial are the domains to start with, and the reason is measurement rather than value. Technical queries and variations are already logged in a register with timestamps, an owner and a contractual meaning — under the NEC (opens in a new tab) forms, early warnings and compensation events carry defined notification and response periods, which is a measurement gift, because the clock is already contractual rather than invented. The baseline therefore exists whether or not anyone has looked at it, and the exclusion arguments are bounded. Site production has the larger prize and the harder measurement problem: programme activities are re-baselined, renumbered and resequenced, so a KPI expressed per 1,000 activities needs the re-baseline register described below or it will report improvement that is purely an artefact of the schedule being rebuilt.
Health and safety deserves a specific warning. It is the one domain where construction already has a properly normalised, externally comparable KPI family — RIDDOR-reportable incident rates (opens in a new tab) expressed per 100,000 worked hours, published in the HSE's industry statistics (opens in a new tab) — which makes it tempting as the first place to demonstrate AI value. Resist that. Reportable incidents are rare enough that a single project cannot produce a statistically meaningful delta inside a reporting year, so a favourable movement is far more likely to be noise than signal. Use leading indicators — observation closure rates, share of high-risk activities with a checked permit — and report them per 100,000 worked hours, in line with HSE's construction guidance (opens in a new tab), while treating the lagging rate as context rather than as evidence.
One constraint sits underneath the whole specification and is easiest to breach in the safety and production domains: several of these metrics are computed from records about identifiable workers. Observation data names people, site imagery captures them, and a coverage metric computed per operative shades quickly into workforce monitoring. The ICO's guidance on AI and data protection (opens in a new tab) sets the expectations for lawful basis, necessity and transparency here, and the EU AI Act (opens in a new tab) treats AI used in worker management as high-risk while prohibiting emotion inference in the workplace altogether. The practical rule for KPI design is simple and costs nothing: aggregate to the package, the shift or the trade, never to the individual, and say so in the definition sheet. A metric that cannot be reported at individual level is also a metric that cannot be gamed by targeting individuals.
The eight tests a construction AI KPI must pass before it enters a board pack
Run these against one number you already report. Most firms tick two or three the first time. Tick as you go — this list works without JavaScript.
0 of 8 ticked
Nothing ticked — which is the honest starting point for most firms
Zero ticks does not mean the AI is not helping; it means nothing you report about it can be checked. The fix is one page, not a programme: pick a single decision on a single live package, write the definition sheet, and pull the before-figure from the register that already holds it. That is the whole distance between Anecdotal and Baselined.
1 of 8 — usually the forum, and rarely the definition
One tick is most often the last item: the number reaches the monthly review. That is worth something, because the audience exists. What is missing is anything that would survive the commercial manager's second question. Write the definition sheet next — numerator, base, exclusions, owner — and take a dated extract of the baseline before anything else changes.
2 of 8 — a number with an audience and no provenance
Two ticks typically means the metric is defined and reported, but computed by hand from the tool's own output. The next move is mechanical: point the calculation at the register the contract recognises — the TQ register, the NCR register, the cost ledger — and let the tool's log only tell you which records it touched.
3 of 8 — computed from the record, still uncomparable
Three ticks is the classic Baselined position on one project. The binding gap is the denominator: while the metric is absolute, it cannot be put next to any other job. Choose the base for this domain, write down why, and reproduce the number on a second, deliberately different project to find out which parts of your definition were project-specific assumptions.
4 of 8 — half the tests, and the missing half is attribution
At four ticks the metric is usually well defined and well sourced, with no baseline discipline or counterfactual behind it. That is the half that commercial people care about. Take a dated pre-go-live extract on the next project before mobilisation, and identify two comparable packages to leave on the previous process. Neither costs money; both need deciding before the tool is switched on.
5 of 8 — credible internally, not yet portable
Five ticks usually leaves recomputability, the gaming note and the cohort. Recomputability is the one to close first, because it forces retention: if the monthly extract is overwritten, nothing before today can ever be checked again. Fix retention this month, then write the gaming mode next to each metric — it takes an afternoon and prevents the most expensive kind of measurement failure.
6 of 8 — two gaps, and they are specific to your estate
At six ticks the remaining items are named pieces of work rather than general advice, and one of them is nearly always the counterfactual. A short session against your live portfolio usually finds a workable matched cohort you had not noticed — schemes on the same framework, same phase, same designer — which is the cheapest attribution route available to a construction business.
7 of 8 — one test left, and it is usually assurance
The last unticked box is most often version control and retention: the definition lives in a spreadsheet and the inputs behind last year's figure are gone. That is what separates a number your board believes from a number a client's assurance team can rely on. Move the definitions into the management system, set a retention rule, and run one recomputation test.
8 of 8 — you are measuring at Assured. Now protect it
All eight ticked means the KPI can be defended by someone who has never met the project. The risk from here is decay rather than failure: definitions age, the portfolio mix shifts, the base stops fitting and the recomputation stops being run. Put the set on an annual review with a change log, and test one figure a year against a live project's records.
The denominator problem, worked through
One AI programme, four ways of dividing the same benefit, four different conclusions — and the rule for choosing a base that survives the next project.
The denominator problem is that the same underlying improvement can be reported as a triumph, a rounding error or a regression depending entirely on what it is divided by, and construction offers more plausible divisors than almost any other industry. Contract value, gross internal floor area, linear metres, programme activities, worked hours, design packages, inspections and certified value are all defensible bases, and they disagree with each other — sometimes about direction, not just magnitude.
| Reported as | What the number says | What it actually means | How it fails |
|---|---|---|---|
| Hours saved (absolute) | 1,400 engineer-hours saved across the framework this year | A count of the tool's own estimated time savings, summed over every scheme | Not comparable with anything; grows automatically with framework size; nobody can check the per-query assumption |
| Percentage vs last year | Median TQ closure down 22% year on year | Two different scheme mixes compared as if they were one population | Last year's mix included two design-and-build schemes with novated designers; this year's did not. The mix moved, not the performance |
| Per £m of certified value | Queries raised per £m down from 14 to 11 | Query intensity normalised by commercial throughput | Reasonable for commercial workload, wrong for design workload — a £5m fit-out generates more queries per £m than a £150m earthworks package by construction, not by performance |
| Per design package, matched cohort | Median closure 61 → 47 working hours on six AI-supported schemes; 58 → 55 on six matched schemes | The delta after the base and the comparison group are both fixed | Only fails if the matching is weak — which is why the matching variables and the unmatched confounders are published with the figure |
The bottom row is the only one that answers the question the business asked, and it is also the least impressive-looking. That trade is the entire discipline. A programme that reports 1,400 hours saved will get applause for two quarters and then be asked to prove it; a programme that reports a fourteen-hour median improvement against a matched cohort will get an argument on day one and be believed thereafter. Construction commercial teams are professionally sceptical of unnormalised numbers because they spend their working lives normalising things — that is what the RICS measurement standards (opens in a new tab) exist to do for quantities, and the same instinct is applied to any benefit claim that arrives without a base.
Choose the base from what drives the numerator, not from what is easy to obtain
Contract value is available for every project, which is why it becomes the default base and why it is usually wrong. Ask what actually generates the events you are counting. Technical queries are generated by design complexity and interface count, so the base is design packages or model elements. Variations are generated by commercial throughput, so the base is certified value. Snags are generated by inspected work, so the base is inspections.
Fix the population before the base
A base is meaningless without a stated population: which packages, which disciplines, which contract forms, over what period. 'Median TQ closure hours' is not a metric until it says whose queries — main contractor to designer only, or including subcontractor queries routed through the site team, which behave completely differently.
Register every re-baseline, because the programme will move
Any metric expressed per programme activity breaks silently when the schedule is rebuilt, and on a live infrastructure job that happens more than once. Keep a re-baseline register with the date, the activity count before and after, and a note on the periods affected. Earned-value practice already treats re-baselining as a controlled event with a documented rationale — see the APM's earned value management resources (opens in a new tab) — and AI KPIs should inherit the same rule rather than inventing a weaker one.
Write down the base's rationale, not just the base
A base with no stated reason gets changed by whoever finds it inconvenient, usually in the quarter it stops flattering the programme. One sentence — 'per design package, because query volume tracks interface count rather than value' — is enough to make a change a decision rather than a drift.
Publish the confounders you could not remove
Weather, ground conditions, design maturity at award and subcontractor capability move construction metrics more than most tools do. Listing them alongside the number is not a hedge; it is the reason the number gets believed. A delta presented as clean invites the room to invent confounders, and the ones they invent will be larger than the ones you would have declared.
There is a public-sector precedent worth borrowing. The UK's Construction Playbook (opens in a new tab) requires should-cost models and benchmarking on public works — that is, an explicit expected value, built before the work, against which the outturn is compared. Whatever one thinks of its application, the discipline is exactly the one an AI adoption KPI needs: state the expected number before the intervention, in a defined base, and compare the outturn against it rather than against a memory. Firms that already produce should-cost models have the muscle for this and rarely apply it to their own digital investment.


