Redefining Technology

Construction & InfrastructureAI Adoption & Maturity Curve

AI adoption KPIs in construction and infrastructure: the denominator problem, and how to fix it

AI adoption KPIs are the metrics that prove an AI programme is changing how a construction business builds, not merely that software has been bought. In construction they fail for one structural reason: every project is a one-off, so a number with no agreed denominator, baseline or comparison cohort cannot be defended.

Site team gathered around a table reviewing performance charts, with a concrete frame under construction and an excavator behind them
Construction & Infrastructure · AI Adoption & Maturity Curve

Key takeaways

  1. A construction AI adoption KPI is only as good as its denominator. Absolute claims — hours saved, documents processed, queries closed — cannot be compared between a £3m fit-out and a £300m viaduct, so they die at the first commercial challenge. Normalise per £m of certified value, per 1,000 programme activities, per 100,000 worked hours or per inspection, and say which.
  2. Construction already owns one properly built KPI family: safety. RIDDOR-reportable rates are expressed per 100,000 worked hours, over a defined period, for a defined population. Copy that grammar — numerator, denominator, population, period, exclusions — for every AI adoption KPI you intend to report twice.
  3. The baseline must be taken before the tool is switched on, from records that already exist in the CDE, the cost ledger and the programme. A baseline reconstructed afterwards inherits the pilot's own effects, and on a project with a re-baselined programme it cannot be reconstructed at all.
  4. Attribution needs a counterfactual, and in construction that is a matched cohort rather than a holdout: comparable packages on the same framework, same phase, same designer, left on the previous process. Without one, weather, design maturity and subcontractor mix will take the credit or the blame.
  5. Five measurement stages — Anecdotal, Counted, Baselined, Normalised, Assured — decide which KPIs you may honestly report. Most firms sit at Counted, reporting licence utilisation and documents processed, which is activity data rather than adoption data and never survives a renewal conversation.

Abbreviations used on this page

CDE
Common data environment (the ISO 19650 information container)
EIR
Exchange information requirements — what the client asks for, contractually
BEP
BIM execution plan — how the supply chain says it will deliver the EIR
AIR
Asset information requirements — what the operator needs at handover
RFI
Request for information (US/global usage)
TQ
Technical query — the UK civils equivalent of an RFI
ITP
Inspection and test plan
NCR
Non-conformance report
CVR
Cost value reconciliation — the monthly commercial truth on a project
EVM
Earned value management, and its indices SPI (schedule) and CPI (cost)
NRM
RICS New Rules of Measurement — the industry's standard measurement basis
RIDDOR
Reporting of Injuries, Diseases and Dangerous Occurrences Regulations

Free · 8 questions · ~3 minutes

Score how well you can actually measure AI adoption

Eight questions, one at a time, about three minutes. They score the measurement system rather than the technology: whether your AI numbers have a denominator, a baseline, a counterfactual and an owner. Answer them and we build your personalised report — your stage on the measurement ladder, your score on each of the four dimensions, and the specific fix that unlocks the next stage — and send it to your inbox.

0 of 8 answered

Question 1 of 8Denominator discipline

When your AI programme reports a benefit, what is that benefit divided by?

An absolute number cannot be compared between a fit-out and a viaduct, so it dies at the first commercial challenge. The base is what makes a KPI portable between projects.

How the score maps to a stage
  • 04 — Stage 1, Anecdotal. AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
  • 59 — Stage 2, Counted. Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
  • 1014 — Stage 3, Baselined. One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
  • 1519 — Stage 4, Normalised. The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
  • 2024 — Stage 5, Assured. The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.

What AI adoption KPIs are in construction — and why they break

A definition, the two KPI families, and the path a number travels from a site event to a board pack — including the shortcut that makes most construction AI numbers unusable.

AI adoption KPIs in construction are the metrics that prove an AI capability has changed how a project is delivered, rather than proving that software has been bought and logged into. They come in two families that only mean something together: adoption KPIs, which measure whether the capability has entered the work — share of technical queries triaged, share of inspections captured through the tool, time from site event to a usable answer — and value KPIs, which measure what changed in the units the project already runs on: days on the critical path, non-conformances per 1,000 inspections, cost of rework, disputed variation value, RIDDOR-reportable incidents per 100,000 worked hours.

The reason these KPIs fail in construction and not, say, in a factory is structural. A production line offers a stable denominator — units, shifts, machine hours — that persists for years. A construction business is a portfolio of temporary organisations, each with its own package breakdown, its own designer, its own subcontract mix, its own weather, and a programme that will be re-baselined at least once. A number computed on one job therefore has no natural counterpart on the next, and the honest answer to 'is our AI adoption improving?' is unavailable until somebody chooses a base, a comparison group and a set of exclusions and writes them down.

Which KPIs you can honestly report depends on where you sit on the measurement ladder set out below. A firm at Counted can report usage and nothing else; a firm at Baselined can report a delta on one project with caveats; only a firm at Normalised can put a portfolio figure in a bid without inviting a challenge it cannot answer. Reporting above your stage is not ambition, it is the mechanism by which a programme loses credibility — and it is much more common than under-reporting.

Credible measurement against time on the ladder

What is plotted here is not value delivered but value that can be defended. It stays close to flat through Anecdotal and Counted — where most construction firms are — and inflects at Baselined, the first stage at which a number survives a commercial challenge. This is why firms with genuinely useful AI tools often cannot show a single defensible figure.

Benefit that survives a commercial challenge by stage

  • Stage 1 · Anecdotal — 34% of operators. AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.
  • Stage 2 · Counted — 38% of operators. Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.
  • Stage 3 · Baselined — 19% of operators. One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.
  • Stage 4 · Normalised — 7% of operators. The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.
  • Stage 5 · Assured — 2% of operators. The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.

Curve shape: logistic, plotted from the stage data above. Distribution: Framing consistent with the UK construction industry KPI programme.

How an AI adoption KPI is manufactured on a construction project

The same site event can produce two completely different numbers. The top lane is the path most reported construction AI figures actually take — through the vendor's dashboard and into a slide, arriving at the board pack with its definition stripped off. The middle and bottom lanes are what it costs to make the number checkable and then portable.

  • Data & feeds
  • AI / model
  • Where value leaks
  • Human in the loop
  • System-of-record action

The process, in words

  • In the site and package lane, an event — a technical query raised, a pour signed off, a snag logged — is picked up by the AI tool, which produces a flag, a score or a drafted response. The tool counts its own work on its own dashboard, using a numerator and a base the vendor chose, and somebody screenshots the result into the monthly pack. The number arrives at the board with its definition stripped off, which is why it cannot be recomputed and does not survive a challenge.
  • In the project-controls lane, the same site event is written to records that already exist — the CDE, the cost ledger, the programme. A metric definition sheet fixes the numerator, the base, the exclusions and the owner; a scheduled job recomputes the figure from retained records; and the result lands in the monthly review and the CVR, a forum with the authority to change scope, spend or sequence.
  • In the portfolio and assurance lane, the definition is expressed in a normalisation base that holds across projects and compared against a matched cohort of packages left on the previous process. What comes out is an attributed delta rather than a raw improvement, and the evidence pack behind it — definition version, retained inputs, recomputation route — is what makes the figure usable in a tender or a framework performance review.
  • The dashed red arrow is the shortcut nearly every construction business takes at least once: the vendor's dashboard number promoted straight into the assured position without a base, a baseline or a trail. It is fast, it looks identical in a slide, and it is the single reason most construction AI benefit claims cannot be repeated a year later.
Step-by-step insights
The site event — one event, two registers, two numbers
The most under-appreciated fact about construction measurement is that the same event is usually recorded twice: once in the tool and once in the contractual register. A technical query exists in the AI assistant's log and in the TQ register in the CDE, with different timestamps, different closure semantics and sometimes different identities, because the tool created a thread that the register treats as one query with two responses. Any KPI you intend to defend has to be computed from the contractual register, with the tool's log used only to establish which records it touched. The moment the tool's log becomes the numerator, the number has left the world the contract operates in.
The vendor dashboard — a marketing instrument doing a management job
Vendor dashboards are not dishonest; they are built to demonstrate that the product is working, which is a different question from whether the project is better off. Their base is almost always the tool's own activity — items processed, items flagged, users active — and their period is the licence period rather than the reporting period. They are perfectly good as an operational health check and completely unusable as a KPI, because the numbers cannot be reproduced from records you retain, and the definition can change with a product release. Treat them the same way you would treat a subcontractor's own productivity claim: useful signal, not evidence.
The definition sheet — where the arguments belong
The definition sheet is one page and it is where the difficult conversations happen. Does a query withdrawn on the day it was raised count. Is a query bounced back for missing information the same query or a new one. Does closure mean the designer's formal response or site acknowledgement of it. Do weekends count. These look like pedantry and they are the whole of the measurement: two honest people computing from the same register with different exclusion rules will produce numbers ten or fifteen per cent apart, and every subsequent argument about the AI tool will actually be an argument about the exclusion rules. Settle them once, in writing, with the commercial team in the room.
Scheduled recomputation — the difference between a metric and a memory
The reporting job should be rerunnable by someone who has never met the project, over retained extracts, producing the same answer. That has two consequences people underestimate. First, it forces retention: if the CDE export is overwritten each month, last quarter's figure can never be checked. Second, it forces the definition into code, where a change is visible in a diff rather than in someone's habit. Construction firms that have done this once usually find the first recomputation disagrees with the originally reported number, and that discovery — early, internal, harmless — is exactly the point of building it.
The normalisation base — chosen per domain, never universally
The base has to reflect what actually drives the numerator. Information-intensive work — design queries, technical submittals, drawing revisions — scales with design packages and model elements, not with contract value. Commercial workload scales with certified value and the number of variations. Production scales with programme activities. Safety already has its base: incidents per 100,000 worked hours, which is the grammar used in RIDDOR-based industry reporting and the reason safety is the one construction KPI family that survives comparison between firms. Copy that grammar per domain and write down why each base was chosen; a base with no stated rationale gets changed by the next person who finds it inconvenient.
The matched cohort — construction's substitute for a holdout
You cannot randomise a viaduct, so attribution in construction is done by matching rather than by holding out. Match on the variables that drive the metric — scheme type, contract form, delivery phase, designer, whether the design was novated, approximate value band — and keep a written register of what you could not match on, typically weather, ground conditions, design maturity at contract award and subcontractor mix. Read the register out with the number. A delta presented with its confounders is believed; the same delta presented as a clean result invites the room to invent the confounders themselves, and they will find bigger ones than you did.

The five stages of the measurement ladder

Anecdotal, Counted, Baselined, Normalised, Assured — what each looks like on a live project, the signals a reviewer can check in an afternoon, and the anti-pattern that traps firms there.

The ladder below measures the measurement system, not the technology — a firm running sophisticated computer vision on every site can sit at Counted, and a firm with one language model and a disciplined quantity surveyor can sit at Baselined. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own registers this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Anecdotal

34% of operators sit here

AI value exists only as project stories — a number somebody remembers, with nothing behind it that another person could check.

Anecdotal is the default state, and it is not a sign that nothing is happening. On most sites something genuinely is: an engineer is using a language model to draft technical query responses, a planner is testing an assistant against the four-week look-ahead, a quantity surveyor is pulling clauses out of a subcontract faster than they used to. The work is real. What does not exist is any artefact a second person could use to verify it.

The failure mode is specific and it is commercial rather than technical. A remembered number has no numerator anybody agreed, no denominator at all, and no before-figure. When the commercial manager asks the obvious question — over how many queries, compared with what, on which package — there is no answer, and the claim is quietly withdrawn. That withdrawal is what teaches a project board that AI benefits are soft, and it is much harder to undo than to avoid.

This stage is cheap to leave. It does not require a platform, a data strategy or a business case; it requires one page of writing. Pick a single decision on a single live package, write down what you are counting, what you are dividing it by, what you are excluding and who owns the number, and take the before-figure from records that already exist. That page is the whole difference between Anecdotal and the stage above it.

In practice

The three-hour claim at the monthly review

On a highways widening scheme, a site engineer tells the monthly review that the AI document assistant saves him about three hours a week finding clauses and specification references. The project director likes it and repeats it to the client. The commercial manager asks two questions: three hours against what, and across how many technical queries? Nobody has the TQ register open, and nobody knows what last quarter's median closure time was. By the following month the claim has been dropped from the pack — not because it was untrue, but because it could not be checked.

What it looks like

  • Benefit is quoted in remembered time — 'about half a day a week'
  • No written definition of what is being counted, or of what it is divided by
  • The only record is a slide in the digital lead's deck
  • The claim does not survive project closeout or a change of project director

Diagnostic signals you can check this week

  • Ask for the definition of any AI benefit number quoted in the last board pack. If it exists only in a slide, you are here
  • Ask what the figure is divided by. Silence, or 'per project', means there is no denominator
  • Ask what the before-figure was and where it came from. 'We think it was about…' is the tell
  • Check whether any AI claim survived a project handover into the next job's bid or lessons-learned pack

Anti-pattern · Commissioning a benefits study

The instinct at this stage is to buy a business-wide benefits assessment — a consultant, a survey of project teams, a modelled figure for the group. It produces a number with a decimal point and no provenance, and the first framework client or auditor who asks how it was derived collapses it. Worse, it teaches the organisation that AI value is estimated rather than measured. Write one metric definition sheet for one decision instead; it costs a morning and it is the only artefact that compounds.

What holds you here

Nothing is written down, so no claim can be checked by a second person — and an unchecked claim is worth nothing commercially.

Highest-leverage next move

Write one metric definition sheet — numerator, denominator, exclusions, owner — for one decision on one live package, and take the before-figure from the CDE.

Cost of leaving

Effort
2–3 months
Team
Digital lead plus half a day a week from a quantity surveyor or planner
Risk
Low — the work is definitional and nothing operational depends on it yet
To next stage
2–3 months

If this is you, the next step is

A half-day workshop with your commercial and digital leads on one live package.

Write your first metric definition sheet

Stage 2

Counted

38% of operators sit here

Activity is counted and reported as adoption — licences, active users, documents processed, models run — with nothing connecting the count to a project outcome.

Counted is where most construction businesses actually sit, and it is a genuine improvement on Anecdotal: the numbers exist, they are produced on a schedule, and they come out of a system rather than a memory. Licence utilisation, active users per project, documents processed, images analysed, drafts generated — all of it is real telemetry. It is simply telemetry about the tool rather than about the business.

The structural problem is that a count has no denominator and no outcome, which makes it unusable in exactly the two conversations that matter. It cannot be compared between projects, because a 400-unit residential job will always process more documents than a substation upgrade; and it cannot be defended at renewal, because nothing in the count tells the commercial director whether a date moved or a cost changed. Counts also invite the worst possible intervention: an adoption campaign that raises usage without raising value, after which the count is not merely uninformative but actively misleading.

Leaving this stage is not a matter of counting better. It is a matter of choosing one decision the tool is supposed to change, naming the outcome that decision moves in the units the project already runs on — days, pounds, non-conformances, worked hours — and taking the before-figure before the next project starts. One decision measured properly outranks forty projects counted.

In practice

The 78% licence utilisation slide

A tier-one contractor reports 78% seat utilisation for its AI document platform across forty live projects, up eleven points on the quarter, with a chart of documents processed per month. At the renewal meeting the commercial director asks what the eleven points bought. The digital team can show that usage rose fastest on the three projects with the most rework, which is either evidence that the tool goes where the pain is, or evidence that the tool creates work — and nothing in the reporting distinguishes the two. The renewal goes through on relationship rather than on evidence, which is the point at which the programme becomes politically fragile.

What it looks like

  • Monthly digital report leads with seat utilisation and active-user counts
  • Volume metrics — drawings scanned, queries drafted, images processed
  • Counts are absolute, so a big project always looks like a success
  • No commercial or delivery metric appears next to any of it

Diagnostic signals you can check this week

  • Open the last AI report. If every metric is a count of activity rather than a rate against a base, you are here
  • Ask whether any AI metric appears next to CPI, SPI, NCR counts or accident frequency rate in the same pack
  • Check whether the number would go up if the project simply got bigger. If it would, it is a count, not a KPI
  • Ask what decision changed as a result of last quarter's report. If the answer is a renewal, the reporting is serving the vendor

Anti-pattern · Running an adoption campaign to make the count go up

When usage plateaus, the standard response is a push — training, mandates, league tables between projects. Usage duly rises, and the relationship between usage and outcome, which was never established, becomes unrecoverable: you can no longer tell whether high-usage projects use the tool because they are well run or because they are struggling. Any incentive attached to a count will be met, and the data it produces afterwards is worse than the data you had before. Define the outcome metric first, then let usage be whatever it needs to be.

What holds you here

The count has no denominator and no outcome, so it cannot be compared between projects or defended when the licence comes up for renewal.

Highest-leverage next move

Choose one decision, define the outcome KPI in the project's own units, and capture the pre-AI baseline before the next project mobilises.

Cost of leaving

Effort
3–6 months
Team
Digital lead, a commercial manager, a planner, and one project willing to be first
Risk
Medium — renewal and rollout decisions are being taken on the wrong number in the meantime
To next stage
3–6 months

If this is you, the next step is

We pick the decision, define the metric and find the before-figure in your existing records.

Turn your counts into one real KPI

Stage 3

Baselined

19% of operators sit here

One decision has a written definition, a pre-AI baseline pulled from the system of record, and a number recomputed from records rather than typed in — on one project.

Baselined is the first stage at which a number can survive a challenge. The definition is written, the before-figure came from a system rather than a recollection, and the current figure is produced by rerunning a query over records that are retained. If the project director is asked in six months where the figure came from, the answer is a script and a record set, not a person's diary.

The character of the work here is unglamorous and mostly commercial. The arguments are about exclusions — does a technical query raised and withdrawn on the same day count; is a query bounced back for missing information a new query or the same one; does a Saturday count as a working day for closure time. Those arguments feel like pedantry and they are the entire value of the stage: an unstated exclusion rule is the most common reason two people compute different numbers from the same register and both are honest.

The constraint that ends this stage is comparability. The number is true for this project, and it says nothing about the next one, because activity identifiers, cost codes, query classifications and package breakdowns differ between jobs — often between phases of the same job. A firm that stops here has one well-measured project and no way to aggregate, which means the KPI cannot inform an investment decision or a bid.

In practice

The technical-query baseline on a viaduct package

Before switching on AI-assisted TQ triage on the viaduct substructure package, the project-controls lead exports eighteen months of the TQ register from the CDE, computes the median and 90th-percentile closure time in working hours, and writes down the rules: withdrawn queries excluded, queries reopened within five days treated as the same query, closure measured to the designer's formal response rather than the site acknowledgement. The current figure is recomputed weekly by a scheduled job. When the number improves, the argument at the review is about whether the design was simply more mature this quarter — a real question, and one the project cannot yet answer.

What it looks like

  • A metric definition sheet exists: numerator, denominator, population, period, exclusions, owner
  • The baseline was taken from the CDE, the cost ledger or the programme before go-live
  • The KPI is recomputed on a schedule, from raw records, by a job rather than a person
  • It appears in the monthly project review alongside the commercial numbers

Diagnostic signals you can check this week

  • Ask to see the metric definition sheet. If the exclusions section is blank, the definition is not finished
  • Ask who could rerun the calculation if the digital lead were on leave, and whether the inputs are retained
  • Check whether the baseline predates the tool's go-live date, from evidence rather than assertion
  • Ask whether the same number could be produced for a second project without new hand-mapping

Anti-pattern · Mandating it across the portfolio before it has been keyed twice

One clean measurement produces immediate pressure to make it a corporate KPI. It gets mandated across thirty projects that classify technical queries differently, use different work-breakdown levels and run three different CDEs, and within a quarter the group figure is an average of incompatible things. Prove the definition on a second, deliberately different project first — different sector, different client, different designer — and fix what breaks. The second project is where you learn which parts of the definition were actually project-specific assumptions.

What holds you here

The number is true for one project and comparable with nothing, because identifiers, classifications and package breakdowns differ across jobs.

Highest-leverage next move

Agree the normalisation base and the classification the KPI is expressed in, then reproduce it on a second, structurally different project.

Cost of leaving

Effort
4–8 months
Team
Project-controls lead, a data engineer part-time, commercial sign-off on the definition
Risk
Medium — the definitional arguments are with the commercial team and cannot be skipped
To next stage
4–8 months

If this is you, the next step is

We take your existing definition to a deliberately different job and report what breaks.

Test your definition on a second project

Stage 4

Normalised

7% of operators sit here

The KPI is expressed in a base that holds across projects and compared against a matched cohort, so a portfolio number means something.

Normalised is the stage at which a construction business can finally answer the question it has been asked since stage one: is this working, across our portfolio, and by how much. That becomes possible only when two things exist together — a base that survives differences in project size and type, and something to compare against that is not simply the last job.

The base is domain-specific and choosing it is real work. Design-side metrics normalise well per design package or per thousand model elements; commercial metrics per £m of certified value; production metrics per thousand programme activities; safety metrics per 100,000 worked hours, which is the base the industry already uses under RIDDOR reporting. One base applied to everything is the classic error: divide a fit-out and a viaduct by contract value and the viaduct will look efficient at everything, because value per unit of information is completely different in the two.

The counterfactual in construction is a matched cohort rather than a holdout, because you cannot randomise a bridge. Matching is on the variables that actually drive the metric — scheme type, contract form, phase, designer, client, and whether the design was novated — and the confounders you cannot match on go in a register that is read out with the number. Weather, design maturity and subcontractor mix will otherwise take the credit or the blame, and everyone in the room will know it.

In practice

The cohort that made the framework case

A civils contractor on a five-year water framework runs AI-assisted technical-query triage on six schemes and leaves six comparable schemes on the previous process — matched on AMP delivery phase, designer, contract form and approximate value band. The KPI is reported as median TQ closure hours and TQs reopened per 100 raised, both per scheme, with a confounder note on two schemes affected by a wet winter. The reported delta is smaller than the pilot claimed and it is believed, which is what allows the client to write the measure into the next framework's information requirements.

What it looks like

  • Every KPI carries a normalisation base — per £m certified, per 1,000 activities, per 100,000 worked hours
  • A matched cohort exists: comparable packages or schemes left on the previous process
  • A re-baseline register records every programme and scope change that breaks a series
  • The KPI sits in the portfolio pack next to CPI, SPI and accident frequency rate

Diagnostic signals you can check this week

  • Ask what base each KPI is divided by, and whether the same base is used across sectors. One base everywhere is a red flag
  • Ask to see the cohort definition and the matching variables, in writing
  • Ask what happens to the KPI series when a programme is re-baselined. If nothing happens, the series is silently broken
  • Check whether the confounder register is read out with the number or filed separately

Anti-pattern · Normalising everything by contract value

Contract value is available for every project, which is exactly why it gets adopted as the universal base — and it is wrong for most metrics. Information intensity does not scale with value: a £5m hospital fit-out can generate more technical queries than a £150m earthworks package. Using value alone makes complex low-value work look like the worst performer in the portfolio and rewards teams for being on big, information-light jobs. Choose the base per domain, write down why, and review it when the portfolio mix changes.

What holds you here

The numbers are credible internally but carry no assurance trail, so they cannot be used in a bid, an insurance discussion or a client audit.

Highest-leverage next move

Make the KPI recomputable by someone outside the team: version the definition, retain the inputs, and build the evidence pack.

Cost of leaving

Effort
9–15 months
Team
Project controls, commercial analytics, a data engineer, and a portfolio owner with the authority to standardise
Risk
Higher — standardising classifications across live projects is change management, not engineering
To next stage
9–15 months

If this is you, the next step is

Two workshops: one on bases per domain, one on cohort matching against your live portfolio.

Design your normalisation base and cohort

Stage 5

Assured

2% of operators sit here

The KPI set is a contractual, auditable artefact — defined in the information requirements, computed from retained records, and recomputable by a third party.

Assured is narrow, and it should be. It does not mean every AI metric in the business is audited; it means a small, named set of KPIs has been promoted to the same status as certified value or reportable-incident rates. Those are the numbers that appear in a prequalification questionnaire, in a framework performance review, or in an insurer's questions about how design-review decisions are being made — and each one carries a definition, a version, retained inputs and a recomputation route.

The engineering at this stage is largely done. The work is management-system work: an owner, a change log, a review cycle, a retention rule, an evidence pack that a person who has never met the project can follow. Firms already running an ISO 19650 information management function and, increasingly, an AI management system in the shape of ISO/IEC 42001 have the machinery for this; the KPI set simply becomes another controlled artefact inside it.

The characteristic failure of Assured is not collapse but slow decay. Definitions age while the operation moves, the portfolio mix shifts and the base stops fitting, personnel change and the recomputation script stops being run. Assurance is a cycle rather than a state: an annual definition review, a change log that says why each version changed, and one live recomputation test a year against a real project's records.

In practice

The recomputation on a framework audit

A client's assurance team picks one reported figure — non-conformances raised per 1,000 inspections on the AI-supported quality process — and asks the contractor to reproduce it. The contractor supplies the definition at the version in force that quarter, the extract of the inspection and NCR registers with record identifiers, and the script. The recomputed figure differs by under two per cent, explained by three records reclassified after the report date, and the difference itself is documented. That exchange takes four days and is the reason the measure is accepted in the next tender.

What it looks like

  • Definitions are versioned and referenced in the EIR and the BEP, not held in a spreadsheet
  • Inputs are retained with lineage for the full contractual retention period
  • Internal audit recomputes a sample and the result is reproducible within tolerance
  • The figures appear in tender submissions and withstand client and insurer questions

Diagnostic signals you can check this week

  • Ask whether the KPI definitions are version-controlled with a change log, and who approves a change
  • Ask whether the inputs behind last year's reported figure are still retrievable today
  • Ask when the recomputation test was last run, and by whom outside the reporting team
  • Check whether the KPI set is referenced anywhere contractual — the EIR, the framework performance schedule, a prequalification response

Anti-pattern · Freezing the definition to protect the trend

Once a KPI has two years of history, changing it breaks the series, and the reflex is to leave it alone. The operation moves anyway: the portfolio shifts toward frameworks, the classification changes, a new CDE arrives. The metric slowly stops describing the business while continuing to look consistent, which is worse than an obvious break. Version the definition, restate the affected periods where you can, and record the discontinuity — a documented break is defensible, a silent drift is not.

What holds you here

Assurance decays quietly — definitions age, the portfolio changes shape, and last year's evidence pack no longer satisfies this year's audit.

Highest-leverage next move

Put the KPI set on the management system's review cycle: annual definition review, a change log, and one live recomputation test a year.

Cost of leaving

Effort
Continuous
Team
Information management, internal audit, commercial, with a standing review forum
Risk
Concentrated — low frequency, high consequence, contractual and reputational in nature

If this is you, the next step is

We attempt to recompute a KPI you have already published, and report exactly where it breaks.

Stress-test one reported figure

Two things are worth noticing about the shape of this ladder. The first is that the biggest single gap is between Counted and Baselined, and nothing in it is technical: it is a page of definitions, an extract taken before go-live, and a commercial manager willing to argue about exclusions. The second is that the top two stages are governed by documents rather than by systems. Normalised depends on a base and a cohort definition; Assured depends on version control, retention and a recomputation route. A firm can climb this ladder without buying anything.

Where construction and infrastructure firms actually sit

The distribution across the ladder, why Counted is the mode, and the external evidence that the industry's measurement problem predates AI entirely.

Most construction and infrastructure firms sit at Counted. They can tell you how many licences are active, how many documents the model has processed and how usage compares between regions, and they cannot tell you what any of it did to a date, a cost or a non-conformance. A minority have one project with a real baseline, and a very small number can put a normalised, cohort-compared figure in front of a client.

Distribution of construction firms across the measurement ladder

Counted is the mode and the plateau. The drop from Counted to Baselined is the largest transition loss on the ladder, and it is a documentation gap rather than a technology gap.

Share of firms

  • 34% — 1 · Anecdotal
  • 38% — 2 · Counted (the plateau)
  • 19% — 3 · Baselined
  • 7% — 4 · Normalised
  • 2% — 5 · Assured

Source: Illustrative distribution, synthesised from the UK construction KPI programme, IPA project-delivery reporting and NAO findings on infrastructure performance

The distribution above is a model-derived illustration rather than a survey, and it is anchored to three external bodies of evidence about how construction measures itself. The UK industry has published normalised headline construction KPIs through Constructing Excellence (opens in a new tab) since the late 1990s, which is why safety, client satisfaction and predictability have a shared grammar and almost nothing else does. The Infrastructure and Projects Authority (opens in a new tab) reports annually on the delivery confidence of the government's major projects, and the National Audit Office (opens in a new tab) has repeatedly found that infrastructure performance claims are hard to compare because baselines move. Every one of those findings applies to AI benefit claims without modification.

5%

Of project cost is directly measurable avoidable error, before indirect and latent costs, in GIRI's published research

Get It Right Initiative

per 100k

Worked hours — the base the industry already uses for RIDDOR-reportable incident rates, and the model for every AI adoption KPI

HSE

1999

The UK construction industry has published normalised headline KPIs annually since the late 1990s

Constructing Excellence

It is worth being precise about what this distribution does and does not say. It is not a claim that two thirds of the industry gets no value from AI; plenty of firms in the Anecdotal and Counted bands have tools that genuinely help engineers and quantity surveyors every day. It is a claim about evidence: that the value is invisible to the commercial function, absent from the cost and output statistics (opens in a new tab) the sector reports, and therefore unavailable in the two conversations where it would matter most — the investment case for scaling, and the tender in which a client asks what your digital capability is actually worth to them.

The KPI specification: which metrics, in which units, from which record

The honest KPI set for each stage, the decision map that says where each metric comes from, and the eight tests a number must pass before it enters a board pack.

The right AI adoption KPI is the one your stage can actually support, computed from a register that already exists and expressed in a base that survives the next project. That gives three requirements, in order: pick metrics your evidence can carry, source each from a named system of record rather than from the tool, and divide each by something that does not change when the job gets bigger. The tables below are the working specification — the first says what each stage may honestly claim, the second says where every metric comes from in a construction estate.

StageAdoption KPIs — has it entered the work?Value KPIs — did anything change?The claim this stage cannot support
1 · AnecdotalNone that are checkable. At best, a written list of which decisions the tool touchesNone. The only honest output is a pre-AI baseline extracted from the CDE, the cost ledger and the programmeAny benefit figure at all — including a conservative one
2 · CountedLicence and active-user counts, items processed, coverage of a register (share of TQs the tool touched)None attributable. Value talk at this stage is projection, and should be labelled as suchThat usage growth means adoption, or that high-usage projects perform better
3 · BaselinedCoverage against a defined population, median time from event to usable answer, share of outputs accepted without rework by the responsible engineerOne named delta against a pre-AI baseline on one project — TQ closure hours, NCRs per 1,000 inspections, days to assess a variationThat the project delta generalises to the portfolio
4 · NormalisedCoverage and acceptance normalised per domain base, reported per scheme and per cohortAttributed deltas against a matched cohort — rework cost per £m certified, SPI at package level, disputed variation value per £mThat the figure is audit-ready, or usable in a contractual performance schedule
5 · AssuredThe same measures under version control, with retention and a recomputation test on recordAttributed deltas with confounders registered, restated across definition changes, referenced in tender and framework reportingThat assurance holds without an annual definition review — it decays quietly
What each stage of the measurement ladder can honestly report. Adoption KPIs prove the capability has entered the work; value KPIs prove something changed in the project's own units. The right-hand column is the claim that stage most often makes and cannot support.

Two disciplines make the whole specification trustworthy, and both are unpopular. The first is that adoption and value KPIs are always reported as a pair: a coverage figure with no delta is theatre, and a delta with no coverage figure is unexplainable, because nobody can tell whether it came from the tool or from the two packages where the tool was never used. The second is that every value KPI names its comparison in the same sentence as the number. 'TQ closure time down 22%' is a claim; 'median TQ closure down from 61 to 47 working hours on the six AI-supported schemes, against 58 to 55 on the six matched schemes' is a measurement.

Operating domainDecision AI touchesSystem of recordKPI it movesNormalisation baseHonest from
Design & engineeringTechnical query triage, drafting responses, clash and change reviewCDE (ISO 19650 information containers), TQ/RFI registerMedian TQ closure hours; queries reopened per 100 raisedPer design package, or per 1,000 model elementsStage 3
Pre-construction & estimatingTake-off checking, subcontract comparison, risk-item extraction from tender documentsEstimating system and tender document setTender coverage achieved; variance between estimate and first CVRPer £m of tender value, by work sectionStage 3
Site productionProgress capture, short-term programme validation, plant and cycle analysisProgramme (P6 / Asta) and the field appSPI at package level; days of float consumed per monthPer 1,000 programme activitiesStage 4
Quality & reworkDefect and non-conformance detection, ITP evidence completenessField quality app, ITP and NCR registersNCRs raised per 1,000 inspections; cost of reworkPer 1,000 inspections, and per £m certifiedStage 3
Health & safetyObservation triage, permit and RAMS checking, high-risk activity flaggingH&S management system and observation registerReportable incidents and near-miss closure ratePer 100,000 worked hours (the RIDDOR base)Stage 4
CommercialVariation and compensation-event assessment, subcontract payment checkingCost ledger, CVR, application and payment recordsDays to assess a variation; disputed value carried at period endPer £m of certified valueStage 3
Handover & asset dataO&M and asset-data completeness checking against the AIRCDE and the asset information modelShare of assets with complete data at handover; post-handover data queriesPer asset record, by classificationStage 4
The construction AI KPI decision map. Every metric is sourced from a register the business already keeps under its information standard, and every one carries the base it is divided by. 'Honest from' is the ladder stage at which the metric first measures something real.

Design and commercial are the domains to start with, and the reason is measurement rather than value. Technical queries and variations are already logged in a register with timestamps, an owner and a contractual meaning — under the NEC (opens in a new tab) forms, early warnings and compensation events carry defined notification and response periods, which is a measurement gift, because the clock is already contractual rather than invented. The baseline therefore exists whether or not anyone has looked at it, and the exclusion arguments are bounded. Site production has the larger prize and the harder measurement problem: programme activities are re-baselined, renumbered and resequenced, so a KPI expressed per 1,000 activities needs the re-baseline register described below or it will report improvement that is purely an artefact of the schedule being rebuilt.

Health and safety deserves a specific warning. It is the one domain where construction already has a properly normalised, externally comparable KPI family — RIDDOR-reportable incident rates (opens in a new tab) expressed per 100,000 worked hours, published in the HSE's industry statistics (opens in a new tab) — which makes it tempting as the first place to demonstrate AI value. Resist that. Reportable incidents are rare enough that a single project cannot produce a statistically meaningful delta inside a reporting year, so a favourable movement is far more likely to be noise than signal. Use leading indicators — observation closure rates, share of high-risk activities with a checked permit — and report them per 100,000 worked hours, in line with HSE's construction guidance (opens in a new tab), while treating the lagging rate as context rather than as evidence.

One constraint sits underneath the whole specification and is easiest to breach in the safety and production domains: several of these metrics are computed from records about identifiable workers. Observation data names people, site imagery captures them, and a coverage metric computed per operative shades quickly into workforce monitoring. The ICO's guidance on AI and data protection (opens in a new tab) sets the expectations for lawful basis, necessity and transparency here, and the EU AI Act (opens in a new tab) treats AI used in worker management as high-risk while prohibiting emotion inference in the workplace altogether. The practical rule for KPI design is simple and costs nothing: aggregate to the package, the shift or the trade, never to the individual, and say so in the definition sheet. A metric that cannot be reported at individual level is also a metric that cannot be gamed by targeting individuals.

The eight tests a construction AI KPI must pass before it enters a board pack

Run these against one number you already report. Most firms tick two or three the first time. Tick as you go — this list works without JavaScript.

0 of 8 ticked

Nothing ticked — which is the honest starting point for most firms

Zero ticks does not mean the AI is not helping; it means nothing you report about it can be checked. The fix is one page, not a programme: pick a single decision on a single live package, write the definition sheet, and pull the before-figure from the register that already holds it. That is the whole distance between Anecdotal and Baselined.

The denominator problem, worked through

One AI programme, four ways of dividing the same benefit, four different conclusions — and the rule for choosing a base that survives the next project.

The denominator problem is that the same underlying improvement can be reported as a triumph, a rounding error or a regression depending entirely on what it is divided by, and construction offers more plausible divisors than almost any other industry. Contract value, gross internal floor area, linear metres, programme activities, worked hours, design packages, inspections and certified value are all defensible bases, and they disagree with each other — sometimes about direction, not just magnitude.

Reported asWhat the number saysWhat it actually meansHow it fails
Hours saved (absolute)1,400 engineer-hours saved across the framework this yearA count of the tool's own estimated time savings, summed over every schemeNot comparable with anything; grows automatically with framework size; nobody can check the per-query assumption
Percentage vs last yearMedian TQ closure down 22% year on yearTwo different scheme mixes compared as if they were one populationLast year's mix included two design-and-build schemes with novated designers; this year's did not. The mix moved, not the performance
Per £m of certified valueQueries raised per £m down from 14 to 11Query intensity normalised by commercial throughputReasonable for commercial workload, wrong for design workload — a £5m fit-out generates more queries per £m than a £150m earthworks package by construction, not by performance
Per design package, matched cohortMedian closure 61 → 47 working hours on six AI-supported schemes; 58 → 55 on six matched schemesThe delta after the base and the comparison group are both fixedOnly fails if the matching is weak — which is why the matching variables and the unmatched confounders are published with the figure
One AI-assisted technical-query programme on a civils framework, reported four ways. The underlying work is identical in every row; only the base changes. Figures are a worked illustration of how bases behave, not measured results.

The bottom row is the only one that answers the question the business asked, and it is also the least impressive-looking. That trade is the entire discipline. A programme that reports 1,400 hours saved will get applause for two quarters and then be asked to prove it; a programme that reports a fourteen-hour median improvement against a matched cohort will get an argument on day one and be believed thereafter. Construction commercial teams are professionally sceptical of unnormalised numbers because they spend their working lives normalising things — that is what the RICS measurement standards (opens in a new tab) exist to do for quantities, and the same instinct is applied to any benefit claim that arrives without a base.

  • Choose the base from what drives the numerator, not from what is easy to obtain

    Contract value is available for every project, which is why it becomes the default base and why it is usually wrong. Ask what actually generates the events you are counting. Technical queries are generated by design complexity and interface count, so the base is design packages or model elements. Variations are generated by commercial throughput, so the base is certified value. Snags are generated by inspected work, so the base is inspections.

  • Fix the population before the base

    A base is meaningless without a stated population: which packages, which disciplines, which contract forms, over what period. 'Median TQ closure hours' is not a metric until it says whose queries — main contractor to designer only, or including subcontractor queries routed through the site team, which behave completely differently.

  • Register every re-baseline, because the programme will move

    Any metric expressed per programme activity breaks silently when the schedule is rebuilt, and on a live infrastructure job that happens more than once. Keep a re-baseline register with the date, the activity count before and after, and a note on the periods affected. Earned-value practice already treats re-baselining as a controlled event with a documented rationale — see the APM's earned value management resources (opens in a new tab) — and AI KPIs should inherit the same rule rather than inventing a weaker one.

  • Write down the base's rationale, not just the base

    A base with no stated reason gets changed by whoever finds it inconvenient, usually in the quarter it stops flattering the programme. One sentence — 'per design package, because query volume tracks interface count rather than value' — is enough to make a change a decision rather than a drift.

  • Publish the confounders you could not remove

    Weather, ground conditions, design maturity at award and subcontractor capability move construction metrics more than most tools do. Listing them alongside the number is not a hedge; it is the reason the number gets believed. A delta presented as clean invites the room to invent confounders, and the ones they invent will be larger than the ones you would have declared.

There is a public-sector precedent worth borrowing. The UK's Construction Playbook (opens in a new tab) requires should-cost models and benchmarking on public works — that is, an explicit expected value, built before the work, against which the outturn is compared. Whatever one thinks of its application, the discipline is exactly the one an AI adoption KPI needs: state the expected number before the intervention, in a defined base, and compare the outturn against it rather than against a memory. Firms that already produce should-cost models have the muscle for this and rarely apply it to their own digital investment.

The four dimensions that set your stage

Measurement maturity is not one number. Four dimensions gate each other, and the lowest one caps the credibility of everything you report.

Four dimensions decide which stage of the ladder a firm is genuinely at — denominator discipline, baseline and counterfactual, instrumentation and recomputability, and assurance and decision rights — and the lowest of the four is the real answer. A perfectly instrumented metric with no counterfactual is precise about nothing; a beautifully matched cohort reported from a spreadsheet nobody can rerun is true and unrepeatable. Scoring them separately is the point, because the total hides the constraint.

  • Denominator discipline

    Whether every reported number carries a base, whether the base was chosen from what drives the numerator, and whether it holds when the job grows. This dimension is where construction differs most from process industries, and it is the one that most often scores zero in firms with otherwise sophisticated data functions. The test is simple: if the metric would move purely because the project got bigger, there is no base.

  • Baseline and counterfactual

    Whether a pre-AI figure was extracted before go-live, and whether anything plays the role of a control. This is overwhelmingly the dimension that determines whether a claim survives a commercial challenge, and it is the one that cannot be retrofitted: a baseline not taken before the tool was switched on is gone, and on a re-baselined programme it is gone permanently.

  • Instrumentation and recomputability

    Whether the number is computed from a system of record by something rerunnable, and whether the inputs are retained. The practical test is a person: if the analyst who built it left tomorrow, could the figure be reproduced? Firms running an ISO 19650 information function already have the retention and naming machinery for this — see the UK BIM Framework (opens in a new tab) guidance on information management — and mostly have not pointed it at their own performance reporting.

  • Assurance and decision rights

    Whether definitions are versioned and controlled, and whether the number reaches a forum that can change scope, spend or sequence. A metric with no decision attached is reporting overhead no matter how well built, and a metric with decision rights but no version control is a liability — the moment it is challenged, nobody can say what it meant last year.

Diagnosing what is actually wrong with your numbers

Plot denominator discipline against baseline and counterfactual. The quadrant names the next investment — and in three of the four cases it is a document rather than a system.

Precise about nothing

  • Well-normalised metrics with nothing to compare against
  • Common in firms with a strong data function and a weak commercial link
  • Fix: take a dated pre-go-live extract on the next project, and pick the cohort before mobilisation

Defensible

  • Base and comparison both fixed and written down
  • The only quadrant whose numbers belong in a tender
  • Fix: version the definitions and prove recomputability before you rely on them

Anecdote

  • Absolute counts against a remembered before-figure
  • Where most construction AI reporting sits today
  • Fix: one definition sheet for one decision on one live package

True but unrepeatable

  • A well-run project experiment nobody can extend
  • Usually a single enthusiastic project team
  • Fix: choose the base per domain and reproduce the metric on a structurally different job
Denominator discipline — top: Normalised per domain base, bottom: Absolute counts
Baseline & counterfactual — left: No before-figure, no comparison group, right: Dated baseline and a matched cohort

One diagnostic is worth running before any of this. Take the last three AI benefit figures your business reported and, for each, write the sentence a sceptical commercial director would say next. If all three sentences begin 'compared with what', the constraint is the second dimension. If they begin 'across how many', it is the first. If they begin 'who worked that out', it is the third. The distribution of those sentences is a faster stage assessment than any maturity model, including this one.

How construction AI KPIs lie, and what stops each one

Six failure modes that produce confident, wrong numbers — each with the cheap preventive measure that removes it.

Construction AI KPIs lie in six recognisable ways, and none of them requires anyone to be dishonest. They are structural properties of measuring a portfolio of one-off projects: the base drifts while the work grows, the schedule is rebuilt underneath the series, the projects that volunteer for pilots are the well-run ones, the only number available comes from the vendor, activity gets reported as adoption, and the weather takes the credit. Each has a cheap preventive measure, and each is far harder to unwind once a figure has been quoted to a client.

Likelihood: highImpact: high

Denominator drift — the scope grows and the base does not

A package grows 30% through variations while the KPI is still divided by the base agreed at contract award. The numerator rises with the work, the rate looks stable or improves, and nobody notices because the base is buried in a spreadsheet nobody opens. On a job with heavy change, this alone can manufacture an apparent improvement of the same order as the effect being measured.

PreventionRecompute the base every reporting period from the current record, and flag any period where it moved more than a stated threshold.

Likelihood: highImpact: medium

The re-baseline that erases the comparison

The programme is rebuilt — activities merged, renumbered, resequenced — and every metric expressed per activity silently changes meaning. The series continues to plot, which is the dangerous part: a broken comparison that looks continuous is worse than an obvious gap, because it is defended rather than investigated.

PreventionKeep a re-baseline register with dates and before/after activity counts, and mark the affected periods on every chart.

Likelihood: highImpact: high

Survivorship — only the well-run projects volunteer

Pilots are hosted by the project directors who are already ahead: good design maturity, a stable supply chain, an engaged client. Their results are then generalised to the portfolio. The measured effect includes the selection, and the rollout to the harder half of the portfolio underperforms by a margin nobody predicted, which reads internally as the technology failing.

PreventionInclude at least one deliberately difficult project in any measured pilot, and report its result separately rather than in the average.

Likelihood: highImpact: medium

The vendor's number is the only number

The reported figure comes from the tool's dashboard, computed with the vendor's numerator, the vendor's base and the licence period as the reporting period. It cannot be reproduced from records you retain, and it can change with a product release without anyone being told. At renewal, the only evidence available is supplied by the party being renewed.

PreventionCompute the contractual metric yourself from the register, and use the vendor's log only to identify which records the tool touched.

Likelihood: highImpact: medium

Activity reported as adoption

Seats, logins, documents processed and prompts issued get reported as adoption metrics. They rise with project size, with training campaigns and with rework, so they can go up for reasons that are neutral, good or actively bad — and nothing in the reporting distinguishes them. Any incentive attached to these counts destroys their remaining information value.

PreventionPair every usage count with a coverage rate against a defined population, and never report a count as a benefit.

Likelihood: mediumImpact: high

The weather, the designer and the subcontractor take the credit

A wet winter, a novated designer who is unusually responsive, or a strong groundworks subcontractor will move closure times, non-conformance rates and float consumption more than most tools do. Without a comparison group these effects are indistinguishable from the intervention, and they cut both ways — programmes have been cancelled on a bad quarter that had nothing to do with the model.

PreventionRun a matched cohort and publish the confounder register alongside the number, every time.

There is a seventh failure that deserves separate treatment because it is deliberate rather than structural: gaming. Any metric with consequences attached will be met, and construction teams are extremely good at meeting metrics. Closure-time targets produce queries split into two so each closes faster; non-conformance rates fall when items are logged as observations instead; coverage rates rise when the tool is run over records nobody needed it for. The preventive measure is not surveillance. It is to write the gaming mode down next to the metric when you define it, report the metric with the counter-signal that would reveal it — query volume next to closure time, observation volume next to NCR rate — and never attach an individual incentive to a metric with a cheap gaming route.

What measured AI programmes look like in public

Three publicly reported programmes read against the measurement ladder. None is an Atomic Loops engagement; each links to the firm's own published material.

Public construction AI material is overwhelmingly about capability rather than measurement, which is itself the finding: firms publish what their tools do far more readily than what the tools changed, because the second claim requires a base and a comparison the first does not. The three programmes below are useful precisely because each illustrates one part of the measurement problem — portfolio scale, instrument provenance, and the one KPI family construction already normalises properly.

Three programmes read against the measurement ladder

Outcomes as published by the firms themselves; verify figures against the linked source before reusing them, and note that none of these is an independently audited benefit claim. The images are generated industry scenes from our library, not photographs of the named firms or their projects.

Scene: project team reviewing a chart on a screen and a digital site model in a site office overlooking a construction siteSuffolk ConstructionUS general contractor · national portfolio · in-house technology function24
Challenge
A national contractor's project data sits in as many shapes as it has projects, so any attempt to say whether a technology helps runs into the fact that no two jobs classify or record the same events the same way.
Approach
Suffolk publishes its approach of building an internal data and technology capability across the business — including a venture arm, Suffolk Technologies, investing in construction technology — rather than buying disconnected point solutions project by project.
Reported outcome
As published by the firm, the effort is framed around applying analytics to its own accumulated project data at portfolio scale, with technology treated as a corporate capability rather than a per-project experiment.
What it shows about the curvePortfolio scale is what makes a matched cohort possible at all. A contractor with dozens of comparable jobs can leave some on the previous process and compare; a firm running three projects a year cannot, and has to fall back on package-level comparison within a single job.

Suffolk Construction (opens in a new tab)

Scene: site engineer with a helmet-mounted camera reviewing captured progress imagery on a tablet in front of a partially built concrete frameBuildotsConstruction progress-tracking technology · deployed by contractors internationally23
Challenge
Progress reporting on site is traditionally a judgement — a percentage complete estimated by the person responsible for delivering it, which is the least independent measurement in construction.
Approach
Buildots publishes material describing automated capture of site state through helmet-mounted cameras, compared against the model and the programme to produce a per-activity view of what is actually built.
Reported outcome
The firm's published customer material reports earlier identification of sequencing problems and reduced schedule slippage on residential and commercial projects. These are vendor-published claims rather than independently audited figures.
What it shows about the curveThe instrument can be bought; the denominator cannot. What makes this approach measurable is that the activity list in the programme becomes the base — a defined population of activities rather than a percentage someone estimates. If you buy the instrument and keep reporting the vendor's summary number, you have bought a better sensor and stayed at Counted.

Buildots resources (opens in a new tab)

Scene: site team on a concrete slab at dusk reviewing data displays, with a steel frame, an excavator and survey equipment around themShawmut Design and ConstructionUS construction management firm · fit-out, academic and healthcare work34
Challenge
Safety data on a construction project arrives as thousands of observations of wildly varying severity, and the numbers that matter — reportable incidents — are rare enough that no single project can show a meaningful trend inside a year.
Approach
Shawmut publishes its safety programme and its use of data and analytics on jobsite observation data to direct attention before incidents occur, alongside its published safety performance.
Reported outcome
As published by the firm, safety performance is reported through externally computed, normalised measures — in the US, the experience modification rate, which is calculated against expected losses for the firm's class and payroll rather than against its own history.
What it shows about the curveSafety is the one construction KPI family that is already normalised, externally computed and audited — and that is exactly why it is comparable between firms. An AI adoption KPI built the same way, with a stated population, a stated base and a third party able to recompute it, will be believed for the same reason.

Shawmut Design and Construction (opens in a new tab)

Read together, the three point at the same conclusion from different directions. Measurement capability is a property of the organisation rather than of the tool: it comes from portfolio scale, from owning the register the metric is computed from, and from adopting a grammar that somebody outside the firm can check. None of the three required a novel model, and none of the improvements available to them was blocked by the state of the art in machine learning.

A 90-day plan: a defensible KPI baseline for AI-assisted technical queries

The Counted → Baselined transition made concrete on one civils package — AI-assisted technical query triage, measured properly. Contains no model development.

Ninety days is enough to produce one defensible AI adoption KPI when the scope is a single decision on a single package, and it is nowhere near enough when the scope is a business. The plan below runs the transition on a specific, near-universal construction problem: technical queries on a civils package, where an AI assistant drafts responses and routes queries to the right discipline, and nobody can currently say whether it has changed anything. The AI already works at most firms attempting this, so the quarter contains no model development at all — it is definition, extraction, cohort design and recomputation.

Counted → Baselined on one package, in one quarter

One package, one register, one owner. If a phase runs long, narrow the scope — one discipline instead of three — rather than extending the plan. The output at day 90 is a number and the evidence behind it, not a rollout.

  1. Days 1–15

    Pick the package and write the definition sheet

    Choose one package with an active technical query register — a substructure or drainage package is ideal, because query volume is high and the interfaces are well understood. Write the one-page definition: numerator (median and 90th-percentile closure time in working hours; queries reopened per 100 raised), base (per design package), population (main contractor to designer queries only), exclusions (withdrawn within 24 hours; queries reclassified as change), owner (the project-controls lead, not the digital lead). Get the commercial manager to sign it.

    A signed definition sheet and a named owner

  2. Days 16–35

    Extract the baseline before anything changes

    Pull at least twelve months of the TQ register out of the CDE, apply the exclusion rules, and compute the baseline distribution — not just the mean, which on query data is dominated by a long tail. Date the extract, retain it in a controlled location, and record the CDE export parameters used. If the register does not hold a formal response timestamp, this is where you discover it, and the fix is a workflow change rather than an analytics one.

    A dated, retained pre-AI baseline with its method recorded

  3. Days 36–55

    Choose the cohort and register the confounders

    Identify two or three comparable packages left on the previous process — matched on discipline, designer, contract form and delivery phase — and write down what you could not match on: design maturity at award, weather exposure, subcontractor, whether the designer was novated. This is the phase most teams skip and the one that decides whether the day-90 number is believed. Agree with the designer that the comparison exists, so it does not surface as a surprise at the review.

    A written cohort definition and confounder register

  4. Days 56–75

    Wire the recomputation and prove it reruns

    Move the calculation off the spreadsheet: a scheduled job over retained CDE extracts, producing both the AI-supported and cohort figures with the same code. Have someone outside the team rerun it and reproduce the number. Set the retention rule for the inputs — a reporting period is not long enough; match the contractual retention period for project records.

    A rerunnable job, reproduced by a second person

  5. Days 76–90

    Report the delta with its caveats, and register the definition

    Report both series — AI-supported and cohort — with the confounder register read out alongside, into the monthly project review and the CVR rather than into a digital update. Then do the thing that makes it compound: lodge the definition in the information requirements for the next project so the second job starts with the base already agreed.

    An attributed delta in the commercial pack, and a reusable definition

The order matters more than the speed

  1. Definition before extraction

    Extracting first and defining afterwards guarantees that the definition is shaped by what the data made easy, which is how exclusions end up quietly optimising the result. Write the definition, then find out what it costs to compute it — and if the answer is that the register cannot support it, that is a finding worth having in week two rather than week ten.

  2. Cohort before go-live

    A comparison group chosen after the results are in is not a comparison group. Pick the packages while nobody knows which way the number will go, and record the choice. This single discipline does more for the credibility of construction AI reporting than any amount of statistical sophistication applied afterwards.

  3. One package before one business

    The pressure at day 90 will be to mandate the metric across the portfolio. Resist it for one more project. Reproducing the definition on a deliberately different job — different sector, different client, different designer — is what reveals which parts of it were assumptions, and it is far cheaper to find that out on a second project than across thirty.

  4. Register the definition where the next project will find it

    A definition that lives with the team dies with the team. Put it where project information requirements live — the ISO 19650 (opens in a new tab) information-management structure the business already runs, referenced from the EIR and the BEP — so the next project inherits an agreed base instead of restarting the argument.

One caution about scope. This plan deliberately measures a decision, not a tool. If the AI assistant is also being used for meeting notes, document search and drafting, those uses are out of scope for these ninety days and should stay out — they have different populations, different bases and different owners, and folding them in produces a composite number that means nothing. Measure one decision properly and the second is a fortnight's work; measure everything at once and nothing survives its first challenge.

The measurement stack, layer by layer

What has to exist under an AI adoption KPI, which layer each stage first requires, and the instrumentation sheet for the core construction metrics.

A defensible AI adoption KPI needs five layers underneath it, and construction firms usually have the first two already — they were built for information management and cost control, not for measuring technology. The stack below is annotated with the ladder stage that first requires each layer, which is the useful reading: a firm trying to report a portfolio figure without the attribution layer is producing a Counted number with a Normalised presentation, which is the most expensive kind of measurement error.

Layers required by measurement stage

Nothing in this stack is a product. Every layer is defined by what it must guarantee, and in most construction businesses the bottom two already exist and have never been pointed at performance reporting.

  1. Registers that already exist

    Stage 1+

    • CDEInformation containers, TQ/RFI register, drawing revisions
    • Cost ledger and CVRCertified value, variations, disputed value
    • ProgrammeActivities, float, baselines and re-baselines
    • Field and quality appsInspections, ITPs, NCRs, observations
  2. Identifiers and classification

    Stage 2+

    • Stable record identityOne query, one identifier, across systems
    • ClassificationConsistent coding of work sections and asset types
    • Package and cost-code mappingThe join between commercial and delivery views
  3. Metric definition layer

    Stage 3+

    • Definition sheetsNumerator, base, population, period, exclusions, owner
    • Version controlChanges visible in a diff, with a stated reason
    • Retention ruleInputs kept for the contractual retention period
  4. Attribution layer

    Stage 4+

    • Dated baselinesExtracted before go-live, retained, method recorded
    • Cohort definitionMatching variables written down before results exist
    • Confounder and re-baseline registersWhat you could not control, and when the schedule moved
  5. Assurance layer

    Stage 5+

    • Recomputation routeA third party can reproduce the figure from retained inputs
    • Evidence packDefinition version, extracts, script, result, differences explained
    • Management-system controlOwner, review cycle, change log, contractual reference

Pipeline described

  1. Registers that already exist (stage 1+) — CDE: Information containers, TQ/RFI register, drawing revisions; Cost ledger and CVR: Certified value, variations, disputed value; Programme: Activities, float, baselines and re-baselines; Field and quality apps: Inspections, ITPs, NCRs, observations
  2. Identifiers and classification (stage 2+) — Stable record identity: One query, one identifier, across systems; Classification: Consistent coding of work sections and asset types; Package and cost-code mapping: The join between commercial and delivery views
  3. Metric definition layer (stage 3+) — Definition sheets: Numerator, base, population, period, exclusions, owner; Version control: Changes visible in a diff, with a stated reason; Retention rule: Inputs kept for the contractual retention period
  4. Attribution layer (stage 4+) — Dated baselines: Extracted before go-live, retained, method recorded; Cohort definition: Matching variables written down before results exist; Confounder and re-baseline registers: What you could not control, and when the schedule moved
  5. Assurance layer (stage 5+) — Recomputation route: A third party can reproduce the figure from retained inputs; Evidence pack: Definition version, extracts, script, result, differences explained; Management-system control: Owner, review cycle, change log, contractual reference
Step-by-step insights
Registers — the measurement asset construction already owns
The single most under-used fact in construction analytics is that the registers required for good AI measurement already exist and are contractually maintained. The TQ register, the NCR register, the inspection record, the variation log and the certified-value ledger are kept because the contract requires them, with timestamps, owners and a formal status model. That means the baseline for most AI adoption KPIs is already sitting in the CDE, whether or not anyone has extracted it. Firms routinely commission new data capture for a metric whose history is already twelve months deep in a register they are contractually obliged to maintain.
Identifiers — you cannot normalise what you cannot key
The layer that quietly blocks everything above it is identity. The same technical query exists as a thread in the AI tool, a record in the CDE and a line in the designer's own log, and unless one identifier ties them together, every cross-system metric becomes a manual reconciliation that nobody repeats. Classification matters for the same reason: if two projects code work sections differently, no portfolio metric can be computed without a mapping table, and mapping tables that live in a spreadsheet decay within two quarters. This is unglamorous integration work and it is the actual prerequisite for stage 4.
Definition layer — put the metric in version control, not in a spreadsheet
Definitions belong under the same change control as any other project information. Version them, record why each version changed, and keep the retention rule with the definition rather than in someone's head. Firms already operating an information management function under ISO 19650 have the naming, status and retention machinery to do this immediately; the step is administrative rather than technical. The test of whether the layer exists is whether you can state, without asking anyone, what the metric meant fourteen months ago.
Attribution — the layer that turns a change into a claim
Attribution is where construction measurement is genuinely harder than in process industries, and the honest response is to be explicit rather than clever. A dated pre-go-live baseline, a cohort chosen before results exist, and a written register of what could not be matched will support a modest claim indefinitely. Statistical adjustment applied after the fact to a poorly designed comparison will not, however sophisticated it is, because the objection is not technical — it is that the choice was made once the answer was known.
Assurance — the layer that lets the number leave the building
The assurance layer exists so somebody outside your organisation can rely on the figure: a framework client, an insurer, an auditor, a prospective partner in a joint venture. Its content is mundane — the definition at the version in force, the retained extracts with record identifiers, the script, the result and an explanation of any difference on recomputation. Firms adopting a formal AI management system in the shape of ISO/IEC 42001 will find this layer maps directly onto its control expectations, and firms adopting the NIST AI Risk Management Framework will recognise it as the measurement function made concrete.

The layer most often skipped is identifiers, and skipping it is what makes the portfolio number arrive two years late. It looks like plumbing that can be deferred until the metrics prove themselves, but every deferred identifier becomes a manual reconciliation, and manual reconciliations do not survive a busy quarter on a live job. The instrumentation sheet below is the practical output of the stack: for each core metric, where it is computed from, what it is divided by, and the stage at which it starts telling the truth.

None of the top three layers requires a new standard to be invented. UK firms already work to ISO 19650 in its national form, published by BSI (opens in a new tab) as BS EN ISO 19650, which supplies naming, status and retention conventions the definition layer can adopt directly. Firms formalising AI governance have two published frames to hang the assurance layer on — ISO/IEC 42001 (opens in a new tab), the AI management system standard, and the NIST AI Risk Management Framework (opens in a new tab), whose measure function is close to a specification for exactly this stack. Neither obliges anyone to report AI adoption KPIs; both make it awkward to report a figure nobody can reproduce.

KPIFormula / readSource registerBaseCadenceHonest from
Register coverageRecords the tool touched ÷ records in the defined populationCDE register plus the tool's record logDefined population, per packageWeeklyStage 2
Output acceptance rateOutputs accepted without material edit ÷ outputs producedTool log joined to the register's final recordPer 100 outputsWeeklyStage 3
TQ closure timeMedian and P90 hours from raise to formal response, working calendarTQ / RFI register in the CDEPer design packageWeeklyStage 3
Query rework rateQueries reopened or superseded ÷ queries raisedTQ / RFI registerPer 100 raisedMonthlyStage 3
Variation assessment timeDays from notification to assessed positionCost ledger and variation logPer £m of certified valueMonthly (CVR cycle)Stage 3
Non-conformance intensityNCRs raised ÷ inspections recordedITP and NCR registersPer 1,000 inspectionsMonthlyStage 3
Package schedule performanceEarned value ÷ planned value at package levelProgramme plus valuation recordsPer 1,000 programme activitiesMonthlyStage 4
Observation closure rateObservations closed within target ÷ observations raisedH&S management systemPer 100,000 worked hoursMonthlyStage 4
Handover data completenessAssets meeting the AIR ÷ assets required at handoverCDE and asset information modelPer asset record, by classificationPer information exchangeStage 4
Instrumentation sheet for the core construction AI adoption KPIs. Every metric is computed from a contractually maintained register, never from the AI tool's own log.

The stage transitions themselves are verifiable, and it is worth checking them against evidence rather than opinion once a year. The four measures below separate each stage from the one beneath it, and all four are answerable from documents rather than from systems — which is a fair summary of this entire page.

MeasureCountedBaselinedNormalisedHow to check it
Definition statusNone writtenOne page, signedVersioned, per domainAsk for the document and look at its revision history
Baseline provenanceEstimatedDated pre-go-live extractExtract plus method, retainedCheck the extract date against the tool's go-live date
Comparison groupNonePackages within one projectMatched cohort across schemesAsk for the matching variables in writing
RecomputabilityNot possiblePossible with the authorDone, by someone elseAsk when the last recomputation ran and who ran it
Verification measures for each transition on the measurement ladder. All four are checkable in an afternoon by someone outside the reporting team.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Adoption KPI
A metric that measures whether an AI capability has entered the work — coverage of a defined register, time from event to usable answer, share of outputs accepted without material edit. It says nothing about value on its own and must always be reported alongside a value KPI.
Value KPI
A metric expressed in the project's own units — days, pounds, non-conformances, worked hours — that measures what changed. In construction it is only honest when reported with its base, its baseline and its comparison group.
Normalisation base
The denominator a KPI is divided by so it survives differences in project size and type: per £m of certified value, per 1,000 programme activities, per 100,000 worked hours, per design package, per inspection. Chosen per domain from what actually drives the numerator.
Denominator drift
The silent failure in which scope grows through variations while the base stays at the figure agreed at contract award, so the reported rate improves without anything improving. Prevented by recomputing the base every reporting period.
Matched cohort
Construction's substitute for a holdout: comparable packages or schemes left on the previous process, matched on the variables that drive the metric — discipline, designer, contract form, phase, value band — with the unmatched variables written into a confounder register.
Confounder register
The written list of things that could not be matched or controlled — weather, ground conditions, design maturity at award, subcontractor capability — read out alongside the reported number. Publishing it is what makes a modest delta believable.
Metric definition sheet
The one-page controlled document that fixes a KPI: numerator, base, population, period, exclusions, owner and review date. It is the artefact that separates the Anecdotal stage from every stage above it.
Recomputability
The property that someone outside the reporting team can reproduce a published figure from retained inputs and a written method. It is the practical test of whether a number is a KPI or a claim, and it forces a retention rule.
Re-baseline register
A record of every occasion the programme was rebuilt, with dates and the activity count before and after, so metrics expressed per programme activity can be restated or flagged rather than silently changing meaning.
Vanity count
An absolute activity figure reported as adoption — seats, logins, documents processed, prompts issued. It rises with project size, with training campaigns and with rework, so it cannot distinguish good causes from bad ones.
Gaming mode
The cheapest way to move a metric without doing the work it is meant to represent — splitting queries to shorten closure time, logging non-conformances as observations, running the tool over records that did not need it. Written down beside the metric at definition time.
Evidence pack
The bundle that allows a third party to rely on a reported figure: the definition at the version in force, the retained extracts with record identifiers, the computation, the result, and an explanation of any difference found on recomputation.

Frequently asked questions

The questions construction and infrastructure teams ask most often when they try to prove an AI programme is working.

What are AI adoption KPIs in construction?

They are the metrics that show an AI capability has changed how a project is delivered, rather than showing that software was bought. They split into adoption KPIs — register coverage, time from site event to usable answer, output acceptance rate — and value KPIs expressed in the project's own units, such as technical query closure hours, non-conformances per 1,000 inspections, days to assess a variation, or cost of rework per £m certified. Neither family means anything alone; reported together, with a base and a comparison, they become defensible.

Why do construction AI KPIs fail more often than in other industries?

Because construction has no stable denominator. A factory measures per unit or per shift for years; a construction business is a portfolio of temporary organisations, each with its own package breakdown, designer, subcontract mix, weather and re-baselined programme. A number computed on one job therefore has no natural counterpart on the next. The failure is not analytical sophistication but agreement: nobody has chosen a base, a population and a set of exclusions and written them down where the next project can find them.

What is the right denominator for a construction AI KPI?

It depends on what generates the events you are counting, and it should be chosen per domain rather than universally. Design-side metrics normalise per design package or per 1,000 model elements, because query volume tracks interface count. Commercial metrics normalise per £m of certified value. Production metrics normalise per 1,000 programme activities. Quality metrics normalise per 1,000 inspections. Safety already has its base — per 100,000 worked hours. Contract value applied to everything is the classic error, because information intensity does not scale with value.

How do you get a counterfactual when every project is different?

Use a matched cohort rather than a holdout, because you cannot randomise a bridge. Identify comparable packages or schemes left on the previous process, matched on the variables that drive the metric — discipline, designer, contract form, delivery phase, value band — and choose them before any results exist. Then write down what you could not match on, typically weather, ground conditions, design maturity at award and subcontractor capability, and read that register out alongside the number. Published confounders are what make a modest delta credible.

Which AI adoption KPIs can we honestly report during a pilot?

Coverage and acceptance, and nothing about value. Coverage is the share of a defined population — say, main contractor to designer technical queries on one package — that the tool actually touched. Acceptance is the share of its outputs used without material edit by the responsible engineer. Both are real, both are computable from the register, and neither claims a benefit. Any value figure during a pilot is a projection and should be labelled as one, because the baseline and the comparison group are usually still being argued about.

How many AI adoption KPIs should a construction programme carry?

Two or three per decision, and no portfolio composite. One coverage measure, one acceptance or quality measure, and one value measure in the project's units is enough to describe whether a decision has changed and whether the change is worth anything. Composite indices — a digital maturity score, an adoption index — are appealing to report and impossible to act on, because a movement in the composite never tells you which component moved or what to do about it. Keep the components visible.

How do you stop construction AI KPIs from being gamed?

Write the gaming mode down when you define the metric, and report the counter-signal beside it. Closure-time targets invite queries to be split so each closes faster, so report query volume next to closure time. Non-conformance rates invite items to be logged as observations, so report observation volume next to the NCR rate. Coverage targets invite the tool to be run over records that did not need it, so report coverage against a defined population rather than a growing one. And never attach an individual incentive to a metric with a cheap gaming route.

What does a defensible baseline look like on a live project?

A dated extract taken from the register before the tool went live, with the exclusion rules applied and the extract parameters recorded, retained somewhere controlled. For technical queries that means twelve months or more of the TQ register from the CDE, with the distribution computed rather than just the mean, because query data has a long tail that averages hide. The two things that make a baseline indefensible are reconstructing it after go-live and taking it from a system whose records have since been reclassified or re-baselined.

Who should own AI adoption KPIs — digital, commercial or project controls?

Project controls should own the computation, commercial should own the definition, and digital should own neither. That split matters because the arguments that decide whether a number is believed — exclusions, population, base — are commercial arguments, and a metric defined by the team that bought the tool is discounted on sight. The digital function's job is to make the calculation possible and rerunnable. In practice the strongest arrangement we see is a project-controls lead as named owner with the commercial manager's signature on the definition sheet.

How do AI adoption KPIs relate to ISO 19650 and the CDE?

Directly, and the relationship is an advantage. ISO 19650 already gives a project a common data environment with information containers, status codes, naming conventions and retention expectations — which is precisely the machinery a recomputable KPI needs. The registers the metric is computed from are contractually maintained inside it, so the baseline usually already exists. The step most firms have not taken is lodging the metric definitions in the same information structure, referenced from the exchange information requirements, so the next project inherits an agreed base instead of restarting the argument.

Should AI adoption KPIs be written into the contract or the EIR?

Into the information requirements first, and into the contract only once the definitions have survived two projects. Putting a measure in the exchange information requirements makes the base and the population part of what the supply chain is asked to deliver, which is where they belong. Putting an immature definition into a performance schedule or a NEC secondary option creates a target before the measurement is trustworthy, and the target will be met — usually through the gaming route the definition had not yet identified. Contractualise last, not first.

How long before an AI adoption KPI is trustworthy enough for a board pack?

About a quarter for a single decision on a single package, and roughly a year before a portfolio figure is worth putting in front of a client. The 90-day path is definition, dated baseline, cohort selection, rerunnable computation and a report into the commercial forum. What extends it beyond that is not analysis but agreement — reproducing the definition on a structurally different project, standardising classification across jobs, and building the retention and version control that let someone outside the team recompute the figure.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for construction, infrastructure and heavy industry — document and query automation, progress and quality capture, commercial analytics — integrated into the CDE, the cost ledger and the programme rather than delivered as a separate dashboard.

  • · Delivery on live civils, rail and building programmes, inside the client's information standard
  • · Measurement design run jointly with commercial and project-controls teams
  • · Recomputable metrics: definitions versioned, inputs retained, third-party checkable
  • · 22 cited sources on this page

Sources

  1. ISOISO 19650-1 — organization and digitization of information about buildings and civil engineering works (opens in a new tab)
  2. ISOISO/IEC 42001 — artificial intelligence management system (opens in a new tab)
  3. UK BIM FrameworkUK BIM Framework — information management guidance (opens in a new tab)
  4. BSIStandards and information services (opens in a new tab)
  5. NECThe NEC suite of contracts (opens in a new tab)
  6. RICSRICS standards and guidance (opens in a new tab)
  7. Constructing ExcellenceConstruction industry KPIs (opens in a new tab)
  8. Get It Right InitiativeResearch on the cost of avoidable error in construction (opens in a new tab)
  9. Association for Project ManagementEarned value management resources (opens in a new tab)
  10. UK Government (Cabinet Office)The Construction Playbook (opens in a new tab)
  11. UK GovernmentInfrastructure and Projects Authority (opens in a new tab)
  12. National Audit OfficeReports on infrastructure and major project delivery (opens in a new tab)
  13. Office for National StatisticsConstruction industry statistics (opens in a new tab)
  14. HSEConstruction health and safety guidance (opens in a new tab)
  15. HSERIDDOR — reporting of injuries, diseases and dangerous occurrences (opens in a new tab)
  16. HSEIndustry health and safety statistics (opens in a new tab)
  17. Information Commissioner's OfficeGuidance on AI and data protection (opens in a new tab)
  18. European CommissionRegulatory framework for AI (the AI Act) (opens in a new tab)
  19. NISTAI Risk Management Framework (opens in a new tab)
  20. Suffolk ConstructionCompany information and published approach (opens in a new tab)
  21. BuildotsPublished resources and customer material (opens in a new tab)
  22. Shawmut Design and ConstructionCompany information and safety programme (opens in a new tab)

Find out whether your AI numbers would survive a challenge

We run the assessment with your commercial and project-controls leads, attempt to recompute one figure you have already reported, and leave you with a definition sheet and a costed 90-day plan for your weakest dimension. You keep both whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.