Redefining Technology

Construction & InfrastructureRegulations, Compliance & Governance

Project AI NIST compliance in construction and infrastructure: running the AI RMF project by project

Project AI NIST compliance is the practice of running the NIST AI Risk Management Framework at the level of one construction or infrastructure project rather than the enterprise. It produces a written project profile — which AI is in use, what each use is for, how it is measured, what happens when it is wrong — that survives handover.

Generated scene: a project delivery team on a concrete-frame site reviewing AI analytics overlays on tablets and laptops
Construction & Infrastructure · Regulations, Compliance & Governance

Key takeaways

  1. The NIST AI RMF is voluntary, non-certifiable and use-case agnostic. There is no such thing as a NIST-certified project — what the framework produces is a written profile and a record, and on a construction project that record is the deliverable, not a badge.
  2. The framework assumes a persistent organisation. Construction delivers through temporary ones, so GOVERN lives in the firm while MAP, MEASURE and MANAGE happen inside a project that will not exist in three years. That mismatch, not the framework's content, is what makes project AI compliance hard.
  3. The AI RMF Core carries 4 functions, 19 categories and 72 subcategories. A usable project profile triages that down to the handful that apply per AI use and names the existing project artefact — the PEP, the RACI, the risk register, the method statement — that already carries each one.
  4. MEASURE is where projects fail. Reliance is placed on the vendor's published benchmark rather than on a number measured against this project's own held-out records and seeded defects, and the gap only becomes visible when something is missed.
  5. The evidence pack is worth more at handover than during delivery. The asset owner inherits every consequence of the project's AI and none of the project's memory, so a profile that is not a contracted information deliverable is a profile that dies at practical completion.

Abbreviations used on this page

AI RMF
NIST Artificial Intelligence Risk Management Framework (NIST AI 100-1, January 2023)
GAI
Generative AI — the subject of NIST's separate Generative AI Profile, NIST AI 600-1
TEVV
Test, evaluation, verification and validation — NIST's term for the evidence behind a performance claim
PEP
Project execution plan
BEP
BIM execution plan (the ISO 19650 information-delivery plan)
CDE
Common data environment (the ISO 19650 project information store)
JV
Joint venture — the temporary company two or more contractors form to deliver one project
EWN
Early warning notice (the NEC contract mechanism for raising an emerging risk)
NCR
Non-conformance report
RFI
Request for information (the formal query from site to the design team)
CDM
Construction (Design and Management) Regulations 2015
O&M
Operation and maintenance information — what the asset owner receives at completion

Free · 8 questions · ~3 minutes

Score one project against the AI RMF ladder

Eight questions about one live project, one at a time, about three minutes. Answer them and we build your personalised profile report — your rung on the ladder, your score on each of the four dimensions, and the specific artefact standing between you and the next rung — and send it to your inbox. Score the project, not the firm: the answers differ, and the project's answer is the one a reviewer will ask for.

0 of 8 answered

Question 1 of 8Register and intended purpose

What does this project hold that lists the AI in use on it?

MAP has nothing to attach to without a project-level inventory, and a firm-level tool list is not one — it omits embedded and client-specified AI, which is where most untracked reliance lives.

How the score maps to a stage
  • 04 — Stage 1, Referenced. Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
  • 510 — Stage 2, Profiled. Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
  • 1116 — Stage 3, Measured. Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
  • 1721 — Stage 4, Managed. Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
  • 2224 — Stage 5, Transferable. Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.

What project-level NIST AI RMF compliance actually is

A definition, the four functions, and the diagram that shows which of them belong to the firm and which belong to a project that will not exist in three years.

Project-level NIST AI RMF compliance is the practice of running the NIST AI Risk Management Framework (opens in a new tab) inside a single construction or infrastructure project, and producing artefacts that a reviewer could ask that project for. Not the firm's policy — the project's register, its intended-purpose statements, its test records, its reliance decisions and its deviation history. The framework is voluntary, non-sector-specific and use-case agnostic by design, which means it tells you what has to be true and leaves you to name the artefact that makes it true. On a project, that artefact almost always already exists.

Two things about the framework matter before anything else. First, it is not certifiable: there is no NIST audit, no certificate and no accredited body issuing one, and any supplier claiming to be 'NIST AI RMF certified' is describing a self-assessment. Certification is the domain of ISO/IEC 42001 (opens in a new tab), which is a different instrument for a different purpose. Second, the framework is organised into four functions — GOVERN, MAP, MEASURE and MANAGE — deliberately echoing the structure of NIST's older and far more widely deployed Cybersecurity Framework (opens in a new tab), which many construction IT and information-security teams already work to. Of the four, GOVERN is explicitly cross-cutting and the other three are context-bound. That distinction is the whole of this page's argument, because in construction GOVERN belongs to a company that persists and MAP, MEASURE and MANAGE belong to a project organisation that does not.

The published Core carries 4 functions, 19 categories and 72 subcategories, with a companion Playbook (opens in a new tab) of suggested actions and a set of crosswalks (opens in a new tab) to other regimes. Read end to end that is a formidable document and a poor project instrument — nobody mobilising a viaduct package is going to work through 72 subcategories per tool. The usable move is triage: for each AI use on the project, identify the handful of subcategories that genuinely bite, and name the existing project artefact that already carries each one. The diagram below is the shape that triage takes.

Where the four AI RMF functions attach to a construction project

Three organisations, one framework. The enterprise persists and owns GOVERN. The project is a temporary organisation and owns MAP, MEASURE and MANAGE. The asset owner sets requirements before the project exists and inherits its consequences afterwards. Compliance fails at the two boundaries, not in the middle lane.

  • Human in the loop
  • Data & feeds
  • AI / model
  • System-of-record action
  • Where value leaks

The process, in words

  • The enterprise lane is GOVERN and only GOVERN: a board-owned AI policy with a risk appetite, an approved tool catalogue that says what may be used and where, and a profile template that says what every project must fill in. None of these is project-specific and none of them, on their own, produces a single piece of project evidence.
  • The project lane is where compliance actually happens, in order. At mobilisation the project writes its AI register (MAP 1). Each row gets an intended-purpose statement and a consequence class taken from the project's own risk matrix (MAP 2–5). Each high-consequence row gets tested against held-out project records and seeded defects (MEASURE). The test result produces a written reliance decision — what the output may and may not be used for — and everything outside that decision routes to the project's existing deviation and early-warning machinery (MANAGE).
  • The asset owner's lane brackets the project at both ends. Before the project exists, the employer's requirements set which AI is mandated, permitted or excluded and what information must be delivered. After it ends, the operating duty holder inherits an asset whose design, inspection and commissioning records were partly AI-assisted, and has no access to the project's reasoning except through the evidence pack.
  • The two dashed returns are what turns a project exercise into an organisational capability: escalations from the project's deviation route reach portfolio review, and the completed evidence pack goes back to the enterprise so the next mobilisation starts from a filled-in profile rather than a blank template.
Step-by-step insights
Why GOVERN in the enterprise lane is necessary and not sufficient
NIST places GOVERN at the centre of the framework deliberately: it is the function that infuses the other three, covering culture, accountability, policy and third-party risk. In a firm with a stable product this works cleanly, because the same people who set the policy also run the systems it governs. In construction the two are separated by a contract, a site and often a joint-venture boundary. The result is a predictable failure: the enterprise completes GOVERN honestly, the projects are told they are covered, and no project artefact is ever created. The policy is a floor, not a profile. Treat it as an input to the project lane — a source of constraints — and judge compliance by what exists on the project.
The register is written at mobilisation because that is the only cheap moment
Mobilisation is the one point in a project's life when everybody expects to fill things in: the PEP is being written, the RACI is being agreed, the CDE is being configured, method statements are being drafted. Adding a register row per AI use costs minutes then and days later, because by month six the tools have accumulated informally and finding them means interviewing package leads. The register also has to catch three arrival routes that nobody thinks of as procurement: embedded features inside tools already in use, client-specified platforms named in the employer's requirements, and individual subscriptions bought on expenses. Two of those three the project did not choose, and both are on its risk.
Intended purpose and consequence class do the triage the framework cannot
72 subcategories is unusable per tool; two sentences per tool is not. The intended-purpose statement says what this output is for on this project — 'flags probable clashes for engineer review, does not approve designs'. The consequence class says what happens if it is wrong, expressed on the project's own existing risk matrix rather than an invented AI-specific scale. Those two fields together determine everything downstream: how much test evidence the use has to carry, whether a DPIA is needed, whether a second human check stays in the method statement, and which subcategories are worth writing against at all. Projects that skip straight to controls end up applying the same weight to a document summariser and a lifting-plan checker.
MEASURE on project data — the step that is almost always skipped
The single most common gap in the middle lane is that no number in the register was produced by the project. Vendor benchmarks are computed on the vendor's corpus, under the vendor's definition of a correct answer, and averaged across customers whose sites, standards and design conventions differ from yours. They are useful for shortlisting and useless for reliance. A project test set is a records exercise, not a research one: pull known-bad cases from your own NCR log, RFI log, snagging list and design-review tracker, hold them out, add a controlled set of deliberately seeded defects, and score. NIST's term for this whole family of activity is TEVV, and the framework is explicit that its parameters must be aligned to actual deployment conditions.
The reliance decision is the artefact a reviewer will ask for
Registers describe and tests measure, but neither says what the project decided. The reliance decision closes that gap in one dated line: given this measured performance, this output may be used for X and may not be used for Y, and here is the residual control. It is the construction equivalent of a design-check certificate — an explicit statement of what has been accepted, by whom, on what evidence. It is also the artefact that protects individuals: without it, an engineer who acted on a flagged clash is carrying a judgement nobody recorded, and a project manager who overrode one is carrying a different one.
Both boundaries fail silently, which is why the dashed edges matter
Nothing on a project announces that the handover boundary has failed. Practical completion happens, the team demobilises, the CDE is archived, and the omission surfaces years later as an unanswerable query. The mobilisation boundary fails the same way: the next project starts, a blank register is issued, and the firm pays the stage-1 cost again while believing it has an AI capability. The fix in both cases is contractual rather than cultural — name the evidence pack in the information-delivery plan so it is a deliverable with a date, and make portfolio review of returned profiles a scheduled item with an owner. Neither is expensive; both are invisible until they are missing.

The ladder: from Referenced to Transferable

Five rungs describing how far the framework has actually travelled into a project — what each looks like on the ground, the signals a reviewer can check in an afternoon, the anti-pattern that traps projects there, and what leaving costs.

The ladder measures how far the framework has travelled from a document into a project's own machinery. It runs Referenced → Profiled → Measured → Managed → Transferable, and the order is not arbitrary: each rung is the function that becomes possible once the previous one has produced an artefact. You cannot measure what is not registered, you cannot rehearse a response to a failure you cannot detect, and you cannot hand over a record that was never assembled.

Read the rungs as descriptions of a project rather than of a firm. A contractor with a mature enterprise policy will have projects spread across three rungs simultaneously, and the useful question is always which rung the scheme in front of you is on. Each panel below is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run this week against your own CDE and risk register, and the anti-pattern is the specific mistake most often made trying to leave that rung.

Defensible reliance released against position on the ladder

The curve is not linear. Referenced and Profiled release very little defensible reliance — a register describes untested confidence — and the inflection is at Measured, when the first number the project can actually interrogate exists. Transferable adds little inside the project and everything after it.

Defensible reliance on project AI by stage

  • Stage 1 · Referenced — 36% of operators. Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.
  • Stage 2 · Profiled — 30% of operators. Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.
  • Stage 3 · Measured — 20% of operators. Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.
  • Stage 4 · Managed — 11% of operators. Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.
  • Stage 5 · Transferable — 3% of operators. Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.

Curve shape: logistic, plotted from the stage data above. Distribution: Structured on the NIST AI RMF Core functions (NIST AI 100-1).

Select a rung

Every rung's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Referenced

36% of operators sit here

Referenced is the stage where the NIST AI RMF is named in a bid answer or a corporate policy and appears nowhere on the project itself.

Referenced is not ignorance. It is usually the opposite: someone in the business has read the framework properly, written a good paragraph about it, and won work with it. The gap is structural rather than intellectual — the paragraph was written by a bid team in a head office, and the thing it describes has to happen inside a project organisation that did not exist when the bid was submitted and that answers to a different director.

What makes this stage hard to see from the inside is that everything about it looks compliant. There is a policy. There is a named framework. There may even be a slide in the induction pack. What is missing is any artefact on the project that a reviewer could ask for: no register, no intended-purpose statement, no measurement plan, no owner. Ask for the project's AI risk position and you get a firm-level document plus an assurance that the project follows it.

The cost of staying here is not a fine — the AI RMF carries no penalty, because it is voluntary. The cost is that reliance accumulates invisibly. Site teams start acting on outputs they have never tested, design teams start submitting drawings a tool helped produce, and commercial teams start valuing progress from an analytic nobody has calibrated. None of that is recorded, so when a question arrives — from a client, an insurer, a coroner or a regulator — the project has to reconstruct from memory what it was relying on, and memory has already demobilised.

In practice

The tender answer nobody read again

A tier-one contractor won a place on a five-year highways framework partly on a digital submission that stated its AI use 'aligns with the NIST AI Risk Management Framework'. Eighteen months into the first scheme, a walk of the project turned up eleven distinct AI tools in daily use: two progress-analytics platforms, a document-summarising assistant licensed by the design team, an embedded scheduling optimiser in the planning software, telematics anomaly alerts on hired plant, and six individual generative subscriptions bought on expenses. The only artefact anyone could produce was the tender paragraph.

What it looks like

  • The framework is cited in a PQQ or tender response; no project artefact exists
  • Nobody on the project team can say how many AI tools are running on it
  • AI arrives by individual subscription, embedded vendor feature and client-specified platform, untracked
  • The words 'AI' and 'machine learning' do not appear in the PEP, the risk register or the RACI

Diagnostic signals you can check this week

  • Search the project execution plan for 'AI' and for 'machine learning', and count the hits. Zero is the common result
  • Ask the document controller how many documents in the CDE were AI-assisted. If the question is novel, you are here
  • Ask the project director who owns AI risk on this project. A firm-level name is not a project answer
  • Ask the bid team what was promised about AI in the winning submission, then ask the project team whether they have seen it

Anti-pattern · Adopting the framework at the enterprise and calling the projects done

The instinctive move is a corporate programme: map the whole AI RMF Core at group level, publish a policy, brief the business. It is genuinely useful work and it does not move a single project off this stage, because the framework's MAP, MEASURE and MANAGE functions are all context-bound — intended purpose, affected people and consequence all differ per project, per package and per contract. A group-level map of 72 subcategories with no project register underneath it is an enterprise artefact describing an empty set. Write the register on one live project first; the corporate map then has something real to generalise from.

What holds you here

There is no project-level record, so none of the framework's functions has anything to attach to.

Highest-leverage next move

Write the register: every AI use on one live project, each with a named owner and a one-paragraph intended purpose.

Cost of leaving

Effort
4–8 weeks
Team
A digital lead and a project manager, part-time, plus an hour each from six package leads
Risk
Low — the output is a register and a template; nothing about delivery changes yet
To next stage
1–3 months

If this is you, the next step is

A two-week walk of one live project: every AI use found, named and owned.

Register one project's AI in a fortnight

Stage 2

Profiled

30% of operators sit here

Profiled is the stage where MAP is done: the project holds a written register of every AI use, what each is for, who it affects and what breaks if it is wrong.

Profiled is the first stage where the framework is doing work rather than being cited. The register is a modest artefact — often a single spreadsheet or a CDE-hosted table with a dozen rows — and it changes the conversation immediately, because for the first time the project can answer the question 'what are you relying on?' with a list rather than an anecdote.

Two entries on that list surprise almost every project. The first is embedded AI: the scheduling optimiser inside the planning tool, the anomaly detection in the plant telematics portal, the auto-classification in the document system. Nobody procured these as AI and nobody thinks of them as AI, so they arrive under the radar and end up carrying real weight. The second is client-specified AI — a platform named in the employer's requirements that the project must use and cannot change. Both belong on the register precisely because the project did not choose them.

The stage's characteristic weakness is that the register describes and does not test. Intended purpose is written from the vendor's description, consequence class is assigned by feel, and the accuracy figure everyone repeats came from a sales deck. That is still a large improvement on stage 1, and it is also a trap: a well-formatted register can make a project feel assured while every number in it is somebody else's. MAP without MEASURE produces confident documentation of untested reliance.

In practice

The register that found the optimiser

A JV delivering a water-treatment upgrade ran a two-week register exercise across its packages. Ten of the twelve rows were expected. The eleventh was the sequencing optimiser built into the planning software, which had been quietly re-ordering resource-levelled activities since mobilisation and whose output fed the programme submitted monthly to the client. Nobody had classified it, nobody owned it, and its intended purpose — as far as the project could establish — was whatever the planner assumed it did. It went onto the register with a consequence class of 'commercial and programme', an owner, and a note that its outputs had never been compared with a hand-levelled baseline.

What it looks like

  • A project AI register exists, dated at mobilisation and owned by a named individual
  • Each row carries a one-paragraph intended-purpose statement written in project language
  • Each row carries a consequence class taken from the project's own risk matrix
  • Embedded and client-specified AI is on the register alongside the tools the project bought

Diagnostic signals you can check this week

  • Open the register and count rows that describe AI the project did not buy. If the answer is zero, the sweep missed the embedded layer
  • Pick a row and read its intended-purpose statement aloud to the package manager. If they disagree with it, the statement was written by the vendor
  • Check whether consequence class uses the project's existing risk matrix or an invented scale. An invented scale will not survive a risk meeting
  • Ask when the register was last updated. If it was mobilisation and the project is now at stage four, it is a snapshot, not a register

Anti-pattern · Grading the register instead of testing it

Once a register exists, the natural next move is to make it better: add columns, add a risk score, add a RAG status, circulate a template across the portfolio. This is administratively satisfying and epistemically empty, because every field in it is still an assertion. A three-column register with one measured number in it is worth more than a fifteen-column register with none. Take the highest-consequence row, build a small test set from the project's own records, and measure it. The register's real value is that it tells you which single row to measure first.

What holds you here

The register describes each AI use but nothing measures it, so reliance still rests on the vendor's numbers.

Highest-leverage next move

Build the project test set — held-out project records plus seeded defects — and set one written threshold per AI use.

Cost of leaving

Effort
2–4 months
Team
A profile owner (usually the digital or design manager), a package manager per row, a data-protection lead for the rows that touch people
Risk
Medium — the first honest test can contradict a claim the project has already relied on
To next stage
3–6 months

If this is you, the next step is

We assemble it from your own NCR, RFI and inspection records — typically three weeks.

Build the test set for your highest-consequence row

Stage 3

Measured

20% of operators sit here

Measured is the stage where MEASURE runs on project data: each AI use has a metric, a threshold and a dated test record taken from this project rather than from a vendor benchmark.

Measured is where the framework starts paying for itself, and the reason is unglamorous: a number measured on this project is arguable in a way a vendor benchmark is not. When a design manager says a checking tool has a 4% miss rate on the project's own seeded non-conformances, the conversation moves immediately to what to do about the 4% — sampling rate, second check, which packages. When the same manager says the vendor reports 96% accuracy, the conversation stops, because nobody in the room can interrogate a number produced on somebody else's data.

The test set is the artefact that defines this stage, and building it is a records exercise rather than a data-science one. Every project already holds a corpus of known-bad examples: the NCR log, the RFI log, the snagging list, the rejected inspection records, the design-review comment tracker. Seeding a held-out set from those, plus a controlled set of defects the team introduces deliberately, gives a recall figure that means something in the project's own units. NIST's term for this whole activity is TEVV, and the AI RMF is explicit that its parameters have to be aligned to real deployment conditions rather than to laboratory ones.

The stage's limitation is that measurement without a rehearsed response is a report, not a control. Projects at this stage can tell you precisely how often a tool is wrong and have no agreed answer for the morning it is wrong about something that matters. The miss gets handled by whoever notices, competently and invisibly, and no record is created — which means the third recurrence of the same failure mode still looks like a first.

In practice

The 4% that changed a sampling rate

On a hospital fit-out, the design team ran an automated standards checker over mechanical and electrical submissions. The project built a held-out set of 120 records — 80 drawn from its own historic NCRs and RFIs, 40 with defects seeded deliberately by two engineers who then swapped sets — and measured recall at 96%, with the misses concentrated in clashes involving builder's work openings. The number itself was reassuring. What changed practice was the concentration: the project kept a manual second check on that one category, dropped the blanket 100% manual re-check everywhere else, and wrote both decisions into the register as a reliance decision with a date.

What it looks like

  • A held-out test set exists per high-consequence AI use, built from the project's own records
  • Each use has a written threshold agreed with the package manager who relies on it
  • Test runs are dated, sampled and recorded, and repeated when the tool or the site changes
  • An override log records what site actually did with the output, not just what was produced

Diagnostic signals you can check this week

  • Ask to see a test record with a date, a sample size and a result. A screenshot of a vendor dashboard is not one
  • Ask where the test cases came from. If the answer is 'the vendor supplied them', the measurement is circular
  • Check whether the threshold is written down and whether the person who relies on the output agreed it
  • Look at the override log. A tool with no overrides at all usually means nobody is checking, not that nothing is wrong

Anti-pattern · Measuring the model instead of the decision

Teams that get good at measurement often start optimising the wrong quantity: overall accuracy across all cases, reported monthly, trending gently upward. On a project, overall accuracy is close to meaningless because the consequences are wildly asymmetric — a missed clash in a riser is not a missed clash in a plant room, and a false positive costs an hour while a false negative can cost a package. Measure per consequence class, report recall on the categories that hurt, and accept a worse headline number in exchange for a defensible one. The AI RMF's own framing is that measurement has to be tied to context and intended use, not to a general benchmark.

What holds you here

Measurement exists but no response is rehearsed, so a bad output is handled by improvisation and leaves no record.

Highest-leverage next move

Wire the AI rows into the project's existing risk machinery — risk register, early warning notice, deviation record — and rehearse one end to end.

Cost of leaving

Effort
4–8 months
Team
Profile owner, a QA or design manager per measured use, two engineers for a week to seed and score the test set
Risk
Medium — measurement can invalidate reliance the project has already priced into its programme
To next stage
4–8 months

If this is you, the next step is

We connect the thresholds to your risk register, EWN route and deviation records.

Wire measurement into the project's risk machinery

Stage 4

Managed

11% of operators sit here

Managed is the stage where MANAGE is live: AI responses are owned, dated and exercised inside the project's existing risk process, and at least one real deviation has been run end to end.

Managed is the stage at which AI stops being a digital-team topic and becomes an ordinary project risk. The signature is boring in the best way: AI items appear on the same weekly risk agenda as temporary works, ground conditions and long-lead procurement, discussed by the same people in the same vocabulary, with the same expectations about owners and dates. Nothing about the framework requires a parallel governance structure, and projects that build one usually find it starved of attention within two months.

The characteristic capability here is rehearsal. Construction is unusually good at this — it already drills fire, rescue from height, confined-space entry and concrete-pour contingency — and the same instinct transfers directly. Running one deliberate exercise, in which a tool is declared unavailable or wrong on a live package and the team executes the fallback, converts a paragraph in the method statement into a known quantity. It also surfaces the thing paper never does: how many downstream artefacts had already absorbed the AI's output before anyone noticed.

The stage's blind spot is time. Everything works while the project exists, because the people who wrote the profile are still in the office and the reasoning is still in their heads. Nothing at this stage has been built to survive demobilisation. The evidence exists but has never been assembled; the register is current but has no consignee; the reasoning behind a reliance decision lives with a design manager who moves to the next scheme in March. A project can be genuinely excellent at MANAGE and still hand over nothing at all.

In practice

The exercised fallback that found six documents

A rail systems project ran a half-day exercise on its document-classification assistant: the tool was declared unavailable at 09:00 and the team executed the manual route in the method statement. The manual route worked. What the exercise found was downstream — six technical submissions already in the client's approval queue carried classification metadata the assistant had produced, and none was flagged as machine-assisted. The fix was a metadata field and a CDE naming rule, agreed the same week. Neither would have surfaced from a document review, because on paper the fallback was already compliant.

What it looks like

  • AI rows sit in the project risk register with owners, dates and movement between reviews
  • A named non-AI fallback exists in the method statement and has been exercised at least once
  • A deviation record and an early warning notice have actually been raised on an AI output
  • Model or version changes from suppliers trigger a re-test rather than a silent update

Diagnostic signals you can check this week

  • Open the last four project risk-register versions and look for AI rows that moved. Static rows mean the register is decorative
  • Ask when the fallback was last exercised, and whether the exercise found anything. 'It has never needed to be' is not an answer
  • Ask what happened the last time a supplier pushed a model update. If nobody noticed, change control does not cover the AI layer
  • Ask who receives the AI evidence at practical completion. At this stage the honest answer is usually nobody

Anti-pattern · Building a parallel AI governance forum

The temptation at this stage is to formalise: a monthly AI risk board, its own minutes, its own escalation path, its own reporting line to the enterprise. On a project it starves. Attendance decays through the second month, the commercial team never joins, and the forum's decisions arrive too late to affect a package that was let six weeks earlier. Everything the framework asks for at MANAGE maps onto machinery projects already run daily — the risk register, the early warning notice, the deviation record, the method statement, the change-control process. Use them. The measure of success is that AI risk becomes indistinguishable in process terms from any other project risk.

What holds you here

Everything works while the project exists; nothing has been built to survive handover, demobilisation or the next mobilisation.

Highest-leverage next move

Make the evidence pack a deliverable: name it in the information-delivery plan and hand it over with the asset information.

Cost of leaving

Effort
6–12 months
Team
Project risk manager, profile owner, package managers, plus commercial input where supplier change control is involved
Risk
Medium — exercising a fallback on a live package costs real hours and occasionally exposes uncomfortable downstream dependency
To next stage
9–18 months

If this is you, the next step is

We draft the information-delivery requirement and the pack's contents list.

Make the evidence pack a contracted deliverable

Stage 5

Transferable

3% of operators sit here

Transferable is the stage where the project's AI record outlives the project: the evidence pack goes to the asset owner at handover and the profile seeds the next mobilisation.

Transferable is a narrow stage and deliberately so. It does not mean the project has automated anything, and it does not mean the AI is more capable than at stage 4. It means one specific thing: that the record of what was relied on, how it was tested and what went wrong has been made into an object that crosses two organisational boundaries — the handover to a permanent asset owner, and the mobilisation of the next temporary organisation.

The handover boundary is the one with statutory weight. In the UK, the Building Safety Act regime made the durability of building information an explicit duty for higher-risk buildings, and the CDM regulations have long required a health and safety file to pass to the client at completion. Neither instrument names AI, and neither needs to: the question a duty holder will ask in year seven is what the design or inspection record was based on, and 'a tool the contractor used, which nobody has any record of' is a bad answer regardless of which regime is asking. The AI evidence pack is the cheapest available insurance against that question, and it costs almost nothing if it is assembled continuously rather than reconstructed at completion.

The mobilisation boundary is the one with commercial weight. A firm whose projects each start from a blank register pays the same setup cost every time and learns nothing across schemes. A firm whose next mobilisation starts from the last project's returned profile — the same tools, the same consequence classes, the same test sets, the same known failure categories — compresses stage 1 and 2 into a fortnight. This is also the only mechanism by which GOVERN becomes real rather than declarative: the enterprise's policy improves because projects hand something back, not because the policy was rewritten.

In practice

The pack that answered a year-three question

A design-and-build contractor handed a school over with an AI evidence pack comprising the project register, seven intended-purpose statements, four test-run records, two reliance decisions and a deviation history with three closed entries. In year three the operator queried an as-built discrepancy in a services zone. The pack showed that the automated clash review had run at RIBA stage 4, that builder's work openings had been a known weak category with a manual second check retained, and that the specific zone had been manually verified on a dated record. The query closed in a fortnight rather than becoming a claim.

What it looks like

  • The AI evidence pack is named in the information-delivery plan as a contracted deliverable
  • The asset owner receives the register, the test records, the reliance decisions and the deviation history
  • Mobilisation on the next project starts from a returned profile, not from a blank template
  • Portfolio review of returned profiles changes the policy, the template and the approved tool catalogue

Diagnostic signals you can check this week

  • Look in the information-delivery plan for the AI evidence pack by name. If it is not a named deliverable, it will not be produced
  • Ask the asset owner's information manager whether they have ever received one, and what they did with it
  • Compare two consecutive mobilisations: did the second start from the first's register or from a blank template?
  • Ask what changed in the corporate policy or the approved tool catalogue as a result of the last project's returned profile

Anti-pattern · Treating the pack as a completion task

The pack looks like a document, so it gets scheduled like one: a task in the handover programme, assigned in the final quarter, to be assembled from whatever can still be found. By then the design manager who wrote the reliance decisions has gone, the test set lives on a demobilised laptop, and half the deviation history was handled verbally. What is produced is a folder that satisfies the deliverable and answers no future question. The pack has to accrete — every register update, test run, reliance decision and deviation filed into it on the day it happens — so that completion is a transmittal rather than an archaeology exercise.

What holds you here

Sustaining transferability is a portfolio discipline — the constraint becomes whether the enterprise actually reads what its projects hand back.

Highest-leverage next move

Close the loop: make portfolio review of returned project profiles a scheduled item that changes the policy, the template and the tool catalogue.

Cost of leaving

Effort
Continuous
Team
Project information manager plus a standing portfolio reviewer in the enterprise
Risk
Concentrated — low frequency, long latency; the failure surfaces years after the team has gone

If this is you, the next step is

We pose a year-three query and see whether the pack can answer it.

Stress-test one handover pack against a real question

Where construction and infrastructure projects actually sit

The distribution across the ladder, and why the Profiled → Measured step loses more projects than any other.

Most projects are at Referenced or Profiled. The distribution is heavily weighted toward the two rungs that produce documents rather than measurements: a large majority of live schemes can point to a firm-level policy, a growing minority hold a project register, and only a small fraction have a number about an AI tool that was produced on their own data. Fewer still have handed anything to an asset owner.

Distribution of construction and infrastructure projects across the ladder

Referenced is the mode. The drop from Profiled to Measured is the largest single loss on the ladder, because it is the first rung that requires the project to produce evidence rather than description.

Share of projects

  • 36% — 1 · Referenced (the mode)
  • 30% — 2 · Profiled
  • 20% — 3 · Measured
  • 11% — 4 · Managed
  • 3% — 5 · Transferable

Source: Illustrative distribution, synthesised from NIST AI RMF adoption material and published UK construction assurance guidance

The shape of the loss is specific to project delivery. In a persistent organisation the Profiled → Measured step is a resourcing decision: someone is asked to build a test set and does. On a project it is a scheduling problem, because the window in which measurement is cheap — before reliance is priced into the programme — is also the window in which the team is busiest. NIST published the framework in January 2023 (opens in a new tab) with an accompanying roadmap (opens in a new tab) precisely because measurement science for AI was the acknowledged weak point, and construction inherits that weakness on top of its own.

It is worth being clear about what a distribution like this does and does not tell you. It is a synthesis, labelled illustrative above, and the honest comparison for any given project is not against the market but against schemes of similar contract form, client type and package mix. A design-and-build hospital under a two-stage contract has a different achievable rung from a five-year highways framework lot, because the framework lot has repeat mobilisations to learn from and the hospital does not.

The project profile: every function against the artefact that carries it

The page's centrepiece — the AI RMF Core crosswalked to construction project artefacts, with the owner on a project and the evidence each one leaves behind.

A project profile is the AI RMF Core answered in project artefacts rather than in new documents. That is the single most important design decision available to a project team: for almost every category in the framework, construction already runs an instrument that satisfies it — the project execution plan, the RACI, the risk register, the method statement, the DPIA, the deviation record, the early warning notice, the information-delivery plan. Creating a parallel AI governance apparatus is both more expensive and less durable, because the parallel apparatus has no natural attendance and no contractual home.

The table below is the crosswalk we use to write a profile. It is organised by AI RMF category rather than subcategory — the subcategory detail belongs in the Playbook (opens in a new tab), which is where you go once you know which categories bite for a given use. Read it as a checklist of artefacts, not of clauses: if the right-hand column is empty for a row that matters to your consequence class, that is the piece of work. Two rows — GOVERN 6 and MANAGE 3 — reach outside the project into supplier paper, and the contractual machinery behind them is a subject of its own, covered by the vendor-governance page in this cell.

Function · categoryWhat it asks of the projectProject artefact that carries itOwner on a projectEvidence it leaves behind
GOVERN 1 · Policies and proceduresThat documented policies and procedures for AI risk are in place and appliedThe firm's AI policy, referenced by name and version in the PEP, with project deviations recordedDigital or technical director; project director countersignsA PEP section naming the policy version in force at mobilisation
GOVERN 2 · Accountability structuresThat roles and lines of responsibility for AI risk are documented and staffedThe project RACI, extended with a named profile owner for each AI useProject directorA RACI row per AI use naming an individual, not a team
GOVERN 4 · Risk cultureThat teams are committed to identifying and communicating AI riskThe standing AI item on the weekly project risk meeting, and the induction slideProject managerRisk meeting minutes carrying AI items alongside temporary works
GOVERN 6 · Third-party riskThat policies address AI risks arising from supplier software and dataThe AI schedule in the subcontract or supply agreement, and the register's supplier columnCommercial managerSigned schedules plus a register row per supplier-brought tool
MAP 1 · Context establishedThat the context of use, and who is in it, is understood and recordedThe project AI register: each use, its package, its site, its data and its usersProfile ownerA register with one row per AI use, dated at mobilisation and updated at gates
MAP 2 · CategorisationThat the AI system is categorised by what it does and how it is usedThe intended-purpose statement — one paragraph per use, in project languagePackage or design managerA signed intended-purpose statement per AI use
MAP 3 · Capabilities and expected benefitsThat targeted usage, goals and expected benefits are understoodThe reliance decision: what the output may and may not be used forProject director with the profile ownerA dated reliance decision recorded against the register row
MAP 4 · Risks and benefits mappedThat risks are mapped for all components, including third-party software and dataConsequence classification against the project's existing risk matrixProject risk managerA consequence class per AI use, on the project's own scale
MAP 5 · Impacts characterisedThat impacts on individuals, groups and communities are characterisedThe DPIA, plus the workforce and neighbour consultation record where site monitoring is involvedData protection lead with the HSE managerDPIA reference and consultation minutes attached to the register row
MEASURE 1 · Methods and metricsThat appropriate methods and metrics are identified, documented and appliedThe measurement plan: metric, threshold, sample size, cadence, who reads itProfile ownerA versioned measurement plan per AI use
MEASURE 2 · Trustworthy characteristics evaluatedThat the system is evaluated for validity, reliability, safety, privacy and bias in contextThe project test set: held-out project records plus deliberately seeded defectsDesign or QA managerTest-run records with dates, sample sizes and results by category
MEASURE 3 · Risks tracked over timeThat mechanisms exist to track identified risks as conditions changeThe monthly AI line in the project performance reportProject controlsA trend by month, not a single snapshot
MEASURE 4 · Feedback on measurementThat feedback on the efficacy of measurement is gathered and assessedThe override log — what site actually did with each outputPackage managerOverride and acceptance rates by package and by month
MANAGE 1 · Risks prioritised and responded toThat AI risks are prioritised, responded to and managed on a documented basisAI rows in the project risk register, with owners, dates and review movementProject risk managerRisk register entries that move between reviews
MANAGE 2 · Benefits maximised, impacts minimisedThat strategies exist to sustain value while limiting negative impactThe named non-AI fallback retained in the method statement, and its exercise recordPackage managerA method statement naming the manual route, plus a dated exercise record
MANAGE 3 · Third-party risks managedThat risks from third-party systems and data are managed in serviceThe model-change notice route in the supply agreement, and the re-test triggerCommercial manager with the profile ownerChange notices received, and the re-test record each one triggered
MANAGE 4 · Treatments documented and communicatedThat treatments are documented, monitored and incidents communicatedThe deviation record and the early warning noticeProject managerDeviation records with dates, owners and closure
The project profile crosswalk: AI RMF Core categories mapped to the construction project artefact that carries each one, its owner inside a project organisation, and the evidence it leaves behind. Categories are grouped by function; the subcategory-level detail lives in NIST's Playbook.

Three triage rules make this table usable in a mobilisation week rather than a quarter. They are the difference between a profile a project actually maintains and a matrix that is completed once and never opened.

  • Triage by consequence, not by tool count

    A project with eleven AI uses does not need eleven full profiles. Two or three will carry safety, structural or statutory consequence and earn every row of the table; the rest carry commercial or programme consequence and earn a register row, an intended-purpose statement and a threshold. Applying uniform depth is the fastest way to guarantee the deep rows are done badly, because the effort was spent on a document summariser.

  • Name the existing artefact before you create a new one

    For every row, ask what the project already produces that could carry it. Almost always something does. The DPIA already exists for site monitoring; the method statement already names manual checks; the early warning notice already exists as a contractual mechanism under NEC. Creating an AI-specific parallel is how governance ends up unattended — and an artefact with a contractual home outlives one with only a policy home.

  • Write the owner as a person, at mobilisation

    Every row on the crosswalk has an owner column for a reason. 'The digital team' is not an owner on a project; a named design manager is. Assign at mobilisation while the RACI is being agreed, and re-assign explicitly at every demobilisation — the framework's accountability requirement is the one construction is best equipped to meet and most likely to let lapse quietly when someone rolls off.

One extension is worth knowing about. Where generative tools are in use — bid text, specification drafting, document summarisation, RFI responses — NIST publishes a separate Generative AI Profile (NIST AI 600-1) (opens in a new tab), which enumerates twelve risks specific to or exacerbated by generative systems and sets out suggested actions against the same four functions. Two of the twelve matter disproportionately on projects: confabulation, which NIST defines as confidently stated but erroneous content, and information integrity. On a project both attach to the same artefact — a document in the CDE that reads as authoritative because it is well-formatted. The practical response is a provenance field on generated documents and a named human check before issue, both of which sit naturally in the information-delivery plan.

The temporary-organisation problem, and how much evidence a use has to carry

Why a framework written for persistent organisations strains against project delivery — and the 2×2 that decides how deep the measurement has to go.

The AI RMF strains against construction because it assumes an organisation that persists, and construction delivers through organisations that dissolve. A joint venture is incorporated to build one thing and wound up afterwards. A framework lot runs for five years with a delivery team that turns over twice. A design-and-build contract hands a permanent asset from a temporary builder to a permanent owner and then removes every person who understood the decisions. None of that is a defect in the framework; it is a translation problem, and it produces four failure points that are specific to project delivery.

  • Ownership dissolves faster than the risk does

    A model's consequences persist in the asset for decades; the person who accepted its output has moved on within eighteen months. This is not new to construction — it is the reason design-check certificates and the CDM health and safety file exist. Apply the same instinct: every reliance decision names an individual and a date, and demobilisation triggers explicit re-assignment rather than silent lapse. The CDM 2015 regime (opens in a new tab) already establishes that duty holders and their information obligations are handed on rather than extinguished at completion.

  • Joint ventures have two policies and one project

    On a JV, each parent brings its own AI policy, approved tool catalogue and risk appetite, and they will not agree. Resolving that at parent level takes longer than the project has. The workable move is to settle it once, at mobilisation, in the JV's own PEP: which parent's catalogue governs, whose DPIA template applies, where the register lives, and who owns each row. Two hours of argument at mobilisation replaces a recurring one at every package let.

  • The client specifies AI the project cannot refuse

    Employer's requirements increasingly name platforms — a progress-monitoring system, a common data environment with built-in classification, an asset-tagging tool. The project cannot change them and still carries their consequences on site. Register them explicitly as client-specified, record the reliance decision anyway, and raise the residual risk through the contract's own early-warning mechanism rather than absorbing it silently.

  • Handover is a cliff, not a taper

    On the day of practical completion, the project's memory ends. Every reliance decision that was not written down, every test that was not recorded, every deviation that was handled verbally becomes unrecoverable. In the UK the building safety regime (opens in a new tab) has already pushed information durability up the agenda for higher-risk buildings, and the Building Safety Regulator (opens in a new tab) is explicit that information must transfer with the asset. An AI evidence pack is the cheapest way to make the project's reasoning transferable, and it is only cheap if it accretes rather than being assembled at the end.

The second half of the translation problem is depth. Once the register exists, the recurring question is not whether to test but how hard, and the answer is a function of two variables: what happens if the output is wrong, and where the number you are relying on came from. The matrix below is how we set that on a project.

How much test evidence does this AI output have to carry?

Plot each register row. The vertical axis is the consequence class you assigned during MAP; the horizontal is the provenance of the performance number you are actually relying on. Only one quadrant is genuinely dangerous, and one is merely expensive.

Unearned reliance

  • High consequence, someone else's number
  • The quadrant every serious incident starts in
  • Fix: build a project test set before the next package relies on it

Warranted reliance

  • High consequence, measured on your own records
  • Thresholds written, second check retained where recall is weak
  • Fix: nothing — re-test when the tool, the site or the standard changes

Proportionate

  • Low consequence, vendor benchmark
  • Register row, intended purpose, a threshold and periodic sampling
  • Fix: sample quarterly and move on — depth here buys nothing

Over-assured

  • Low consequence, fully measured
  • TEVV effort spent where it cannot pay back
  • Fix: move the effort to the top-left row that is still untested
Consequence if the output is wrong — top: Safety, structural or statutory, bottom: Programme or commercial only
Where the performance number came from — left: The vendor's published benchmark, right: Measured on this project's own data

It is the joint responsibility of all AI actors to determine whether AI technology is an appropriate or necessary tool for a given context or purpose, and how to use it responsibly.

That sentence is the reason the matrix has two axes rather than one. Appropriateness is contextual: the same clash-detection tool is warranted on a warehouse and unearned on a hospital riser, not because the model changed but because the consequence class did. It is also why a project cannot inherit its depth decision from a corporate policy — the policy does not know which package you are letting next month.

What the enterprise-and-project split looks like in public

Two publicly reported programmes — one contractor, one asset owner — read against the ladder. Neither is an Atomic Loops engagement; each links to the organisation's own published material.

The clearest public evidence for the enterprise-and-project split is in what large organisations have chosen to build and to require. A contractor that builds AI centrally still has to profile each arrival on each project; an asset owner that mandates digital delivery is writing the requirements that make a project's evidence pack exist. The two cases below sit on opposite sides of the same boundary, and both are read here against the ladder rather than presented as endorsements.

Two reference points read against the ladder

Outcomes as published by the organisations themselves; we have not independently audited them — verify figures against the linked source before reusing them. Both card images are generated industry scenes from this page's image library, not photographs of the named organisations.

Generated scene: a governance meeting in a boardroom overlooking a city construction skyline, with data panels on screenSkanskaGlobal contractor and project developer · Nordics, Europe, US13
Challenge
Delivering AI capability to a business that executes through hundreds of separate project organisations, where a tool approved centrally still arrives on each project with a different intended purpose, a different consequence class and a different client's requirements attached.
Approach
Skanska USA Building established a Digital Transformation and Solutions Team uniting its data, emerging technology and AI capabilities, and built internal tools — the Sidekick suite and Skanska Metriks cost modelling — on the firm's own accumulated project record rather than licensing general-purpose products.
Reported outcome
Skanska reports that the Sidekick suite scaled from an initial 2024 pilot to more than 1,000 employee users supporting work across 500-plus projects, with the tools intended to surface safety and operational risks earlier and reduce administrative burden.
What it shows about the curveBuilding centrally is a strong GOVERN position and does not discharge MAP or MEASURE. A tool used across 500 projects has 500 intended-purpose statements to write and 500 consequence classes to assign, because what it is for and what breaks if it is wrong differ by contract, package and client — which is exactly why the profile is per project and not per tool.

Skanska — Digital Transformation and Solutions Team press release (opens in a new tab)

Generated scene: an elevated highway viaduct under construction at dusk with a project team reviewing digital model overlaysNational HighwaysStrategic road network operator · England · asset owner and client24
Challenge
Operating a permanent network that is built and modified by a rotating population of temporary supplier project organisations, each of which demobilises and takes its reasoning with it while the operator retains the asset for decades.
Approach
National Highways publishes its Digital Roads programme, setting out how digital design, construction and operation — including connected data and digital twins of the network — are to be delivered by its supply chain, so requirements are stated by the permanent organisation rather than negotiated per scheme.
Reported outcome
The published programme positions digital delivery and data as network-level requirements on suppliers across design, construction and operation, making information transfer an explicit expectation of the client rather than a project-team preference.
What it shows about the curveThe rung above Managed is bought by the client, not by the contractor. Where the asset owner names the information it will receive, the evidence pack becomes a deliverable with a date and a recipient; where it does not, even a well-run project hands over nothing, because nobody asked and nobody would have read it.

National Highways — Digital Roads (opens in a new tab)

Read together, the two cases describe the boundary this page is about. The contractor case shows that enterprise capability, however good, stops at the project gate: the framework's context-bound functions have to be executed by whoever is standing on the site. The asset-owner case shows the other side: the durability of a project's AI record is largely determined by what the permanent organisation asked for before the project began. A project team that wants to reach Transferable and has a silent client will have to write the requirement itself, into the information-delivery plan, and get it agreed.

The evidence stack, layer by layer

What has to exist at each rung — five layers, annotated with the rung that first requires them, and the one layer everybody defers.

A transferable project profile requires five layers, and the order in which they are built determines whether the profile compounds or has to be reconstructed. The stack below is deliberately unglamorous: nothing in it is a product, every layer is defined by what it must guarantee, and four of the five are made of artefacts a project already produces. Only one layer — the transfer layer — has no natural home in existing practice, which is precisely why it is the one that goes missing.

The project AI evidence stack, rung-annotated

Each layer is annotated with the rung that first requires it. A project trying to reach Measured without the register layer is testing tools it has not defined; a project trying to reach Transferable without the response layer is handing over a description of untested confidence.

  1. Enterprise floor

    Stage 1+

    • AI policy and risk appetiteBoard-owned, versioned, referenced by the PEP
    • Approved tool catalogueWhat may be used, in what setting, with what data
    • Profile template and register schemaThe fields every project must fill in, identically
  2. Project register layer

    Stage 2+

    • AI register rowOne per use, including embedded and client-specified tools
    • Intended-purpose statementOne paragraph, project language, signed by the package owner
    • Consequence classAssigned on the project's existing risk matrix, not a new scale
  3. Measurement layer

    Stage 3+

    • Project test setHeld-out NCR, RFI and inspection records plus seeded defects
    • ThresholdsWritten per use, agreed by whoever relies on the output
    • Test-run recordsDated, sampled, scored by category — the TEVV trail
  4. Response layer

    Stage 4+

    • Reliance decisionWhat the output may and may not be used for, dated and signed
    • Named fallbackThe manual method retained in the method statement, and exercised
    • Deviation and early-warning routeThe project's own machinery, carrying AI items unchanged
  5. Transfer layer

    Stage 5+

    • Evidence packAccreted continuously; a transmittal at completion, not an assembly
    • Information-delivery requirementThe pack named as a deliverable in the BEP or its equivalent
    • Portfolio returnThe completed profile read at portfolio review and reused at mobilisation

Pipeline described

  1. Enterprise floor (stage 1+) — AI policy and risk appetite: Board-owned, versioned, referenced by the PEP; Approved tool catalogue: What may be used, in what setting, with what data; Profile template and register schema: The fields every project must fill in, identically
  2. Project register layer (stage 2+) — AI register row: One per use, including embedded and client-specified tools; Intended-purpose statement: One paragraph, project language, signed by the package owner; Consequence class: Assigned on the project's existing risk matrix, not a new scale
  3. Measurement layer (stage 3+) — Project test set: Held-out NCR, RFI and inspection records plus seeded defects; Thresholds: Written per use, agreed by whoever relies on the output; Test-run records: Dated, sampled, scored by category — the TEVV trail
  4. Response layer (stage 4+) — Reliance decision: What the output may and may not be used for, dated and signed; Named fallback: The manual method retained in the method statement, and exercised; Deviation and early-warning route: The project's own machinery, carrying AI items unchanged
  5. Transfer layer (stage 5+) — Evidence pack: Accreted continuously; a transmittal at completion, not an assembly; Information-delivery requirement: The pack named as a deliverable in the BEP or its equivalent; Portfolio return: The completed profile read at portfolio review and reused at mobilisation
Step-by-step insights
Enterprise floor — a constraint set, not a compliance claim
The enterprise layer's job is to remove decisions from projects, not to make claims on their behalf. A good approved tool catalogue answers three questions before a project asks them: may this tool be used at all, with what categories of data, and in what settings. A good profile template makes register rows comparable across schemes, which is the only reason portfolio review is possible later. What the layer must not do is assert that projects are compliant. The moment the enterprise treats its policy as the evidence, projects stop producing any, and the framework's context-bound functions quietly go unexecuted across the whole portfolio.
Register layer — the cheapest artefact and the highest leverage
Everything above this layer is impossible without it, and it is a week of work at mobilisation. The two fields that carry the weight are intended purpose and consequence class, because together they set the depth of everything downstream. Write intended purpose in the project's own language and in the negative as well as the positive — 'flags probable clashes for engineer review; does not approve designs; is not evidence of compliance with the specification'. Negative clauses are what stop scope creep, which on projects is not a governance abstraction but the ordinary process of a useful tool being asked to do slightly more each month.
Measurement layer — a records exercise, not a research one
The obstacle here is imagined difficulty. Teams assume a test set requires data science, and it requires an afternoon in the NCR log. Pull known-bad cases from your own records, hold them out, have two engineers seed additional defects into clean records and swap sets so neither scores their own, then run the tool and score by category rather than in aggregate. What you learn is not a headline accuracy number but a shape: which categories the tool is weak on. That shape is what changes practice, because it lets you keep a manual second check where it matters and drop it where it does not — which is where the time saving actually comes from.
Response layer — reuse the machinery, always
The response layer is where projects are tempted to invent, and where invention is least survivable. Construction already runs a deviation process, an early-warning mechanism, a risk register with owners and dates, and a method-statement discipline that names manual checks. Every MANAGE requirement maps onto one of these. The only genuinely new component is the reliance decision, and even that has an obvious precedent in the design-check certificate. Reuse buys attendance, contractual standing and audit trail for free; a parallel AI forum buys none of them and starves within a quarter.
Transfer layer — the layer with no natural owner, which is why it is missing
Every other layer has someone whose job it obviously is. The transfer layer does not: the project team is demobilising, the asset owner has not asked, and the enterprise is looking at the next bid. It therefore has to be created by contract rather than by intention — named in the information-delivery plan with a date and a recipient, and accreted continuously so completion is a transmittal. Projects that schedule the pack as a completion task produce a folder that satisfies the deliverable and answers no future question, because by then the people who knew why a reliance decision was made have gone.

The layer most often skipped is the response layer's fallback, and it is the one that determines whether anyone will accept the AI at all. A package manager will adopt a tool they can stop using on a Tuesday morning without stopping work; they will resist one whose removal has no plan. Naming the manual method in the method statement costs a paragraph. Exercising it once costs half a day and is the only way to discover how many downstream artefacts have already absorbed the output — which, on every project where we have run the exercise, is more than anyone predicted.

A 90-day plan: profile the design-assurance checker on one package

The Profiled → Measured transition made concrete on a single, common construction problem — the automated standards checker running over design submissions before issue for construction.

Moving one rung takes about 90 days when it is scoped to one AI use on one package, and multiple years when it is scoped to a firm. To make that concrete, the plan below runs the transition on a specific problem most design-and-build projects now have: an automated checker that reviews design submissions against the project specification and the applicable standards before they are issued for construction. The tool already exists and is already being used — the quarter contains no procurement and no model development, only profiling, measurement and the decisions that follow.

Profiled to Measured on one design package, in one quarter

One package, one tool, one named owner. If a phase needs more than its window, narrow the scope — one discipline rather than four, one submission type rather than all — rather than extending the plan.

  1. Days 1–15

    Register the use and write its intended purpose

    Add the checker to the project AI register with a named owner — normally the design manager for the package. Write the intended-purpose statement in one paragraph, including what the output is not: it flags probable non-conformances for engineer review; it does not approve a submission and it is not evidence of compliance. Assign a consequence class on the project's existing risk matrix, and record where the tool's outputs currently reach: which reviewers, which CDE workflow, which approval status.

    A register row, a signed intended-purpose statement, a consequence class

  2. Days 16–40

    Build the test set from your own NCR and RFI log

    Pull 80 historic cases from the project's own non-conformance reports, RFIs and design-review comments where the issue would have been detectable in the submission. Have two engineers seed 40 further defects into clean submissions and swap sets so neither scores their own. Hold the whole set out of the tool's normal workflow. Score by defect category, not in aggregate — the categories are what change practice.

    A 120-case held-out set with a recall figure per category

  3. Days 41–65

    Set the threshold, the reliance decision and the fallback

    Agree a written threshold with the package manager who relies on the output. Record the reliance decision: which categories the checker may screen unaided, which retain a manual second check, and which it must not be used for at all. Name the manual route in the method statement and resource it. Where the tool is a supplier's, confirm the model-change notice route and the re-test trigger with the commercial manager.

    A dated reliance decision and a resourced manual fallback

  4. Days 66–90

    Rehearse the deviation route and start the pack

    Run one deliberate exercise: declare the checker unavailable on a live submission and execute the manual route, then trace which downstream artefacts had already absorbed its output. Raise one real deviation record end to end. Open the evidence pack in the CDE and file everything produced so far — register row, statement, test records, reliance decision, exercise note, deviation — and name the pack in the information-delivery plan.

    A rehearsed response and an evidence pack that has started accreting

The order matters

  1. Intended purpose before metrics

    Deciding what to measure requires knowing what the output is for. Teams that start with metrics measure the tool's headline accuracy, which averages across categories with wildly different consequences and tells the package manager nothing they can act on. Write the purpose first, in the negative as well as the positive, and the measurement plan writes itself.

  2. Project data before vendor benchmarks

    A supplier's benchmark was computed on their corpus, against their definition of a correct answer. It is fine for shortlisting and worthless for reliance, because the person carrying the risk cannot interrogate it. Two engineers and an afternoon in your own NCR log produce a number that can be argued about in a design meeting, which is the only kind that changes a sampling rate.

  3. One package before the project

    Profiling every AI use on the project simultaneously guarantees all of them are done shallowly. Take the highest-consequence row, run it to the end, and use what you learn — how long the test set took, which categories were weak, what the fallback exercise found — to size the rest. The register already told you which row to pick.

Two practical notes. First, the test set is a project asset, not a one-off: keep it, and re-run it when the supplier ships a model change, when the design standard is revised, or when the package moves to a materially different building type. Second, if the checker touches anything that identifies individuals — and some document tools do, through metadata and comment authorship — the DPIA belongs in the same fortnight, not later; the ICO's AI guidance (opens in a new tab) sets out what a data-protection assessment has to cover when AI is in the processing path.

Verifying the profile: measures, thresholds and the readiness checklist

What to read from the project's own records to prove the rung, and the eight-item check that separates a described profile from an evidenced one.

A profile is verified from records the project already keeps, not from a self-assessment. Every measure below reduces to something countable in the register, the CDE, the risk register or the test log — which matters because the whole point of the exercise is producing evidence somebody else can check. The table is the read sheet: what to count, where it lives, how often to read it, and the rung at which each measure first means anything.

MeasureHow to read itWhere it livesCadenceHonest from
Register coverageAI uses on the register ÷ AI uses found in a package-by-package sweepRegister vs a walk of the packagesAt each stage gateRung 2
Embedded-arrival shareRegister rows the project did not procure ÷ all rowsRegister, supplier columnQuarterlyRung 2
Intended-purpose coverageRows with a signed statement ÷ all rowsRegister, CDEMonthlyRung 2
Measured-reliance shareHigh-consequence rows with a project test record ÷ high-consequence rowsRegister vs test logMonthlyRung 3
Recall by defect categoryDetected seeded and historic defects ÷ total, per categoryTest-run recordsPer test run, and on supplier changeRung 3
Override rateOutputs overridden or ignored ÷ outputs produced, by packageTool logs and reviewer recordsMonthlyRung 3
Deviation closure timeDays from AI deviation raised to closed, medianDeviation recordsMonthlyRung 4
Fallback exercise ageMonths since the manual route was last exercised, per useMethod statement and exercise notesQuarterlyRung 4
Re-test lag on supplier changeDays from model-change notice to completed re-testChange notices vs test logPer noticeRung 4
Pack completeness at gateRequired pack artefacts present ÷ required, at each gateEvidence pack in the CDEAt each stage gateRung 5
Reading the project profile from project records. 'Honest from' is the rung at which the measure first describes something real; below that rung it will read as zero or as noise.

Two of these deserve emphasis because they are counter-intuitive. A very low override rate is usually bad news rather than good: it normally means nobody is checking, not that nothing is wrong, and a tool that is never overridden has stopped generating the reviewer feedback that MEASURE 4 exists to capture. And fallback exercise age is the measure most likely to be dismissed as bureaucratic — right up to the first supplier outage, at which point it is the only thing that determined whether the package kept moving.

Project profile readiness checklist

Eight items. If you cannot tick all eight for one live project, that project is not at Managed regardless of how good the tools are. Tick as you go — this list works without JavaScript.

0 of 8 ticked

0 of 8 — Referenced, and now it is explicit

Nothing ticked is the honest starting position for most live projects, and it is not a judgement on the team — it is what happens when a framework is adopted at head office and the project is told it is covered. Do not start with the framework. Start with a two-week sweep of one project's packages and write the register. Everything else on this list becomes possible once rows exist.

Failure modes that unwind a project profile

Profiles decay, and they decay for reasons specific to temporary organisations. Five patterns account for most of it.

Profiles decay, and on a project they decay faster than in a persistent organisation, because the conditions that sustained them — the people, the contract, the phase — all change on a schedule. Five patterns account for most of the regression we see, and four of the five are invisible from the documents alone: the register still reads as current, the test record still exists, and the thing that changed was outside the paper.

Likelihood: highImpact: high

The profile owner rolls off and the reliance stays

A design manager who wrote four reliance decisions moves to the next scheme in March. The decisions remain in force, the tools keep running, and nobody can now explain why a particular category retained a manual second check. On a project this is not an edge case — it is the normal rhythm of resourcing, and it usually happens between stage gates when nobody is reviewing anything.

PreventionProfile ownership transfer goes on the demobilisation checklist next to system access and vehicle returns, with an explicit named successor per register row.

Likelihood: highImpact: medium

The supplier ships a model change and nobody re-tests

A cloud-delivered tool is updated silently. The interface is identical, the outputs shift, and the test record on file was produced against a model that no longer exists. The project continues to rely on a number that has quietly become historical, and the drift is only detected when a miss reaches site.

PreventionA model-change notice obligation in the supply agreement, with a standing re-test trigger and a named recipient on the project — not the enterprise.

Likelihood: mediumImpact: medium

The project moves phase and the test set stops representing it

A test set built from substructure records is still in force when the project reaches façade and fit-out. The measured recall was real and is now about a different distribution of work, different standards and different failure categories. Nothing on paper changed, so nothing triggered a review.

PreventionTie re-test to stage gates rather than to the calendar, and record the phase the test set was drawn from on the test record itself.

Likelihood: mediumImpact: high

Scope creeps past the intended-purpose statement

A tool profiled for one purpose is gradually asked to do more, because it is useful and available. The checker used to screen submissions starts being cited in approval comments; the summariser used for internal notes starts producing client-facing text. Each step is small and reasonable, and the cumulative position is outside everything that was tested.

PreventionWrite intended purpose in the negative as well as the positive, and re-read the statements at each stage gate against what the tool is actually being used for.

Likelihood: mediumImpact: high

The pack is scheduled as a completion task

Everything needed for the evidence pack was produced during delivery and never assembled. In the final quarter, a coordinator collects whatever can still be found: the register, some test screenshots, no reliance decisions and no deviation history. The deliverable is satisfied and the asset owner receives something that cannot answer a year-three question.

PreventionFile into the pack on the day each artefact is produced, and check pack completeness at every stage gate rather than at completion.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Project AI profile
The set of artefacts that answer the AI RMF for one project: the register, the intended-purpose statements, the consequence classes, the measurement plan, the test records, the reliance decisions and the deviation history. It is per project, not per firm or per tool.
AI RMF Core
The structured body of the NIST AI Risk Management Framework: four functions (GOVERN, MAP, MEASURE, MANAGE) containing 19 categories and 72 subcategories. GOVERN is cross-cutting; the other three are context-bound and therefore project-bound in construction.
TEVV
Test, evaluation, verification and validation — NIST's collective term for the evidence behind a performance claim. On a project it means a dated test run against held-out project records, scored by defect category, not a vendor benchmark.
Intended-purpose statement
A one-paragraph project-language statement of what an AI output is for and, critically, what it may not be used for. It sets the consequence class and therefore the depth of everything downstream, and it is the control that stops scope creep.
Consequence class
The severity assigned to an AI use if its output is wrong, expressed on the project's own existing risk matrix rather than an invented AI-specific scale. It determines how much test evidence the use has to carry before anyone relies on it.
Reliance decision
The dated, signed record of what a measured AI output may and may not be used for on this project, and what residual control remains. The construction equivalent of a design-check certificate, and the artefact a reviewer is most likely to ask for.
Project test set
A held-out collection of known-bad cases drawn from the project's own NCR log, RFI log, snagging list and design-review tracker, plus defects seeded deliberately by engineers who do not score their own. Scored by category, not in aggregate.
Seeded defect
A fault introduced deliberately into an otherwise clean record so a tool's recall can be measured without waiting for real failures. Two engineers seed and swap sets so neither scores their own work, which removes the obvious bias.
AI evidence pack
The assembled project profile handed to the asset owner at completion: register, statements, test records, reliance decisions and deviation history. Accretes during delivery so that handover is a transmittal rather than an archaeology exercise.
Temporary organisation
The project, joint venture or framework delivery team that exists only to build one thing and is dissolved afterwards. The unit that actually executes MAP, MEASURE and MANAGE, and the reason the AI RMF needs translating rather than adopting verbatim.
Profile owner
The named individual accountable for one register row while the project exists — usually a design or package manager rather than anyone in the digital team. Ownership transfer is a demobilisation task, not an assumption.
Mobilisation inheritance
What a new project starts from: either a blank register template, or the previous scheme's returned profile with its tools, consequence classes, test sets and known weak categories already populated. The mechanism by which GOVERN becomes real.

Frequently asked questions

The questions project directors, design managers and digital leads ask most often when applying the NIST framework to a live scheme.

What is the NIST AI Risk Management Framework, and does it apply outside the US?

It is a voluntary framework published by the US National Institute of Standards and Technology in January 2023 (NIST AI 100-1) that organises AI risk management into four functions — GOVERN, MAP, MEASURE and MANAGE — containing 19 categories and 72 subcategories. It is deliberately non-sector-specific and use-case agnostic, carries no legal force anywhere including the US, and is used internationally because it is the most practical structured articulation of what AI risk management involves. UK and EU contractors use it as a working structure while their legal obligations come from elsewhere.

Can a construction project be NIST AI RMF certified?

No. The framework is voluntary and non-certifiable: there is no NIST audit, no accredited certification body and no certificate. Any supplier claiming NIST AI RMF certification is describing a self-assessment, and it is reasonable to ask what evidence sits behind it. Certification against a management-system standard is a different instrument — ISO/IEC 42001 is the one designed for that, with accredited bodies and audit cycles. On a project the useful output of the NIST framework is not a badge but a record: a register, measured performance, dated reliance decisions and a deviation history.

Who owns the project AI profile on a joint venture?

One named individual in the JV, appointed at mobilisation and written into the project RACI — normally the design or digital manager rather than anyone at either parent. The harder question is whose corporate policy governs, because each parent brings its own AI policy, approved tool catalogue and risk appetite. Settle that once in the JV's project execution plan: which catalogue applies, whose DPIA template is used, where the register lives and who signs reliance decisions. Two hours at mobilisation prevents a recurring argument at every package let.

How does the AI RMF map onto ISO 19650 information management?

They are complementary and meet at the information-delivery plan. ISO 19650 governs how project information is produced, exchanged and handed over, including the common data environment and the delivery-plan structure; the AI RMF governs how the risk of AI-produced information is managed. In practice the AI evidence pack should be named as an information deliverable in the BIM execution plan, and AI-generated or AI-assisted documents should carry a provenance field in the CDE. Neither standard requires the other, but running the framework without an information home is how the record dies at handover.

How much test evidence does MEASURE actually require on a project?

As much as the consequence class earns, and no more. Plot each register row on two axes: what happens if the output is wrong, and where the performance number you rely on came from. High-consequence rows relying on a vendor benchmark need a project test set — typically 80 historic cases from your NCR and RFI logs plus 40 seeded defects, scored by category. Low-consequence rows relying on a vendor benchmark need a register row, a threshold and periodic sampling. Applying uniform depth guarantees the rows that matter are done badly.

How is this different from having an AI governance charter?

A charter is the firm's rulebook; a profile is one project's record. The charter states what is permitted, who approves what and how the business audits itself, and it is written once and maintained centrally. The profile answers the framework for a specific scheme — these eleven tools, these intended purposes, this measured recall, these dated decisions — and is produced by a team that will disband. Both are needed and they fail differently: a firm with a charter and no project profiles has policy without evidence, which is the most common position in the sector.

What does the NIST Generative AI Profile add for construction teams?

It extends the same four functions to generative systems, enumerating twelve risks that are novel to or exacerbated by them, with suggested actions against each. Two matter disproportionately on projects: confabulation — confidently stated but erroneous content — and information integrity, because both attach to the same artefact, a well-formatted document in the CDE that reads as authoritative. Where generative tools touch bid text, specification drafting, RFI responses or document summaries, the practical response is a provenance field on generated documents and a named human check before issue.

What happens to the AI evidence at practical completion?

It should transfer with the asset information, and on most projects today it does not. The asset owner inherits a building or a network whose design, inspection and commissioning records were partly AI-assisted, and has no access to the project's reasoning once the team demobilises. The fix is contractual rather than cultural: name the AI evidence pack in the information-delivery plan with a date and a recipient, and file into it continuously so completion is a transmittal. Packs assembled in the final quarter satisfy the deliverable and answer no future question.

How does the NIST framework relate to the EU AI Act and ISO/IEC 42001?

They do different jobs and are not substitutes. The EU AI Act is binding law with classification-driven obligations; ISO/IEC 42001 is a certifiable management-system standard; the NIST AI RMF is a voluntary risk-management structure with no legal force and no certificate. Their practical relationship on a project is that the NIST profile generates most of the evidence the other two ask for — inventory, intended purpose, testing records, human oversight arrangements, incident handling. NIST publishes crosswalks to other regimes for exactly this reason. Do the profile once; report from it into whichever regime applies.

How long does it take to write a project AI profile?

About two weeks to reach Profiled on a live scheme, and about a quarter to reach Measured on one high-consequence use. The register itself is a two-week sweep: interview package leads, catch the embedded and client-specified tools nobody thinks of as AI, and write one paragraph of intended purpose per row. Measurement is the part with real cost, roughly three weeks for the first test set because the records work is unfamiliar. Doing it at mobilisation costs a fraction of doing it at month six, when finding the tools means interviewing everyone.

Does the profile cover AI our subcontractors bring onto the project?

Yes — every AI use that affects the project's decisions belongs on the project register, regardless of who bought it. At profile level this means three things: record the arrival, assign a consequence class, and make an explicit reliance decision about what the output may be used for on your packages. The contractual machinery behind that — screening at tender, the AI schedule in the subcontract, acceptance trials, flow-down depth and offboarding at completion — is a substantial subject in its own right and is covered in the vendor-governance page in this cell.

What is the smallest useful version of this on a project already halfway through?

One register, one measured row, one reliance decision. Do not attempt a retrospective full profile on a scheme in delivery; it competes with production and will not finish. Spend a fortnight on a package-by-package sweep to build the register, pick the single highest-consequence row, test it against records the project already holds, and write one dated reliance decision. That takes about six weeks, produces something a client or insurer could actually be shown, and tells you honestly how much the remaining rows would cost.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for construction, infrastructure and industrial operators — design assurance, progress and quality analytics, forecasting and decision support running against live project data. We build the measurement, the fallback route and the evidence trail alongside the model, because on a project the record is what outlives the team that made it.

  • · AI risk profiles written against the NIST AI RMF for live project delivery
  • · Test sets built from operators' own NCR, RFI and inspection records
  • · Evidence packs contracted into information-delivery plans and handed over at completion
  • · 20 cited sources on this page

Sources

  1. NISTAI Risk Management Framework (opens in a new tab)
  2. NISTAI Risk Management Framework 1.0 (NIST AI 100-1) (opens in a new tab)
  3. NISTGenerative AI Profile (NIST AI 600-1) (opens in a new tab)
  4. NIST Trustworthy and Responsible AI Resource CenterAI RMF Playbook (opens in a new tab)
  5. NIST Trustworthy and Responsible AI Resource CenterAI RMF crosswalks to other frameworks and regulations (opens in a new tab)
  6. NISTAI RMF roadmap (opens in a new tab)
  7. NISTNIST risk management framework aims to improve trustworthiness of AI (opens in a new tab)
  8. NISTCybersecurity Framework (opens in a new tab)
  9. ISOISO/IEC 42001 — AI management systems (opens in a new tab)
  10. ISOISO 19650-1 — organisation and digitisation of information about buildings (opens in a new tab)
  11. UK BIM FrameworkUK BIM Framework (opens in a new tab)
  12. Health and Safety ExecutiveConstruction health and safety (opens in a new tab)
  13. Health and Safety ExecutiveConstruction (Design and Management) Regulations 2015 (opens in a new tab)
  14. Health and Safety ExecutiveBuilding safety (opens in a new tab)
  15. Health and Safety ExecutiveConstruction statistics in Great Britain, 2025 (opens in a new tab)
  16. GOV.UKBuilding Safety Regulator (opens in a new tab)
  17. Information Commissioner's OfficeGuidance on AI and data protection (opens in a new tab)
  18. European CommissionRegulatory framework for AI (opens in a new tab)
  19. SkanskaSkanska USA Building establishes Digital Transformation and Solutions Team (opens in a new tab)
  20. National HighwaysDigital Roads (opens in a new tab)

Find out what one live project could actually evidence today

We walk one scheme against the framework with your project director, design manager and information manager, mark which artefacts exist and which are assumed, and leave you with the register, the intended-purpose statements and a costed 90-day plan for the highest-consequence AI use on it. You keep the artefacts whether or not we build with you.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.