AI in construction

AI in Construction: A Practical Guide for UK Contractors

AI in construction is moving from experiment to operations. This guide cuts through the noise to explain what AI is actually doing on UK projects right now, which roles benefit most, how to evaluate tools without getting burned, and where the technology genuinely falls short.

01

What AI in Construction Actually Means (and What It Doesn't)

Let's be blunt. Most of what gets called "AI in construction" right now is either basic automation dressed up in marketing language or genuinely useful tools that are narrower in scope than their vendors admit.

Actual AI in construction falls into three categories:

1. Pattern recognition on historical data. Predictive analytics that looks at past project data to forecast cost overruns, schedule slippage, or safety incidents. Balfour Beatty used this approach on civil and rail projects and reported a 20% reduction in material waste with budget accuracy hitting 94%. These tools work where you have large, consistent data sets. They struggle on novel project types.

2. Natural language processing (NLP) on documents. AI that reads contracts, site diaries, correspondence, and reports to extract structured information. This is where the genuinely new value is for commercial teams. An NLP model can read 500 site diary entries and identify which ones contain compensation event triggers under NEC4 clause 60.1. A human doing the same job takes weeks and misses roughly 40% of legitimate events.

3. Generative AI for drafting and analysis. Large language models used to draft correspondence, summarise reports, or generate first-cut analyses. Useful for productivity. Genuinely risky if you let the model draft contractual notices without a human expert reviewing them.

What AI is not doing on UK construction projects right now: autonomously managing contracts, replacing QSs, or making procurement decisions. Anyone telling you otherwise is selling something.

02

The Commercial Case: Where the Money Is

The construction industry has a chronic underrecovery problem. On a typical £50M NEC4 Option C package, commercial teams identify roughly 60% of the compensation events they're entitled to notify. The other 40% slips through because of poor records, time pressure, staff turnover, or the simple reality that no human can read 18 months of site diaries with the forensic attention required.

40%

of legitimate compensation events slip through. On a project with 3% variations, that is £600,000 in unrecovered revenue.

That 40% isn't a rounding error. On a project with 3% variations, it represents £600,000 in unrecovered revenue. On a programme like HS2 or East-West Rail, it's multiples of that across dozens of packages.

AI addresses this at the data layer. The problem isn't that commercial teams are incompetent. The problem is that the data they need is scattered across:

  • Handwritten site diaries in paper files or scanned PDFs
  • Daily allocation sheets capturing labour and plant
  • Foreman reports with informal language ("boss said stop" = potential instruction under clause 27.1)
  • Email chains referencing verbal instructions
  • Weather records and access constraints

AI can read all of this, cross-reference it against the contract and programme baseline, and surface the entries that warrant commercial action. That's not replacing QS judgement. That's making sure QS judgement is applied to the right data.

I've seen projects where the AI review uncovered £340,000 in legitimate compensation events that the manual review had filed as "no action." Every single one was defensible. The records existed. They just hadn't been read by someone who knew what to look for.

Worked example: compensation event recovery on a civils package£42M · NEC4 Option C · East Midlands

On a £42M NEC4 Option C highways package in the East Midlands, the commercial team had been running monthly manual diary reviews. By the end of construction in February 2025, the compensation event register showed 34 agreed events totalling £1.1M. Standard for a project of that size.

When the QS ran an AI review of all site diaries and foreman reports from the preceding 18 months, the model flagged 19 additional events worth reviewing. Eleven turned out to be legitimate entitlements that had been either missed or filed as non-recoverable. The eleven events broke down as follows:

  • Six were physical conditions claims under clause 60.1(12): unexpected services, contaminated materials, and one genuinely novel ground interface that matched the contractor's risk threshold under the Scope
  • Three were Client instruction events under clause 60.1(1): verbal instructions at site meetings that appeared in the foreman's report but had never been elevated to the CE register
  • Two were weather-related disruption events under clause 60.1(13) that the team had recorded but not notified because they assumed the weather compensation event threshold hadn't been reached

Total value of the eleven events: £287,000 after the Project Manager's assessment. Eight-week notification windows were still open on seven of them at the time of the AI review because the team had caught it during the final three months of the project. The other four were assessed through the compensation event mechanism on the basis that the Contractor had given early warning of the conditions even if the formal CE notice was late.

Additional events flagged for review19
Legitimate entitlements confirmed11
Notification windows still open7
Recovered after PM assessment£287,000

Lesson: the records were good. The gap was the systematic review that AI provided.

03

How AI Works with Construction Data

Understanding the mechanism matters. If you're evaluating an AI tool and the vendor can't explain this clearly, walk away.

The three-layer model

Layer 1: Data ingestion. AI needs structured input. Site diaries, allocation sheets, correspondence, and programme data all need to be in a format the model can process. This is where most implementation projects get stuck. The data exists. It's not clean or consistent.

Layer 2: Contract context. The model needs to know the contract terms to make useful judgements. For NEC4 projects, this means loading the contract data, the compensation event register, the Accepted Programme, and the key clauses that govern time and cost. A general AI tool without this context will produce generic output. A specialist tool trained on NEC4 produces actionable analysis.

Layer 3: Output and action. The AI surfaces findings. A human expert decides what to do with them. The best tools present findings with the underlying evidence, a confidence score, and a suggested action. The worst tools give you a score out of 10 with no explanation.

What good AI output looks like

AI finding · site diary, 14 March 2025CE14 · cl. 60.1(12)

"Site diary entry 14 March 2025. Foreman's report references a two-hour concrete pour stoppage due to unexpected underground services. This aligns with compensation event category CE14 under clause 60.1(12) (physical conditions). The eight-week notification window under clause 61.3 closes 9 May 2025. Recommended action: raise early warning and prepare CE notification."

That's actionable. Compare it with "potential compensation event identified in Week 11." The second is useless.

Gather QS AI Agent reviewing site records and surfacing compensation events
Gather's QS AI Agent surfacing compensation event triggers from site records.
04

AI by Role: Who Gets the Most Value

AI doesn't deliver the same value to everyone on a construction project. Here's an honest summary of who benefits most, and why. Each section links to a dedicated guide.

04.1Quantity Surveyors

QSs get the clearest commercial benefit. The core job of a QS on an NEC4 contract, among other things, involves reviewing records, identifying compensation events, preparing and agreeing quotations, and managing the change register. AI tools that work with site records and contract data directly compress the most time-intensive part of that job.

The specific win: instead of manually reviewing 6 months of site diaries to prepare a compensation event assessment, the AI surfaces the relevant entries, cross-references them against the contract baseline, and drafts a first-cut narrative. The QS reviews, edits, and submits. What took 12 hours takes 2.

There are risks. A QS who relies on AI output without understanding the underlying contract mechanism will miss things the AI misses too. AI amplifies QS capability. It doesn't replace QS judgement.

Read the full guide: AI for quantity surveyors

04.2Commercial Managers

Commercial managers are primarily concerned with the overall commercial position: cost vs. budget, revenue vs. entitlement, risk exposure, and cash flow. AI helps them see the complete picture faster.

The specific win is in reporting latency. On a traditional project, the commercial manager sees the position weekly or monthly, compiled by the team. AI tools that read live data from site diaries, cost systems, and the CE register can surface the current position daily. Trends that would otherwise only appear in a month-end report become visible in time to act.

Read the full guide: AI for commercial managers

04.3Site Engineers

Site engineers don't use AI for commercial analysis. Their benefit is in records quality and site operations. AI tools that assist with site diary creation, prompt engineers to capture the details that matter commercially, and flag when records are incomplete or inconsistent make a real difference.

The problem I keep seeing: site engineers write perfectly adequate records from an operations perspective that are commercially worthless. "Delayed 2 hours. Weather." Great for operations. Useless for a compensation event claim. AI can prompt the engineer to add the causal chain, the resources affected, and the contract reference. That record is now commercially defensible.

Read the full guide: AI for site engineers

04.4Project Managers

Project managers need AI primarily for programme management and risk. AI tools that track the Accepted Programme against actual progress, flag emerging delays with a causal analysis, and model the programme implications of compensation events save significant time in the weekly reporting cycle.

On NEC4 contracts specifically, the programme is contractually critical. Failure to maintain and update the Accepted Programme has commercial consequences. AI that monitors programme health and surfaces issues before they become disputes is genuinely valuable.

Read the full guide: AI for project managers

04.5Project Directors

Project directors are managing portfolios of risk, not individual packages. AI at this level is about aggregated intelligence: which packages are most exposed commercially, where are the systemic risks, which teams are performing well and why.

A project director running three concurrent packages across a framework contract doesn't have time to read every weekly commercial report in detail. AI that summarises the position across all packages, flags the outliers, and predicts where the next dispute is likely to emerge is extremely valuable. The risk is that aggregated AI output hides detail that matters. Good project directors use AI for triage, not for decision-making.

Read the full guide: AI for project directors
05

NEC4 Contract Management and AI: A Natural Fit

NEC4 is the dominant contract form for major UK infrastructure projects. It's used on HS2, Network Rail, National Highways, and most water company frameworks. If you're working on projects over £10M in the UK, there's a good chance you're working under NEC4.

NEC4 creates specific obligations around records, notifications, and timelines that make it an ideal environment for AI assistance. A few examples:

Clause 61.3: The eight-week time bar. The Contractor must notify a compensation event within eight weeks of becoming aware of it. Miss that window and the entitlement is gone, regardless of how legitimate the event was. AI that monitors site records in real time and flags potential CE triggers before the window closes is a genuine risk mitigation tool.

Compensation events under clause 60.1. There are 19 categories of compensation event in NEC4, ranging from physical conditions (60.1(12)) to Client instructions (60.1(1)). Each has a slightly different evidence requirement. AI trained specifically on NEC4 clause 60.1 can match site diary entries to the relevant categories with reasonable accuracy, reducing the chance of events being missed or miscategorised.

The Accepted Programme. NEC4 places more emphasis on the programme than almost any other contract form. The Accepted Programme is the baseline for both delay analysis and compensation event assessment. AI tools that track programme compliance and model the impacted programme for each CE reduce the workload considerably.

Disallowed Cost under clause 11.2(26). Poor records are one of the fastest routes to Disallowed Cost on NEC4 Option C and D contracts. If you can't prove resource was deployed as claimed, the Project Manager can disallow it. AI tools that cross-reference site diary records against cost submissions reduce this exposure.

For a deeper dive, see our NEC4 guide.

06

Site Diaries: The Data Problem AI Solves

The site diary is where the commercial story of a project gets written, one entry at a time, usually by someone who has no idea that's what they're doing.

A foreman writing "access delayed 3 hours, waiting for Client's rep to clear the area" is describing a compensation event. They don't know that. The QS reviewing that diary six months later might spot it, might not. By the time it reaches final account, the window is closed.

This is the core data problem that AI addresses. The site diary contains the evidence. The evidence is unstructured, informal, and spread across hundreds of entries. AI reads it systematically, every entry, every day, looking for the patterns that indicate commercial events.

Gather Record capturing a structured site diary entry
Structured site diary capture in Gather Record.

Good site diary practice and AI work together. An engineer who writes detailed, consistent records gives the AI better data to work with. The AI, in turn, can prompt engineers to capture the specific information that makes records commercially useful: the cause of a delay, the resources affected, the duration, and any verbal instructions received.

The combination reduces the risk of the most common commercial failure on UK construction projects: legitimate events that existed in the records but were never acted on.

07

Evaluating AI Tools: A Practical Framework

The AI tools market in construction is noisy. Every platform with a dashboard claims to use AI. Here's how to cut through it.

AI vendor evaluation scorecard

Use this table when evaluating any AI tool for commercial construction use. Score each criterion 1-3. A total below 10 is a reject. Above 12 is worth trialling.

Evaluation CriterionScore 1Score 2Score 3
Training data specificityGeneric documents / no disclosureConstruction documentsNEC4 contracts + UK site records specifically
Output traceabilityScore only, no evidenceSummary with document referencesSpecific entry citation with clause mapping
Confidence handlingNo confidence flagsLow-confidence items flaggedEvidence-weighted confidence with override workflow
Integration approachRequires new data entryFile upload / API availableWorks with your existing data format as-is
Implementation timeline6+ months2-6 monthsUnder 8 weeks for pilot package
Vendor construction expertiseGeneric SaaS teamSome construction advisersBuilt by or with practising QSs / contract experts

Scoring guide: Total 15-18: strong candidate. Total 10-14: trial with conditions. Total below 10: not ready for commercial construction work.

Five questions to ask any AI vendor

1. What is the model actually trained on?
A general LLM fine-tuned on generic documents is not the same as a model trained specifically on NEC4 contracts, UK construction site diaries, and compensation event case law. The training data determines the output quality. Ask to see examples.

2. What does the output look like, specifically?
Ask for a demo with real data, not a polished presentation. What does a flagged compensation event output look like? What evidence does it cite? Can you trace the AI's reasoning? If the answer is "it gives you a risk score," that's not enough.

3. How does it handle data it hasn't seen before?
AI models fail when they encounter something outside their training distribution. On a complex project, that happens regularly. What does the tool do? Does it flag low confidence? Does it hallucinate a plausible-sounding answer?

4. What are the integration requirements?
Every AI tool needs data. Where does that data come from? If the answer involves a six-month integration project and a dedicated data engineer, factor that into the evaluation. The best tools work with the data you already have in the format you already use.

5. Who is responsible when the AI is wrong?
The answer should be "the human who acted on it." AI output is an input to professional judgement. Any vendor suggesting their tool removes the need for professional judgement is either mistaken or selling to someone who doesn't understand the liability.

The specialist vs. generalist decision

There are two types of AI tools in construction: horizontal platforms that do many things adequately, and specialist tools that do one thing extremely well.

For general productivity, document management, and reporting, horizontal platforms often make sense. Microsoft Copilot integrated with your existing Office stack is genuinely useful for drafting and summarising.

For commercial work on NEC4 contracts, specialist tools win. The difference in output quality between a general LLM and a model trained specifically on NEC4 clause 60.1 and UK site diary conventions is significant. The general model will identify some compensation events. The specialist model will identify substantially more, with better evidence citations and fewer false positives.

08

AI Construction Limitations: What Doesn't Work Yet

Honest assessment. This is where most AI guides in construction stop. They should start here.

Complex commercial judgement. Deciding whether a compensation event quotation should be challenged, whether an early warning creates a legal obligation, or how to handle a disputed disallowed cost assessment requires NEC4 expertise and commercial experience. AI can inform the decision. It can't make it.

Novel physical conditions. AI models trained on existing site data learn from what has happened. Genuinely novel ground conditions, unusual site interfaces, or first-of-type engineering challenges are exactly the areas where AI is least reliable.

Relationship dynamics. The relationship between the Project Manager and the Contractor's commercial team shapes how NEC4 contracts run in practice. AI doesn't understand that a particular Project Manager routinely issues late assessments and that this creates a pattern of compensation events. An experienced QS does.

Final account negotiation. The final account is a human process. It involves judgement, relationship, commercial pressure, and often a degree of pragmatic settlement. AI tools that read 24 months of records and produce a final account schedule are useful inputs. They're not a substitute for the negotiation.

Adversarial environments. If the other party is actively suppressing evidence or constructing a misleading paper trail, AI working from available data will be misled by it. Forensic analysis in disputed situations requires human expertise.

Be sceptical of any AI tool that doesn't acknowledge these limitations. The technology is genuinely useful within its domain. Overselling it creates the expectation failures that damage adoption.

09

Implementation: How Tier 1 Contractors Are Actually Adopting AI

Most Tier 1 contractors are at the early majority stage. They've done the pilots. Some are now rolling out to specific contract types or business units. Here's what's working.

What's working

Narrow, high-value use cases first. The contractors getting results have picked one specific problem, usually compensation event identification or programme monitoring, and built the AI workflow around that. They haven't tried to boil the ocean.

Starting with data hygiene. Before deploying AI, they've standardised their site diary formats, allocation sheet templates, and cost coding. AI applied to clean, consistent data produces dramatically better output than AI applied to five different diary formats from five different sub-contractors.

QS-led implementation. The projects that have gone well have been led by commercial teams, not IT. The QSs define what good output looks like, what the false positive rate is acceptable at, and when to override the AI recommendation. IT enables the data flow. They don't run the implementation.

Treating AI as a junior QS. The most useful mental model I've encountered: treat AI output the way you'd treat work submitted by a capable but inexperienced QS. Review it. Check the reasoning. Don't submit it to the Project Manager without applying professional judgement.

What's not working

Trying to automate everything at once. Teams that deploy AI across 15 functions simultaneously end up with mediocre output across all 15 and no champion for any of them. Focus wins.

Skipping the integration work. AI tools that sit outside your existing data flows require people to enter data twice. People don't do that consistently. Garbage in, garbage out.

Deploying AI without training the commercial team. AI output is only useful if the people receiving it understand what it means and what to do with it. A compensation event flag means nothing to a foreman who doesn't know what a compensation event is.

For earned value management specifically, AI integration follows the same pattern: start with the data layer, establish clean baselines, then add AI analysis on top.

10

Common Mistakes in AI Adoption

These are the mistakes I see repeatedly across UK contractors adopting AI for commercial work.

1. Confusing AI with automation. Automation removes human steps from a process. AI assists human judgement. Conflating the two leads to inappropriate trust in AI output for decisions that require professional expertise.

2. Buying the platform before defining the use case. I've seen commercial teams spend six figures on an AI platform and then spend a year trying to work out what problem they're solving. Define the use case first. Evaluate tools against it.

3. Ignoring data quality. AI output quality is bounded by input quality. A £200M programme with inconsistent site diary formats, patchy allocation sheets, and no standardised cost coding will produce poor AI output regardless of how good the model is.

4. Not involving the end users in implementation. QSs who weren't consulted in the selection process won't use the tool. Commercial managers who don't trust the output will ignore it. AI adoption is a change management problem as much as a technology problem.

5. Expecting AI to find everything. Even good AI tools miss legitimate compensation events. The goal is to increase the proportion you identify from 60% to 85% or 90%, not to achieve 100%. Setting unrealistic expectations creates the conditions for abandonment after the first missed event.

6. Neglecting the eight-week time bar. The most commercially damaging mistake on NEC4 contracts is missing notification windows. AI tools help here, but only if they're deployed before the window closes. Implementing AI for the final account review six months after a project is in dispute is too late.

11

Reference Table: AI Use Cases by Role

RolePrimary AI Use CaseSecondary Use CaseKey Risk MitigatedRelevant Resource
Quantity SurveyorCompensation event identification from site recordsCE quotation drafting assistanceEight-week time bar (clause 61.3)AI for QSs
Commercial ManagerReal-time commercial position reportingDisallowed Cost risk monitoringRevenue underrecoveryAI for commercial managers
Site EngineerSite diary quality promptingAllocation sheet completionRecords gap at final accountAI for site engineers
Project ManagerProgramme compliance monitoringRisk register maintenanceDelay to Accepted ProgrammeAI for project managers
Project DirectorPortfolio commercial dashboardsSystemic risk identificationLate visibility of commercial exposureAI for project directors

AI readiness by contract type

Contract FormAI ReadinessKey AI ApplicationNotes
NEC4 Option AHighCE identification, programme monitoringFixed price: CE management critical
NEC4 Option CVery highCE identification + cost verificationTarget cost: Disallowed Cost risk significant
NEC4 Option DVery highAs Option CRemeasurable target: programme tracking critical
NEC4 Option EModerateCost verification, early warning monitoringCost-reimbursable: less CE pressure
JCT StandardModerateVariation identification, extension of timeDifferent change mechanism; still benefits from records AI
FIDIC Red BookModerateClause 20 notification monitoringDifferent dispute mechanism; AI less specialised

Frequently Asked Questions

What is AI in construction?

AI in construction refers to the application of machine learning, natural language processing, and predictive analytics to construction management tasks. In practice, this includes analysing site records to identify missed compensation events, monitoring programme compliance, predicting cost overruns from historical data, and assisting with commercial document drafting. UK applications are particularly focused on NEC4 contract management, where clause obligations create specific, high-value use cases for AI-assisted record review.

How is AI being used in UK construction?

UK contractors, particularly Tier 1 firms like Balfour Beatty, Costain, Murphy, and Amey, are using AI primarily in three areas: commercial management (compensation event identification, cost verification), programme management (schedule variance detection, early warning monitoring), and site operations (safety monitoring, records quality). The most commercially significant application is NEC4 contract management, where AI tools applied to site diaries and cost records recover revenue that manual processes miss.

Can AI replace quantity surveyors?

No. AI amplifies QS capability rather than replacing it. The tasks AI does well are data-intensive but relatively mechanical: reading large volumes of records, cross-referencing them against contract terms, and flagging items that need professional attention. The tasks that require QS expertise, commercial judgement, clause interpretation, negotiation, and dispute management, remain firmly human. The more realistic question is whether QSs who use AI effectively will outperform those who don't. The answer to that is yes.

How accurate is AI at identifying compensation events?

This depends heavily on the quality of the underlying site records and the specificity of the AI model. Tools trained specifically on NEC4 contracts and UK site diary conventions perform considerably better than general AI tools applied to the same data. In practice, a well-implemented specialist tool identifies 80-90% of legitimate compensation events in site records, compared with roughly 60% for manual review. The remaining 10-20% typically involves events with very limited contemporaneous documentation or complex factual circumstances.

What is the eight-week time bar under NEC4?

Under clause 61.3 of NEC4, the Contractor must notify a compensation event within eight weeks of becoming aware that the event has occurred. Failure to notify within this period bars the entitlement, regardless of how legitimate the event was. This is one of the highest-value AI use cases in construction: real-time monitoring of site records to flag potential CE triggers before the notification window expires.

What data does AI need to work effectively in construction?

The minimum data set for useful AI output on an NEC4 project includes: site diaries (daily), daily allocation sheets (labour, plant, materials), the contract documents (including the scope, conditions, and programme), the Accepted Programme baseline, and the existing compensation event register. The more consistent and structured this data is, the better the AI output. Projects that have standardised diary formats, allocation sheet templates, and cost coding before deploying AI get significantly better results.

How long does it take to implement AI on a construction project?

For a specialist NEC4 commercial tool, a focused implementation on a single package typically takes four to eight weeks from data onboarding to live output. The time is driven primarily by data preparation, not technology. If your site diaries are in a consistent digital format and your cost coding is standardised, implementation is fast. If you're working from scanned handwritten diaries in five different formats, plan for longer. Enterprise-wide roll-out across a programme or framework adds integration and change management time.

Is AI suitable for smaller construction projects?

The commercial case for AI is strongest on larger, more complex projects because the volume of records and the value of missed events justifies the implementation. On projects under £5M, the value recovery may not justify the cost of specialist AI tooling. That said, the market is moving quickly. Tools that are being deployed on £50M+ infrastructure packages today will be accessible to £5M projects within a few years.

AI-powered QS

See What Your Site Records Are Missing

Gather's QS AI Agent analyses every site diary entry for potential compensation events under NEC4 clause 60.1. Most teams identify 60% of their entitlement manually. Gather finds the rest. £287,000 recovered on a single project from missed CEs.