Every software company you deal with has an AI agent now. Your contract administration platform has one. Your document management system has one. Your accounting software has one.
Now think about your commercial team. Count how many of them used an agent this week to do a piece of work start to finish. Not asked ChatGPT a question. Not had Copilot tidy an email. Actually handed over a task and got a finished piece of work back.
I suspect the answer is close to zero, and I think that's worth taking seriously rather than blaming the profession being slow.
First, what an agent actually is
The word has been stretched to the point of meaninglessness, so here's the version that matters on a job.
A chat tool answers. You ask it something, it gives you words back, and you do something with those words.
An agent does. You give it a task, it goes and looks at things, uses tools, takes several steps, and comes back with the work done. The difference isn't intelligence. It's access and permission.
Ask a chat tool what should go in a compensation event quotation, and you get a decent general answer that you already knew. Give an agent the same question with access to your records, the accepted programme and the contract, and it can go and assemble the quotation. Same underlying model. Completely different outcome.
That distinction is the whole point. Most QSs have only ever used the first kind, and have quite reasonably concluded that AI is a better search engine.
Why nobody in your office is using one
Four reasons. Three of them are good ones.
1. The review problem
If an agent drafts you a compensation event notice, and the consequence of getting it wrong is a time bar or a rejected quotation, you're going to check every line of it against the contract. You now have a drafting job replaced by a checking job. Checking someone else's work is often slower than doing it yourself, and it's certainly less enjoyable.
For a lot of QS tasks, the bottleneck was never the typing. It was the judgement, and the judgement hasn't moved.
Where this objection breaks down is when the agent is working from your own project record rather than generating from nothing. Reviewing a summary of what your site team actually recorded is fast, because you're checking against a source you trust and can click straight into. Reviewing a plausible paragraph invented by a general-purpose model is slow, because you have to verify every claim from scratch.
That's the dividing line. Agents grounded in your records are worth the review. Agents grounded in nothing are not.
2. The access problem
An agent with no reach into your systems is a chatbot with better manners.
This is where most construction AI dies. The site records are in one system, the correspondence is in Outlook, the programme is in a file someone owns, the valuations are in a spreadsheet on a network drive, and IT has quite sensibly not given a third-party AI tool permission to touch any of it.
You can't agent your way around a permissions problem. And the honest position is that IT are usually right to be cautious, because the failure mode of an agent with write access to your commercial records is considerably worse than the failure mode of a chatbot that says something silly.
The practical answer is narrow, read-only, logged access to one system at a time, rather than waiting for someone to hand an AI the keys to everything.
3. The definition problem
Very few QSs have been shown what an agent looks like in their own work. They've been shown demos of agents booking flights and writing code.
Nobody has sat down with a mid-career QS and said: here's your Thursday afternoon, here's the twenty minutes of it that a machine can do reliably, here's exactly how to set that up. Until that happens, AI agent stays an abstraction, and abstractions don't get adopted.
4. The incentive problem
The fourth reason is the uncomfortable one, and I'll pose it rather than answer it.
If you're working under a target cost or cost-reimbursable arrangement and your commercial team gets faster, who captures that value? If you're a consultant billing hours, and the work takes half as long, what happens to the fee?
I don't think this is the main blocker in contracting. But I've heard enough versions of why would we make ourselves more efficient on this one to think it's not zero either, and pretending otherwise isn't honest.
Where agents actually earn their keep on a job
Not the interesting work. The opposite.
The tasks where agents work today share a shape: repetitive, evidence-based, high volume, and painful to do properly under time pressure. That's a very good description of a lot of commercial administration, and of why so many QS teams are drowning in it.
Finding events in the record. Going back through weeks of site records to identify what might constitute a compensation event, an early warning, or a delay to a planned operation. A human does this badly under deadline pressure, not because they lack skill, but because there's too much to read. It is also where the money leaks.
Assembling the evidence bundle. Pulling together the diary entries, photographs, labour returns and correspondence relating to a single event, in date order, with the gaps flagged. This is hours of clicking and no judgement at all.
Checking the record against the contract clock. Which early warnings haven't been followed up. Which quotations are approaching a reply deadline. Which events were notified but never quantified. Miss that clock and you are into 8-week time bar territory.
Drafting from the record. Once the evidence is assembled, a first draft narrative that cites what actually happened rather than what the model imagines usually happens. That is the difference between a quotation that survives PM scrutiny and one that comes straight back.
Notice what's not on that list: deciding whether something is a compensation event, agreeing a valuation, negotiating with the other side. That's your job, and it will remain your job, and any vendor telling you otherwise should be shown the door.
What to watch out for
Anything an agent produces about contractual entitlement needs a named human to check it and a record that they did. The RICS professional standard on responsible use of AI becomes mandatory on 9 March 2026, and its requirements around human oversight, competence and documentation apply to exactly this kind of output. If you can't say who reviewed it, you have a problem regardless of whether the output was correct. We covered what the new RICS AI rules mean for commercial teams separately.
Watch for confident contract references that don't exist. General purpose models are notably bad at clause numbers and will produce plausible ones. Every clause reference gets checked against the actual contract, every time.
Dip-sample the outputs even when they look right. The failure mode of a good agent isn't obvious nonsense; it's a subtly incomplete evidence bundle that nobody notices until the adjudication.
And be realistic about the setup cost. If configuring the agent takes longer than the task it saves, and you only do that task twice a year, do it by hand.
Try this today
Pick one recurring commercial task that takes you more than an hour, happens at least weekly, and involves reading rather than deciding. Chasing outstanding early warnings is a good candidate.
Write down, in plain English, the exact steps you take. Where you look, what you look for, what you do with what you find, what the finished output looks like. Ten minutes, on paper.
That document is the specification for your first agent. It's also, on its own, the most useful thing most commercial teams could produce this month, because half the time writing it down reveals that the process isn't actually agreed between the people doing it.
Use an enterprise-grade tool with a data processing agreement in place, such as Claude Team or your organisation's licensed Copilot deployment. Don't paste project records into a free consumer account.
At Gather we've been building in exactly this direction, on the basis that an agent is only as good as the record it can reach. Agents that can search the site record, find events and assemble the evidence behind them are considerably more useful than agents that write confident prose about NEC4 in the abstract.
The wider point stands whatever tools you use. The agent moment in construction won't arrive as a single product everyone starts talking about. It will arrive as a hundred small jobs that quietly stop being done by hand, and most people will only notice afterwards.
Key Takeaways
- A chat tool answers; an agent does. The difference is not intelligence, it is access and permission.
- Agents grounded in your own project records are worth reviewing. Agents grounded in nothing are not.
- Four real barriers: the review problem, the access problem, the definition problem and the incentive problem.
- Agents earn their keep on repetitive, evidence-based, high-volume tasks, not on judgement.
- Deciding entitlement, agreeing valuations and negotiating stay with the QS.
- Write down the steps of one recurring task. That document is the specification for your first agent.
Gather turns your site diaries into commercial evidence and flags compensation events before the eight week time bar closes. Book 15 minutes and see it run against one of your own projects.
Book a 15 minute demo



.webp)




