Skip to article
AI and trust

An AI finding without a source is just a rumour

Confidently wrong output is a documented, named failure mode of language models. On a construction project the fix is not a better model — it is refusing to publish what cannot be checked.

The gate does not ask whether a finding sounds right. It asks whether a reviewer could check it without taking the model's word for anything.

Ask a language model why an activity is at risk and you will get an answer. It will be specific, it will be plausible, it will name a drawing, and there is no way to tell from reading it whether the drawing exists.

That is not a prompt problem. It is a property of the technology, it has a formal name, and the standards bodies have already written it down.

#The failure has a name, and it is not “hallucination”

NIST's Generative AI Profile for the AI Risk Management Framework calls it confabulation: the phenomenon in which systems "generate and confidently present erroneous or false content in response to prompts", noting that these are "colloquially also referred to as 'hallucinations' or 'fabrications'".1

The word choice matters more than it looks. Hallucination suggests a malfunction — something that happens when the system goes wrong. Confabulation is what the system does when it is working exactly as designed and has nothing to say. It fills the gap, fluently, at the same confidence as everything else. There is no tell.

NIST files the risk against the "Explainable and Interpretable" trustworthy-AI characteristic,1 which is the useful part: the mitigation is structural, not linguistic. You do not fix confabulation by asking the model to be careful.

#The number that should end the argument

The first systematic study of this on document-grounded professional work looked at law. Asked specific, verifiable questions about real federal court cases, the models tested — GPT-4 among them — "hallucinate at least 58% of the time, struggle to predict their own hallucinations, and often uncritically accept users' incorrect legal assumptions."2

≥58%
hallucination rate on specific, verifiable questions about real federal cases
Poor
ability of the models to predict their own hallucinations
Yes
tendency to accept a user's incorrect premise rather than correct it

That study is about law, not construction, and we are not going to pretend it measured a schedule. But the task shape is the same one project controls asks for: retrieve a specific fact from a large corpus of dense professional documents and state it precisely. And the second and third findings are the ones that should worry a project team more than the first. A tool with a 58% error rate that knew when it was guessing would be usable. A tool that cannot tell, and that will agree with a wrong premise a superintendent brings to it, is worse than the error rate suggests.

Construction already knows what to do with an assertion nobody can trace. It rejects it.

This is not a new standard the industry has to invent for AI. Forensic schedule analysis has a standards-body recommended practice with a five-layer method taxonomy specifically so that two analysts can say which method they used and be checked against it.3 Schedule quality has named, quantitative checks with published thresholds. The discipline's entire posture toward expert opinion is show your work. Software does not get an exemption for being clever.

#So the gate goes before the output, not after

The design consequence is simple and slightly unfashionable: every candidate finding has to carry a reference into the project record, that reference has to resolve, and a finding whose reference does not resolve does not get published. It is dropped, and the drop is logged against the run so it is auditable rather than invisible.

A finding is therefore one of three things and never a fourth: anchored to a source document, computed from schedule math, or explicitly labelled as an assumption. "The model thinks so" is not on the list.

#What it costs

Dropping findings is not free. Some of what gets dropped is true. A real supply-chain risk the model correctly inferred from something it could not point at gets thrown away, and nobody ever learns it was right.

We take that trade knowingly, because the two errors are not symmetrical. A dropped true finding costs you one insight. A published unverifiable finding costs you the reviewer's trust in every other finding on the page — and once a scheduler has been burned by a confident citation to a drawing that does not exist, they are right never to trust the tool again. The expensive failure is not the miss. It is becoming the thing people stop reading.

The regulators are converging on the same place. ISO/IEC 42001:2023, the first AI management system standard, lists traceability, transparency and reliability among its core benefits.4 The EU AI Act requires high-risk systems to be "designed and developed in such a way as to ensure that their operation is sufficiently transparent to enable deployers to interpret a system's output and use it appropriately."5 Whatever a construction platform is, "interpret the output and use it appropriately" is unachievable if the output cannot be traced.

How the same principle plays out on a specific problem is in the weather baseline post — where the entire argument is about which records a third party authored — and on the construction delay risk page.

Frequently asked

Why do AI tools invent drawing numbers and specification sections?

Because generating fluent, plausible text is what the underlying models do, and they do it at the same confidence whether or not the source exists. NIST's AI Risk Management Framework names this confabulation and describes it as confidently presented erroneous or false content, which is why the mitigation has to be a verification step outside the model rather than a better-worded prompt.

Does citing a source mean the AI finding is correct?

No, and it is important not to oversell it. A resolvable reference proves the cited document exists and is relevant; the reasoning about it can still be wrong. What it changes is who carries the verification burden: a reviewer can open the source and check in seconds instead of having to establish whether the citation is real at all.

What happens to findings that cannot be sourced?

In SP4N they are dropped before publication and the drop is recorded against the analysis run, so the filtering is auditable rather than hidden. That means some true findings are discarded. We consider that the cheaper of the two available mistakes.

Sources

  1. 1AI Risk Management Framework: Generative AI Profile (NIST AI 600-1)NIST · 2024
  2. 2Large Legal Fictions: Profiling Legal Hallucinations in Large Language ModelsJournal of Legal Analysis (Oxford University Press) · 2024A legal-domain study. Cited here for the task shape, not as a construction measurement.
  3. 3Recommended Practice 29R-03, Forensic Schedule AnalysisAACE International · 2011
  4. 4ISO/IEC 42001:2023 — Artificial intelligence management systemISO · 2023
  5. 5EU AI Act, Article 13 — Transparency and Provision of Information to DeployersRegulation (EU) 2024/1689 · 2024
Share this
LinkedInPostEmail