Why does AI make up numbers, and how do you stop it in reporting?
A language model produces plausible text, and a plausible-looking number is easy to produce. The fix is not a better prompt. It is not letting the model produce the number at all: query the figure from the system of record, pass it in, and restrict the model to explaining it. If a number in a report cannot be traced back to a specific query, it should not be in the report.
Generation and arithmetic are different jobs
Asking a model to both compute and narrate is asking one component to do two things it is unevenly good at. Language models are strong at expressing a relationship in clear prose and unreliable at deriving a figure from raw records, particularly across joins, date boundaries and deduplication rules.
So split the job. A query engine computes; the model writes. This is the architectural choice that separates useful AI reporting from confidently wrong reporting.
Compute first, narrate second
In practice this means the pipeline resolves a defined set of metrics against your data, hands the model a small structured payload of those metrics, and asks for interpretation only. The model never sees the raw tables and is never asked what a total was.
It also means the metric definitions live in one place, versioned, rather than being re-derived by whatever prompt happens to run. Two reports disagreeing because they defined "lead" differently is a far more common problem than a model inventing a figure.
A useful side effect: once metrics are computed separately, the same definitions serve the narrative, the chart and the alert. When a written summary disagrees with the dashboard next to it, the cause is almost always two independent computations rather than anything the model did.
Three checks that catch the rest
- Every figure carries its source. Each number in the output maps to a named metric and the query that produced it. Nothing unsourced is allowed through.
- Automated numeric consistency. Extract every number from the generated text and confirm each one appears in the input payload. This is a simple, mechanical check and it catches invented figures and transposed digits reliably.
- An explicit no-data path. When a metric is missing or the period is incomplete, the correct output is a statement that the data is unavailable. If the system has no way to say that, it will fill the gap.
Where it still goes wrong
Grounding fixes fabricated numbers. It does not fix fabricated causation. A model handed accurate figures will still happily assert that a spend increase caused a revenue increase when the two merely moved together, and that sentence reads exactly as authoritative as the correct ones.
The defense is to keep causal language out of the generated layer, or to require that any claimed relationship be backed by an analysis step that ran. Our briefs are built around that constraint, and revenue reporting stays honest for the same reason.
Topics: grounding · hallucination · reporting · data integrity
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.