Skip to main content

How do you keep bad data out of an AI-generated report?

APIs & Data Published August 29, 2026
Short Answer

With validation at ingest and a rule that unclear records are excluded and reported rather than guessed. Practical gates: required fields present, values within expected ranges, foreign keys resolvable, no duplicate source keys, totals reconciling to the source. Failures go to a quarantine table a human reviews. AI written over unvalidated data produces confident, fluent sentences about wrong numbers.

The gates worth having

Data quality is not a philosophy, it is a list of checks that run on every load.

  • Required fields. A job without a completion date or a business unit cannot be reported on correctly. Catch it at the door.
  • Range checks. An invoice total that is negative, or thousands of times the median, is either a credit memo you have not modeled or a typo.
  • Referential checks. Every job points to a customer that exists, every invoice to a job. Orphans indicate a sync gap.
  • Uniqueness. One row per source identifier. A violation means a pagination or retry defect.
  • Totals reconciliation. Your sum for a closed window matches the source's own number, within a stated tolerance.

Quarantine beats coercion

The tempting shortcut is to fix bad values in flight. Null revenue becomes zero. Missing source becomes "other." An unmatched lead gets assigned to the largest campaign. Each of these turns a visible problem into an invisible one, and the report still renders, which is exactly why nobody catches it.

Routing the record to a quarantine table instead keeps the problem countable. Someone can look at fifty quarantined jobs and see that a technician has been filing under a retired business unit since April, which is a fixable process issue rather than a permanent distortion in the numbers.

The errors AI cannot detect

A language model reading your data has no way to know that a tech logged a membership call as a service call, or that one location has been entering leads a day late all quarter. The values are plausible, well-formed and wrong. The model will build a coherent narrative around them.

This is the difference between generative output and operational AI. The safeguard is not a better prompt; it is validation upstream and a system that can say what it does not know. Any AI-generated analysis worth trusting should be able to state its coverage and its exclusions alongside its conclusions.

Measure quality, do not just assert it

Track a few numbers over time and publish them where the people using the reports can see them: quarantine rate, unmatched rate on joins, null rate on key fields, and reconciliation variance. They should be small and stable.

When one of them moves, something upstream changed, and you have a specific place to look before the effect reaches a decision. Data quality that is measured tends to improve; data quality that is merely claimed tends not to.

Topics: data quality · validation · quarantine · AI reporting

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001