How would we know if the AI's output is still good six months after launch?
Keep a fixed set of real examples with agreed correct answers, and re-run it whenever anything changes - the model, the instructions, or the data feeding it. Without that, quality change is invisible. Nobody announces that output got worse; people just quietly stop relying on it. The test set is a permanent asset, and the model is a replaceable part.
Drift arrives from three directions
Output quality changes without anyone touching the thing that changed.
- The model moved. Providers update models. Behavior on your edge cases can shift even when general capability improves.
- Someone edited the instructions. A tweak to fix one complaint frequently degrades three cases nobody was looking at.
- Your business moved. New service lines, new terminology, a new pricing structure, a new market. The system is answering correctly about a company that no longer exists.
Building the frozen set
Take a few dozen to a couple hundred real records, with identifying details handled appropriately, and have humans agree on the correct output for each. Weight it deliberately: include the ambiguous cases, the ones the system got wrong in the past, the unusual job types, and the records with missing fields. A test set of easy examples always passes and tells you nothing.
Freeze it. The value comes from comparability over time, so resist the urge to keep adding examples in a way that makes this month incomparable to last month. Add a new version when you need one, and keep running the old one alongside it for a while.
Run it on change, not just on a calendar
The trigger should be any change to the model, the instructions, the retrieval logic or the upstream data mapping - plus a periodic run to catch silent provider-side shifts. Store the results with a date and the configuration that produced them, so that when someone asks whether output is worse than it was in March, the answer is a comparison rather than an opinion.
This is ordinary software regression testing applied to a probabilistic component. Nothing about it is novel, and skipping it is the most common reason teams cannot say whether a system is degrading. It belongs in the same operational layer as your platform monitoring.
The live signal that costs nothing
Alongside the frozen set, watch what users do with output in production: how often they override it, how often they correct a specific field, which categories generate the most edits. That correction rate is a continuous quality measurement generated for free by people doing their jobs.
A rising correction rate in one category is usually the earliest warning of drift, and it points at the specific area to investigate. Pair it with the frozen set and you can distinguish "the model changed" from "our business changed" - two problems with completely different fixes. This is the same reasoning that governs how we build analysis people are supposed to trust.
Topics: evaluation · drift · quality · monitoring
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.