What happens when a rep says the AI scored their call wrong?
You need a defined path: the rep flags the item, a human listens to the cited moment, and the record is either corrected or explained. Both outcomes are useful. If the rep is right, fix the score and check whether the rubric wording caused it. If the rep is wrong, the quoted evidence makes the coaching specific. Track dispute rates by criterion — clusters point at a bad rubric line, not a bad rep.
A dispute path is a quality instrument, not a complaint box
Most teams treat disputes as friction to minimize. That is backwards. A rep who listens to their own call and argues about one criterion is doing free, high-quality error detection on the exact cases your sampling would most likely miss.
A dispute rate of zero is not a sign that everything is right. It usually means people have concluded that flagging changes nothing, which is the point at which you stop hearing about problems without having fewer of them.
Make it easy: a flag on the specific criterion, not an email to a manager. The lower the friction, the better your signal.
Three causes, three different fixes
- Transcription error. The model scored text that does not match the audio — background noise, crosstalk, a name misheard. Fix belongs upstream, in capture quality.
- Rubric ambiguity. The criterion is genuinely readable two ways. Fix belongs in your wording, and the same dispute is probably about to arrive from other reps.
- Borderline judgment. The call really is on the line. Fix is a documented tie-breaking convention so the next one is decided the same way.
Correct the record without erasing it
Never overwrite the original result. Store the AI output, the human override, who made it, when, and why. The history is what makes the system auditable later, and it is what lets you measure whether overrides cluster around one manager or one criterion.
This is the same discipline that makes any operational AI system defensible: the record shows what the machine said and what the human decided, separately.
Read the dispute rate per criterion
Aggregate dispute volume tells you almost nothing. Dispute rate broken out by criterion tells you where the standard is unclear. One criterion generating a large share of flags is a wording problem you can fix in an afternoon.
Watch the reverse signal too. A criterion that never gets disputed and never fails is probably not measuring anything — it is a line in the rubric that always passes. Pruning those makes the scorecard shorter and more honest.
Topics: disputes · rubrics · QA · fairness
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.