What evidence should an AI call score include so a rep can actually dispute it?
The criterion exactly as written, the score given, and the specific moment that produced it: a timestamp and the quoted line, or an explicit statement that the expected behavior never occurred anywhere in the call. If a rep cannot move from the score to the seconds of audio behind it, the score is an assertion. Assertions do not survive their first serious challenge, and they should not.
A score is a claim, and claims need support
The moment a number affects how someone is treated at work, it has to be reviewable. Not because reps are litigious, but because unexplained numbers destroy the coaching conversation. A rep who cannot see why they scored what they scored has nothing to work on except the number itself.
Explainability is also the fastest quality control you have. Most rubric defects surface when someone reads the evidence and says the quoted line does not mean what the criterion assumed.
What a defensible score record contains
- The criterion text and its version. Rubrics change. A score means nothing without knowing which wording produced it.
- The value and the anchor it matched. Not just a two, but the written description of what a two requires.
- A quoted span with a timestamp. Clickable back to the audio, not paraphrased.
- An absence rationale for zeros. Where the system looked and what it did not find.
- An override field with the reviewer's name and reason. Human correction has to be a first-class part of the record, not an edit that erases history.
Negative evidence is the hard part
Quoting a line to justify a positive score is easy. Justifying a zero means proving a thing did not happen, and "it is not in the transcript" is a weak claim if the transcript itself is imperfect.
The honest approach is to state scope. A good absence rationale says the call was searched for any request for a time commitment and none was found, and flags when transcript quality was poor enough that the finding is uncertain. Uncertainty flagged is better than confidence faked, and any system that never expresses doubt is hiding something. This is a design requirement we build into rubric scoring, not an optional feature.
Disputes are a rubric improvement engine
Track which criteria get disputed and how often disputes are upheld. Upheld disputes cluster. When they concentrate on one criterion, the criterion is ambiguous. When they concentrate on one rep, look at their audio quality before you look at their behavior. When they concentrate on one call type, the rubric is being applied to conversations it was not written for.
That feedback loop is worth more than a marginally better model. It is also what makes the program defensible when someone asks how a coaching decision was reached, and it keeps the humans in the coaching process in charge of the judgment.
Topics: explainability · evidence · disputes · AI scoring · fairness
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.