How do I know why the AI scored a call the way it did?
Because a usable score arrives with its evidence: the rubric item, the mark, and the exact words from the call that produced it, with a timestamp. A score without a quote is an opinion you cannot verify, cannot coach from, and cannot defend when a rep disagrees. Evidence is also how you catch a rubric item the system is reading differently than you meant it.
A bare number fails the first coaching conversation
A rep is told she scored a six on discovery. Her first question is which part. If the answer is a shrug, the session is over — nothing gets changed because nothing specific was identified. Worse, the rep concludes the score is arbitrary, and every future score inherits that suspicion.
The same score with a timestamp and the caller's actual sentence attached produces a different conversation entirely. Both people listen to twenty seconds of audio and then argue about the behavior, not about the number.
There is a second reason to insist on this. Managers are busy, and a scorecard without evidence quietly trains them to accept the number. Once that happens nobody is checking the system, and errors in the rubric propagate for months before anyone notices.
What an explainable score record contains
- The rubric item, in the exact wording the team agreed to.
- The mark, and how it rolled into the weighted total.
- A verbatim excerpt from the transcript with a timestamp you can jump to in the recording.
- The gap — what was missing, stated as a behavior rather than a judgment.
- The counterfactual — what would have earned full credit on that item on this specific call.
Evidence is a debugging tool for the rubric itself
When you read the quotes behind a low-scoring item across twenty calls, one of two things happens. Either you see a real, repeated behavior gap, or you see that the item is being applied to situations you never intended — emergency calls scored for a value pitch, existing customers scored for identity capture they already gave.
That second case is a rubric bug, not a rep problem, and you only find it by reading evidence. This is a large part of what coaching data is for: tuning the standard as much as tuning the people.
The disagreement path has to exist
Reps should be able to flag a score. The manager opens the recording at the timestamp and either the mark stands, with the reasoning stated out loud, or the item's wording changes because it was ambiguous. Both outcomes are good; a system with no dispute path is one where quiet resentment accumulates instead.
Scores support human coaching decisions. They do not make them. Keeping evidence attached to every mark is what makes that claim real instead of a slogan — see how operational AI is meant to sit under human judgment rather than above it.
Topics: explainability · evidence · call scoring · trust
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.