Can the AI explain why it scored a call the way it did?
It can produce an explanation, but that text is a plausible account written after the fact, not a readout of how the score was computed. The version you can actually audit is different: require the system to cite the specific transcript spans that triggered each rubric criterion. Evidence you can go read beats reasoning you have to trust. If it cannot point at the sentence, it cannot defend the score.
Plausible is not the same as faithful
When you ask a language model why it produced an answer, it generates a reason. That reason is optimized to be convincing, not to be a true description of the computation that occurred. It will usually sound excellent, and it may be right, and you have no way to tell from the text alone.
This is the single most misunderstood point in AI governance conversations. Buyers ask "is it explainable" and accept a paragraph of reasoning as the answer. The paragraph is the least reliable part of the output.
What auditable explanation actually requires
Three things, and none of them are the model's self-narration.
- Evidence citation. Every criterion outputs the transcript span that satisfied or failed it, with timestamps. A supervisor can play those seconds and agree or disagree in about a minute.
- Structured criteria. A single overall score is unauditable by construction. A score decomposed into named criteria written by your own leadership can be argued with line by line. That is the whole design idea behind rubric-based evaluation.
- Stored input snapshot. The exact transcript, rubric version and model version used. Without it, re-examining a six-month-old score is guesswork.
The test that separates real from theatrical
Pick a score somebody disagreed with. Ask the system for the evidence, open the recording at the cited timestamp, and listen. Either the cited moment supports the criterion or it does not.
If the citation is vague — "the representative did not build rapport throughout the call" — you have narration, not evidence. Real evidence is a quote at a timestamp. Systems that cannot produce it tend to be defended with adjectives instead.
Why this matters more for coaching than for reporting
A miscategorized marketing source costs you a slightly wrong chart. A wrongly justified evaluation costs you a rep's trust, and once that is gone the program stops working regardless of how accurate it is.
Evidence-linked scoring is what makes an evaluation conversation about behavior rather than about the software. The manager plays the moment, the rep hears it, and the discussion moves on to what to do differently. That is the intended use of AI-supported coaching — the system assembles evidence, the human makes every judgment that affects a person.
Topics: explainability · rubric · auditability · evidence
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.