Why do two managers score the same call differently, and how do I fix it?
Because the criterion is written as a judgment instead of an observation. "Built rapport" invites disagreement; "used the customer's name before asking for the address" does not. The fix is calibration: have several reviewers score the same small set of calls independently, compare criterion by criterion, and rewrite every criterion where they diverge. Then test the rewrite on a fresh set of calls.
Disagreement is a rubric defect, not a people problem
When two experienced managers land on different scores, the instinct is to decide who is right. That is the wrong question. If a criterion can be read two ways by two qualified people, it will be read two ways by twenty reps, and no amount of retraining reviewers fixes the underlying ambiguity.
Treat every disagreement as a bug report against the wording. The reviewer who scored low and the reviewer who scored high are both telling you what they think the sentence means. The repair is in the sentence.
How to run a calibration session that actually produces changes
- Pick six to eight calls that span outcomes. Include two obvious wins, two obvious losses, and several ambiguous ones. Ambiguous calls are where the rubric breaks.
- Score independently and blind. No discussion first. If reviewers hear each other's reasoning before scoring, you learn nothing.
- Compare per criterion, never on totals. Two reviewers can agree on an 82 while disagreeing on six separate line items that happen to cancel out.
- Only discuss divergent criteria. Agreement needs no meeting time.
- Rewrite, then re-test on new calls. Re-scoring the same calls after discussion proves nothing except that people remember the discussion.
The criteria that always drift
A short list causes most of the trouble: tone, empathy, urgency, professionalism, and anything phrased as whether the rep "controlled" the call. Each of these is a summary of many small behaviors, and reviewers weight the behaviors differently.
The rewrite pattern is to name the behavior that made you feel the thing. "Empathy" becomes acknowledging the problem in the customer's own words before moving to scheduling. "Control" becomes ending each exchange with a question rather than a pause. That is the same rewriting discipline that rubric analysis depends on to produce scores that hold up.
Why AI scoring makes calibration matter more, not less
Automated scoring removes reviewer-to-reviewer variance, which is a real gain. It does not make the standard correct. A consistently applied vague criterion produces a consistently wrong result, applied to every call instead of a handful.
So the human work moves upstream. People argue about what good looks like, write it down precisely, and check periodically that the applied standard still matches their intent. The system applies it uniformly and shows its evidence. Your standards, machine-applied evaluation, human decisions about what to do next. Coaching workflows sit on top of that, not underneath it.
Topics: calibration · rubric design · inter-rater reliability · QA · scoring
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.