Skip to main content

Two managers scored the same call differently. Who is right?

Coaching & QA Published August 7, 2026
Short Answer

Usually neither — the rubric is ambiguous. Disagreement between two honest scorers is a measurement problem, not a judgment problem. Fix it by having both score the same calls blind, comparing item by item rather than on the total, and rewriting the wording of whichever items they split on. If agreement stays low, no coaching built on those scores can be trusted.

Compare items, never totals

Two managers can land within a few points of each other on the total score while disagreeing on half the individual items, with the errors cancelling out. The total looks reassuring and hides the problem completely.

Calibration only works at the item level. Line up both scorecards side by side and count how many items matched. That number, not the score gap, tells you whether your rubric is functioning.

Track the agreement number over time and treat it as a health metric for the rubric itself. If two managers agree on nine of twelve items, the three they split on are your entire backlog.

How to run a calibration session

  • Pick five calls that span the range: one clean booking, one clear miss, one angry caller, one price shopper, one messy edge case.
  • Score independently, without discussion, and without seeing each other's marks.
  • Compare item by item and stop on every disagreement.
  • Rewrite the item, not the person's opinion. If two reasonable people read it differently, the sentence is at fault.
  • Re-run on new calls a month later to confirm the rewrite worked.

Items that stay contested after two rewrites

Some items resist definition because the underlying idea is genuinely subjective. Showed empathy is the classic case. You have three options: replace it with an observable proxy, demote it to a note that carries no score, or drop it.

Dropping an item feels like lowering standards. It is not. An item that cannot be scored consistently was never enforcing a standard, only creating noise.

Be honest about the cost of keeping a contested item. Every ambiguous line on the scorecard adds noise to every rep's total, which makes real differences between people harder to see.

Where automated scoring helps, and where it does not

A model applying the same rubric to every call is consistent by construction. That is not the same as being correct, and the distinction matters: consistency means the comparison between two reps is fair, but the standard itself can still be wrong.

So the calibration work does not disappear when you automate — it moves. Instead of calibrating managers against each other, you periodically sample scored calls and check them against a human read, then adjust the rubric wording where the two diverge. That audit loop is a permanent part of running AI rubric analysis responsibly, and it belongs to the same coaching process everything else does.

Topics: calibration · inter-rater agreement · rubric design · management

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Coaching & QA

What counts as a bookable call?

A call from someone who could have been booked on that call: a real prospect or customer, in your service area, asking for work you actually perform, who …

Aug 7, 2026Read answer →

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001