Will AI call scoring unfairly penalize reps with accents or fast, short call styles?
It can, if nobody checks. Transcription accuracy varies with accent, audio quality and background noise, and rubrics carry unstated length assumptions that mark efficient reps down. The defenses are checking transcript quality per rep before trusting any score, writing criteria around whether something happened rather than how long it took, and auditing score distributions by rep, shift and location for gaps you cannot explain.
Two different failure paths, with different fixes
The first is upstream. If speech recognition performs worse on a particular voice, headset or noisy branch office, the rubric is scoring a degraded transcript and every criterion suffers. No amount of rubric tuning fixes a bad transcript.
The second is in the rubric itself. Criteria written by people who talk a certain way encode that way of talking as correct. Both problems produce the same symptom, a rep who scores lower than their results justify, so you have to test for them separately.
Audit the transcript before you audit the rep
Compare transcript quality signals per rep: rate of inaudible markers, unusually short transcripts relative to call duration, garbled or nonsense tokens, and recognition confidence where the system exposes it. If one rep's calls show materially worse signals than peers, that is where to look first.
Then spot-check by hand. Pull three of their calls, read the transcript alongside the audio, and count the errors that actually changed meaning. A transcript can look messy and still support scoring, or look clean and drop the one sentence a criterion depended on.
Rubric assumptions that quietly penalize
- Thoroughness criteria that reward talking. A rep who books in ninety seconds with everything captured is better, not worse.
- Rapport criteria that reward small talk. Some customers want efficiency, and reading their preference is a skill, not a deficit.
- Interruption penalties applied symmetrically. Overlapping speech is normal in some regions and conversational styles.
- Vocabulary expectations. Requiring particular phrasing rather than particular meaning turns a language preference into a score.
Distribution audits and the human backstop
Periodically compare score distributions across reps, shifts, locations and call types. You are not looking for perfect equality, which would be its own red flag. You are looking for gaps you cannot explain by mix, tenure or measured outcomes.
Where you find one, investigate before acting on any of the affected scores. And keep the structural protection in place regardless: scores inform coaching, a human reviews before any consequence, and no system decides who is a good employee. Employment practices vary by state and change over time, so have your own counsel review how evaluation data is used in your organization. We build evaluation to support human coaching decisions, not to replace them.
Topics: fairness · bias · transcription · accents · human oversight
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.