Where does AI call scoring get things wrong?
Mostly where the audio is poor or the truth is not in the audio at all. Crosstalk, background noise, speakerphone and weak cell connections degrade transcription, and names, street names and model numbers are the first casualties. Anything requiring outside context — whether a slot was actually open, what a technician promised last week — cannot be scored from the recording alone.
Transcription failure modes
Speech-to-text is strong on clear conversational speech and weak everywhere else. The predictable trouble spots are worth knowing because they map directly onto rubric items you might otherwise trust.
- Crosstalk. Two people speaking at once produces garbled or dropped segments, and interruption-heavy calls are exactly the ones you most want to review.
- Proper nouns. Customer names, street names, neighborhood names and equipment model numbers are frequently mangled.
- Environment. A caller in a truck, in a mechanical room, or on a hands-free speaker is much harder to transcribe than one at a kitchen table.
- Very short calls. Ten seconds of audio rarely supports a meaningful score, and forcing one produces noise.
Judgment failure modes
Beyond transcription, some things are genuinely hard to read from text. Sarcasm and dry humor. Regional idiom. A terse, efficient rep who books the job in ninety seconds and scores poorly against a rubric written for longer conversations.
Bilingual and mixed-language calls are their own category — a call that switches languages mid-conversation can defeat both transcription and scoring. Bluefrog's voice translation work is still in development, and until that ships, mixed-language calls should be flagged for human review rather than scored automatically.
What is simply not in the audio
Whether the schedule actually had an opening that day. Whether the customer had been told something different by a technician on Thursday. Whether the rep was covering another queue. Whether the caller had already booked with a competitor before dialing.
Rubric items that depend on those facts cannot be evaluated from a recording. Some of them become answerable once the call is joined to the operational system — availability and prior job history live there — and some never do.
Designing a rubric around the limits
Two habits handle most of this. First, allow an explicit unknown mark instead of forcing a score when the evidence is not there, and exclude unknowns from the rate rather than counting them as failures. Second, flag low-confidence calls — bad audio, very short, language mismatch — into a human review queue.
A system that admits what it could not determine is more useful than one that always produces a number. The admissions are also a diagnostic: a sudden rise in unreadable calls usually means something changed in the phone system, not in the reps.
Topics: limitations · transcription · accuracy · call scoring
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.