How does call AI know who is talking when the customer and the rep are on the same line?
It depends on how the call was recorded. If your phone system records each side on its own channel, speaker separation is close to trivial and highly reliable. If everything is mixed into one mono file, the system has to infer speaker turns from voice characteristics, and it degrades badly on speakerphone, three-way calls and interruptions. Fixing the recording configuration solves more speaker problems than any model change.
Two-channel recording is the single biggest quality lever
Most telephony platforms can record a call as stereo: the customer on the left channel, the agent on the right. When that is turned on, the question of who said a given sentence is not a prediction at all. It is a fact carried in the audio file, and every downstream metric that depends on speaker identity inherits that reliability.
When the platform writes a single mixed mono track instead, the system must perform diarization: cluster the audio into voices, then decide which cluster is the agent. That inference is usually good and occasionally wrong, and it is wrong in ways that are easy to miss. Before anyone debates model quality, check which format your recordings are actually stored in. This is a recurring first step in any call intelligence deployment.
The four situations where speaker separation breaks
- Speakerphone in a truck or a shop. Both voices arrive through one microphone with room echo, so the acoustic difference the model relies on shrinks.
- Three-way and conference calls. A dispatcher who conferences in a technician adds a third voice mid-call, often on a worse connection than either original party.
- Heavy interruption. When two people talk over each other for several seconds, the boundary between turns is genuinely ambiguous in the audio.
- Warm transfers. The agent voice changes partway through the recording while the channel stays the same, which mono diarization frequently reads as one continuous speaker.
Why it matters more than it sounds like it should
Speaker attribution is load-bearing for a surprising number of metrics. Talk ratio, interruption counts, who asked for the appointment, whether the agent or the customer raised price first, and whether a required disclosure was actually spoken by the employee all depend on it. Get the speaker wrong and the transcript is still readable, but the scorecard built on top of it is quietly inverted.
This is also why evaluation data deserves a sanity check before it is used in coaching conversations. If a rep disputes a talk-ratio number, the first thing to inspect is not the model output but whether that particular call was on speakerphone or transferred.
What to do about it
Turn on dual-channel recording where the platform supports it, even if it costs storage. Where it does not, treat mono calls as a separate population: still transcribe and classify them, but be cautious about publishing per-speaker behavioral metrics from them. Flag conference and transferred calls so they can be excluded from speaker-sensitive reporting rather than silently averaged in.
Most of this is configuration work at the phone system, not modeling work, which is why it belongs in the integration phase of a project rather than being discovered later when someone questions a number.
Topics: diarization · recording setup · transcription · audio
Have a version of this question about your own business?
The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.