Your Support Bot Talks Over Customers. Native Voice Agents Don't

Why removing the transcribe-think-speak pipeline changes enterprise voice support, and what breaks if you deploy it without the right guardrails.

SB

SmartBuddy Engineering Team

Autonomous Systems & AI Industry Solutions
Your Support Bot Talks Over Customers. Native Voice Agents Don't

Old-school IVR meant pressing numbers into a phone tree. The first generation of conversational bots improved on that, barely, by stitching together speech-to-text, a text model, and text-to-speech, three separate systems handing work to each other in sequence. That chain added two to four seconds of dead air before every reply, stripped out tone and emotion in the transcription step, and fell apart completely the moment a caller interrupted mid-sentence.

Native audio-to-audio models skip the stitching. Speech goes in, speech comes out, through one integrated system, and the difference is immediately obvious the first time you talk to one.

Two Architectures, Side by Side

code
Stitched pipeline (real, noticeable lag):
[Audio in] β†’ [Speech-to-text] β†’ [Text context] β†’ [Text model] β†’ [Text-to-speech] β†’ [Audio out]
                                                        β”‚
                                    tone, cadence, and inflection get lost here

Native audio-to-audio (the actual improvement):
[Audio in / video feed] ─────────► [Realtime neural engine] ─────────► [Audio out]
                                          β”‚
                        pitch, pauses, and natural interruption survive

Removing the transcription step doesn't just save time, it means the model can actually hear frustration, hesitation, a raised voice, and adjust its pacing in response, the way a person would. A transcript alone can't carry any of that.

What Actually Changes for Support

Stitched IVR (STT + LLM + TTS) Native audio engine A human agent, tier 1
Time to respond Several seconds, noticeable lag A few hundred milliseconds Comparable to a fast human
Handling interruption Talks over the caller Cuts off instantly when interrupted Natural
Can it see the screen No, voice only Yes, screen share or camera feed Yes, with the right tools
Cost to scale Cheaper per minute, but limited experience More expensive per minute Limited by headcount, not cost

What Actually Needs Tuning If You're Building This

Real-time voice support over WebRTC has a handful of settings that determine whether it feels natural or actively annoying.

Turn detection needs a sensible threshold, too low and normal background office noise triggers false interruptions, too high and it stops responding to real speech. A moderate setting is usually the right starting point, then tune against your actual call environment.

Silence padding matters more than it sounds like it should. Set it too short and the agent barges in the moment a caller pauses mid-sentence to think, which reads as rude and cuts off information you needed.

And the agent should be able to call backend functions, looking up an account balance, checking a ticket status, without breaking the conversational flow. Returning that data quickly keeps the exchange feeling continuous instead of stalling on a lookup.

code
const sessionUpdate = {
  type: "session.update",
  session: {
    modalities: ["audio", "text"],
    instructions: "You are an enterprise technical support specialist. " +
      "When a user shares their screen, reference specific UI elements " +
      "by their label. Keep responses under two sentences unless " +
      "explaining a multi-step troubleshooting procedure.",
    voice: "alloy",
    input_audio_format: "pcm16",
    output_audio_format: "pcm16",
    turn_detection: {
      type: "server_vad",
      threshold: 0.5,
      prefix_padding_ms: 300,
      silence_duration_ms: 500,
    },
  },
};

Three Things That Will Actually Bite You

Audio tokens cost meaningfully more than text tokens, and an unbounded voice session adds up fast. Set a hard session timeout, ten minutes is a reasonable starting point, and route anything unresolved to a human rather than letting the meter run indefinitely.

If a caller is on speakerphone without headphones, the microphone picks up the assistant's own voice output and the system can end up interrupting itself in a feedback loop. Acoustic echo cancellation in the browser media stream isn't optional here, it's the difference between a usable deployment and an unusable one.

And low-quality microphones or loud background environments can cause the model to hallucinate words out of static, which in a support context means confidently acting on something the caller never actually said. A high-pass filter on the incoming audio, cutting anything below roughly 80 Hz, catches a lot of that before it ever reaches the model.

Frequently Asked Questions

Can it look something up while still talking?

Yes, function calls can run asynchronously while the conversation continues, so it can check a billing record mid-sentence and fold the answer back in once it returns, without an awkward pause.

Is this usable for healthcare, compliance-wise?

It can be, with the right enterprise agreement in place, a signed BAA and zero data retention configured on the API side. Without those two things explicitly set up, don't treat it as compliant by default.

Does the customer need special hardware on their end?

No. Any modern browser that supports WebRTC and the MediaStream API can handle this, there's nothing to install on the caller's side. Getting the turn-detection tuning and the echo cancellation right on the first deployment saves weeks of after-the-fact complaints. [SmartBuddy builds and tunes these voice pipelines for support teams β†’](https://smartbuddy.cloud/start-a-project.html)

Did you find this technical breakdown helpful?

Tap to rate this guide · 8 views

Comments

Comments are reviewed before appearing publicly.

No comments yet β€” be the first.

πŸš€ Ready to Deploy Autonomous Skills in Production?

Get this skill (and 29 more) in the SmartBuddy Shop, or work with our engineering team to architect custom multi-agent workflows for your company.