5 Voice Agent Failure Modes You'll Hit in Week One — Venky B, Plivo
Read the talk
5 Voice Agent Failure Modes You'll Hit in Week One
Venky B of Plivo explains why a working voice demo can fail on a real call, and how model selection, transcript cleanup, typed fields, speech normalization and conversation timing change the result.
From a talk by Venky B
At a glance
Ideas worth remembering
Optimize time to first audio across the complete response path. Fast token throughput and a good median model latency can still leave callers waiting.
Venky reports a 2.5–3× multilingual token-fertility advantage for Gemma 4 over Qwen 3.5 in his team's evaluations. Fewer tokens per word can improve word-generation speed; the reported comparison does not establish a reproducible advantage for unspecified checkpoints or workloads.
Use call state to narrow recognition and interpretation: boost relevant keywords dynamically, clean transcripts with context and normalize multilingual scripts.
Define typed fields before collecting data. Validation should turn suspicious input into confirmation or repetition, and field-level evaluations should identify which collection behavior fails.
Prepare model output for speech with formatting cleanup, pronunciation dictionaries, entity-specific pacing and application-owned normalization.
Turn detection, barge-in and backchanneling remain separate engineering concerns. A modular pipeline can support them without requiring a dedicated speech-to-speech model.
The pipeline works; the call still fails
A voice agent can sound great in development and start failing as soon as real callers arrive. Venky B, founder of Plivo, opens with that familiar transition. Plivo began building voice and SMS APIs in 2011; he reports that its platform carries over a billion voice calls each month. That is the telephony scale behind these observations, rather than a count of AI-agent conversations. 0:34
Plivo's agent offering sits above its own SIP trunking and audio-streaming infrastructure. It includes a programmable speech pipeline and a no-code visual studio. The programmable offering combines separate components rather than using a single model that takes speech in and produces speech out.
The first implementation usually connects speech-to-text, an LLM and text-to-speech, with turn detection deciding when the conversation should advance. LiveKit and Pipecat are examples of frameworks used to orchestrate those pieces. Each component can appear fast enough in isolation, and the proof of concept can work. Production exposes the gaps between them: waiting too long, misreading an identifier, collecting an invalid value or saying a correct answer badly.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure one: intelligence has to fit inside the pause
Time to first audio measures the interval between the caller finishing and the agent starting to speak. Venky reports that teams often aim for less than 550 milliseconds, while many deployments land between 750 and 1,200 milliseconds. Beyond 1.2 seconds, he sees callers begin hanging up. These are production observations from his experience, rather than a universal threshold for every conversation. 5:15
That pause forces a three-way decision between cost, intelligence and latency. Longer model reasoning can improve an answer, but it also keeps the caller waiting. The speaking model therefore usually needs thinking turned off. Better instruction following and tool calling still help; gains that depend on spending extra time reasoning are difficult to use on this immediate response path.
The LLM is the largest latency contributor in the pipeline Venky describes. His frontier-model examples have a median time to first token of roughly 450–500 milliseconds on a good day, with P90 or P95 reaching 1.2–1.3 seconds. Those percentiles describe the slower end of the response distribution. Time to first token is only one part of time to first audio: the system still has to turn generated text into audible speech. A satisfactory median can conceal pauses that make the conversation frustrating.
Fast token generation does not automatically solve the wait for the first token. Cerebras and Groq enter the discussion as high-throughput options, but Venky describes dedicated capacity as expensive and, at the time of his experience, booked twelve months ahead. That commitment creates another risk: the team must choose infrastructure for a model it expects to keep using despite rapid model changes.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose models for spoken words and successful actions
Self-hosted open-source models are Plivo's answer to the LLM latency problem. Running models on its own GPUs adds operational work, but Venky describes this as a practical way to target consistently less than 300 milliseconds at the model layer while balancing cost and capability. That target remains separate from the complete caller-to-agent audio delay. 9:27
Venky names Qwen 3.5 and Gemma 4 as the models his team evaluated. He says both worked well for English and reports Gemma as 2.5–3× better in their multilingual token-fertility evaluations—how many tokens it takes to generate one word in a language. He connects this to faster word generation under otherwise equal conditions. This is his team's reported comparison; the supplied material does not specify the exact checkpoints, language set or evaluation configuration, or provide independently reproduced results. 10:28
The mechanism connects tokenization to the caller's wait. If producing a word requires more tokens, the model must generate more pieces before it has produced the same spoken content. Under otherwise equal conditions, fewer tokens per word can mean faster word generation; that is the explanation Venky gives for the multilingual contrast he observed. Tokens per second alone therefore misses an important part of voice performance: how quickly those tokens become words in the caller's language.
Model size depends on what the agent must do:
- Generic conversation: Venky describes mixture-of-experts models using a three- or four-billion shorthand as useful out of the box. The talk does not specify checkpoints or whether that shorthand refers to active parameters. Active and total parameter counts differ for MoE models: the supplied Gemma 4 model card lists 25.2B total and 3.8B active parameters for its MoE model, so the shorthand should not be read as a 3–4B total model size for GPU hosting. The drawback comes when fine-tuning: his team has found that modifying these models can damage their behavior.
- Industry-specific work: For applications that need deeper adaptation, such as healthcare, his starting point is an eight- or twelve-billion model. The selection criteria include fast generation, good instruction following and a high tool-calling success rate.
- Separate speaking and action models: Another architecture uses a small conversational model, perhaps three billion, for talking and a larger model for tool calls. It assigns the more demanding action work to a model chosen for that capability.
Plivo runs two flavors: fine-tuned models for specific industries and an out-of-the-box mixture-of-experts model for more general uses. The useful decision comes before fine-tuning: establish whether the model follows the required instructions and calls tools successfully. A model that already does those jobs may need changes to the surrounding pipeline more urgently than changes to its weights.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure two: transcription errors travel downstream
Even a strong transcription engine needs help on real calls. Venky contrasts word error rates of 4–6 percent on known evaluation sets with double-digit rates on noisy calls involving accents and domain vocabulary. The errors concentrate in consequential places: proper nouns, jargon, missing phone-number digits and omitted pieces of long addresses. 13:35
Code-switched speech creates a different kind of failure. English words can arrive written in the script used for Hindi; Hindi can arrive written in Latin characters. The LLM may continue in the incoming script, and the speech synthesizer can then pronounce the response incorrectly. The mismatch begins in transcription but changes both the generated text and the eventual audio.
Three mechanisms clean up the input before it reaches the main conversational model:
- Dynamic keyword boosting: Boost relevant proper nouns during the phase of the call where they are expected. Keeping a large keyword list active throughout the call can encourage unwanted substitutions; changing the list with the call state narrows what the recognizer should listen for.
- Contextual post-processing: An LLM with domain context can interpret a suspicious transcription. Venky's example is an
Einside a phone number, where3is a plausible correction. The field-collection step that follows must decide whether to confirm that interpretation or ask again. - Transliteration: An LLM or a neural transliteration engine can normalize the script of multilingual transcripts before downstream generation.
The goal is a consistent cleaned transcript regardless of which transcription engine produced it. Recognition supplies an imperfect reading of the audio; the application supplies the context needed to interpret that reading. Cleanup alone is insufficient, though: a plausible string still needs to become a valid value.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure three: decide the field's shape before asking
Data collection is a voice UX problem. Venky estimates that 50–60 percent of agents struggle badly here. His remedy borrows from Python dataclasses, Pydantic, TypeScript's Zod and ordinary form fields: decide what kind of value the application needs before asking the caller for it. Plivo reports collection accuracy improving from roughly 30 percent to 95 percent, later citing 95–97 percent without fine-tuning. The evaluation population and scoring method are unspecified, so those figures describe Plivo's reported experience rather than a transferable benchmark. 18:00
Follow the phone-number example through that change. As free text, a transcript containing an E is simply a string the LLM has to interpret. As a phone-number field, it has allowed characters and an expected length. The E now creates a detectable validation error. The agent can propose 3 and confirm it with the caller, or reject the value and request a repetition. The visible change is a different next action: the suspicious character triggers a correction conversation instead of passing silently into collected data.
Where does the field definition change the phone-number flow? The diagram follows the same E through validation and the two recovery paths. The field supplies the rules that expose the error. A likely correction then remains a proposal for the caller to confirm; the alternative is to collect the value again.
Other fields need their own collection behavior:
- Names: A difficult name may require letter-by-letter spelling and confirmation. Repeating the same recognition attempt does not supply the structure that spelling does.
- Relative dates: “Next week, Wednesday, eight” needs the current date to resolve the calendar day and clarification to distinguish 8 AM from 8 PM. Treating it as a datetime field makes the missing information explicit. Tool calls can do much of the date-resolution work alongside the LLM.
The same decomposition changes evaluation. Test phone-number collection, name spelling and datetime resolution at the field level, as small units of behavior. A failed field test identifies the broken collection step directly, instead of requiring hundreds of complete conversations to reveal it. Those checks target collection reliability; they do not establish that the rest of the call behaves correctly. 20:51
Call state ties the approach together. At each point, the agent knows which field it is collecting and what that field permits. This gives recognition, interpretation, validation and confirmation a narrower job. Plivo's reported gains came from structuring that context, rather than endlessly adjusting prompt wording and hoping the model would become more obedient.
Define allowed characters and expected digit count before asking.
The phone-number type exposes the invalid character. The agent then confirms a proposed correction or asks the caller to repeat.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure four: prepare the answer for speaking
Correct generated text can still make bad audio. A normalization layer between the LLM and text-to-speech gives the application control over how the answer is spoken. Passing model output directly to the synthesizer leaves formatting, entity pronunciation and reading conventions to the engine. 22:24
The output layer has several separate jobs:
- Remove display formatting: Strip emojis and Markdown before synthesis. Venky notes that orchestration frameworks can handle this with configuration flags, which still need to be set.
- Specify pronunciation: Use custom dictionaries for proper nouns, brands and acronyms.
- Slow down for entities: Reduce speaking speed to 0.8× or 0.7× when reading an email, phone number or a name letter by letter, so the caller can hear the individual parts.
- Normalize reading conventions: Prepare emails, currencies and dates in the application's own layer rather than depending entirely on a TTS engine's interpretation.
This also makes switching synthesizers less disruptive. If the first provider is unavailable, or the application changes providers, its own normalization rules can remain in place. Each engine may still need pronunciation configuration, but the application retains responsibility for preparing the content it wants spoken.
Venky's first pronunciation check is deliberately personal: can the agent say his surname, Balasubramanian, and his company's name, Plivo? These are concrete words the product should handle, rather than an abstract voice-quality score. For a customer-facing platform, the same control should be available to customers, so they can specify the names and terms their own callers need to hear.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Failure five: conversation timing needs its own engineering
The ending returns to the timing layer around the pipeline: end-of-turn detection, barge-in and backchanneling. These concern when the agent takes its turn, how the caller interrupts it and how conversational acknowledgments fit into the exchange. Generating a good sentence quickly does not settle those interaction decisions. 25:06
Venky closes with a specific architectural claim: a dedicated speech-to-speech model is not required to support these behaviors; Plivo has found ways to implement them in a pipeline. He moves through the final slides without explaining the detection or interruption algorithms, so this establishes an implementation option rather than a recipe for reproducing it. The modular pipeline remains viable, provided conversation control receives attention alongside recognition, generation and synthesis.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Your Voice Agent Doesn't Need a Frontier Model
A companion topic for choosing models under the conversational latency budget.
- Why ChatGPT Keeps Interrupting You
Continues the conversation-timing question introduced briefly at the end.
- Pydantic is all you need
Connects the typed-field approach to structured LLM outputs.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
Let's just do a couple of quick
- 0:14
questions and then we'll jump right in.
- 0:16
Uh, how many of us in the room here have
- 0:20
built voice AI agents?
- 0:24
Okay, that's a that's a pretty good
- 0:27
audience here. And how many of you guys
- 0:28
have built AI agents that have been
- 0:30
deployed in production?
- 0:34
Not bad. Okay, cool. So, uh we'll talk
- 0:38
about what typically happens, right?
- 0:40
Like everyone's talking about wise AI
- 0:42
agents. Uh the
- 0:47
you know, one pill solution to pretty
- 0:48
much everything in the world today uh is
- 0:50
is wise agents. So, everyone's building
- 0:52
one and trying to deploy that. They
- 0:54
sound great when you're sort of building
- 0:56
that in your dev sort of landscape and
- 0:59
then the moment you take this to from a
- 1:02
proof of concept to production things
- 1:04
start failing. Uh so we'll walk through
- 1:06
these five different angles of like how
- 1:09
uh or what we have seen uh at PO with
- 1:13
VIA agents but just before that a quick
- 1:16
uh intro from from my side uh I am Wenke
- 1:19
the founder and CEO uh she used the
- 1:23
title agent engineering manager u I'm
- 1:25
calling myself chief agent officer uh
- 1:28
from a from a title standpoint u okay so
- 1:32
what is what is uh you know why why are
- 1:34
even qualified for this this discussion
- 1:37
and and like uh what what are we seeing
- 1:39
that a lot of companies don't get to
- 1:41
see? I I'll talk a bit about our journey
- 1:43
in terms of like how we've uh come along
- 1:46
so far and then jump right in. Uh you
- 1:49
know we were we've been around for about
- 1:51
14 years. Our journey has been a
- 1:53
developer API platform and then now an
- 1:56
uh you know an AI agent business. We
- 1:58
started with voice and SMS APIs back in
- 2:00
the day uh 2011 and then uh you know now
- 2:04
we are primarily focused on our AI agent
- 2:07
offering uh the the full stack on our
- 2:10
platform. We we see over a billion voice
- 2:14
calls each month across the globe. Uh
- 2:16
and which is where we've seen a lot of
- 2:18
these uh you know patterns emerge in
- 2:20
terms of like how when we work with our
- 2:21
customers what happens on their voice
- 2:24
agents in in production. Uh we're a uh
- 2:27
90 member team and uh we've we have uh
- 2:31
50 million funding in the bank. Fun
- 2:33
fact, this is not from external VC
- 2:36
investors. This is all from being a
- 2:38
profitable company having put that cash
- 2:39
in the bank over over these years. Uh
- 2:44
some customers we we power across the
- 2:46
globe. Uh you know, we've just left some
- 2:48
some logos in there. But primarily from
- 2:50
an offering standpoint, u I would sort
- 2:53
of cohort this into three different
- 2:54
buckets. One is a programmable AI agent
- 2:58
offering. We call it uh I mean it's a
- 3:00
speech pipeline, not a true
- 3:01
speech-to-pech product yet, but that's a
- 3:04
that's a programmable offering. We also
- 3:05
have an AI agent studio. It's a no code
- 3:08
visual uh builder. And then we like I
- 3:11
said, we started with voice APIs. So we
- 3:14
obviously have built this out over the
- 3:16
last 14 years, the SIP trunking and the
- 3:18
audio streaming layers. So we don't rely
- 3:20
on other folks for the telefony or the
- 3:22
carrier layer. Like that's the
- 3:24
breadandbut business we've built over
- 3:25
all these years and and that's on top of
- 3:28
which our AI uh agent platform sits.
- 3:32
Okay, with that uh let's get into this,
- 3:36
right, which I'm I'm sure since you guys
- 3:38
have all built AI agents, you've all
- 3:40
seen this or you know built this in in
- 3:43
one manner or another and we'll spend
- 3:44
more time on this in terms of like how
- 3:47
uh the entire pipeline looks, right? Uh
- 3:50
what we see with customers is and and
- 3:52
I'm sure you guys can all relate to this
- 3:54
is you know anyone thinking about AI
- 3:56
agents what they do is they pick a bunch
- 3:58
of these orchestration frameworks and
- 4:01
they do a pretty good job live kit or a
- 4:02
pipecat you know build their AI oen on
- 4:05
top of that uh they think they can just
- 4:07
sort of orchestrate these different four
- 4:09
layers speechtoext lm uh and and TTS
- 4:13
with turn detection in between and we're
- 4:16
off to the races like my AI engine agent
- 4:17
works in a in a P and it's good to work
- 4:20
in production. Uh typically that's what
- 4:22
happens. They sort of measure their
- 4:24
latencies and you can see some
- 4:26
indicative latencies on on this slide at
- 4:28
at each layer and they're like yeah this
- 4:30
this uh seems good for me for what I
- 4:33
need. So let let's position production
- 4:36
and then the the production w uh sort of
- 4:39
start to kick in and and you see all
- 4:41
sort of failure modes which we are going
- 4:42
to spend you know most of the time on on
- 4:45
in this talk at least. Uh I've kept some
- 4:48
time at the end for Q&A if you guys want
- 4:50
to have uh you know questions but we'll
- 4:52
jump right in from from this to uh you
- 4:55
know different failure modes we see.
- 4:57
Let's start with
- 5:00
you know the first one which everyone
- 5:01
talks about like this is the most spoken
- 5:03
about failure mode which is latency. U I
- 5:06
think we have a few AI agent talks today
- 5:09
or AI agent talks today. Um I'm pretty
- 5:11
sure like everyone everyone's going to
- 5:12
touch upon this specific failure mode
- 5:15
which is why I'm bringing this right up
- 5:17
uh in in terms of uh you know some like
- 5:21
how this entire experience is for uh
- 5:24
users right uh typically most folks
- 5:27
measure this by time to first audio so
- 5:30
the time when you user stop speaking to
- 5:34
your agent starts speaking right and I
- 5:37
think you've you've probably seen this
- 5:38
if you guys have built voice agents on,
- 5:40
you know, what uh good or natural feels
- 5:43
like, what uh sort of annoying feels
- 5:47
like or noticeable feels like, and then
- 5:48
what annoying feels like, which is, you
- 5:51
know, different tiered steps. Uh we
- 5:53
notice, you know, most people want to be
- 5:57
under 550 cuz that's what's advertised
- 6:00
by, you know, platforms or uh you know,
- 6:03
solutions or or or or layers. But I
- 6:06
think most end up between 750 to 1.2. uh
- 6:08
that's where most of the folks end up
- 6:10
at. Uh the really bad performing ones
- 6:12
end up you know more than 1.2 and then
- 6:15
you start to see users uh hang up. Uh
- 6:18
now I I'll share with you like what
- 6:20
we've seen practically in in uh
- 6:23
production with uh customers using this
- 6:26
with at at different layers and then you
- 6:29
know solutions to uh some of these. The
- 6:32
way we want to think about this layer is
- 6:35
sort of a balance between these three
- 6:37
which is cost, intelligence and latency,
- 6:42
right? And and and why do I bring these
- 6:45
three up? Because they're sort of
- 6:47
interrelated. I think one of the things
- 6:48
I was just chatting with uh you know a
- 6:50
couple of folks outside one of the
- 6:51
things last one year we've seen lot of
- 6:53
innovations lot of intelligence spike on
- 6:56
the LLM side of uh things right and most
- 7:00
of the you know intelligence has come in
- 7:02
in terms of thinking or uh you know
- 7:05
reinforcement learning and and so on and
- 7:07
so forth the irony with voice agents is
- 7:10
like almost always your the the LLM or
- 7:13
the agent that's talking has to have
- 7:16
thinking turned
- 7:17
Right. So all the advancements we've had
- 7:20
in the LLM layer in the last one year
- 7:23
like none of that even apply here now.
- 7:25
Right? You obviously you have you know
- 7:27
better models that can do you know
- 7:29
better instruction following or tool
- 7:30
calling but pretty much all of your
- 7:32
intelligence that's been built in on the
- 7:34
thinking layer is all off by default if
- 7:36
you want it to be fast enough. So so
- 7:38
that's one of the ironies that we come
- 7:40
up with. So then how do you sort of
- 7:41
balance intelligent cost and latency?
- 7:44
Let's let's look at some of these uh you
- 7:46
know options uh that are out there in
- 7:48
the market right so and I'm specifically
- 7:50
picking LLM because if you looked at the
- 7:52
previous chart LLM is u you know sort of
- 7:57
your highest latency bucket that adds to
- 8:00
this right and uh if you look at you
- 8:02
know frontier models which I think most
- 8:05
folks start by default your your openi
- 8:09
your clouds your geminis u you know p50
- 8:11
ttfftd is roughly around 450 to 500 on
- 8:15
on a good day and it can get spiky,
- 8:18
right? It can it can uh you know P90 P95
- 8:21
can go easily upwards of 1.2 1.3 seconds
- 8:24
even uh and and that's not good for the
- 8:27
overall agent experience.
- 8:30
So so that so that's your frontier
- 8:31
model. Now there's another options which
- 8:34
is your your cerebrus or or the gro that
- 8:37
is famous and popular for spitting out a
- 8:39
lot of tokens or or tokens very fast,
- 8:42
right? uh these work but for you to get
- 8:45
dedicated latency or time to first token
- 8:48
on these you need dedicated capacity and
- 8:50
that is really expensive that's where I
- 8:52
spoke about the cost uh as as being one
- 8:54
of the things to balance right it's
- 8:56
really expensive and then like you talk
- 8:58
to anyone from the gro team or the
- 8:59
cerebrus team they'll tell you you need
- 9:01
to book 12 months in advance for
- 9:03
dedicated capacity they're booked out
- 9:04
for the next 12 months so so that's
- 9:07
that's a pretty expensive option and
- 9:08
then you really need to be sure that the
- 9:11
model you're deploying on some of these
- 9:12
infra layers uh will be here 12 months
- 9:17
from now and and it's a it's a big
- 9:18
investment and a big unknown. So, so
- 9:21
what's a realistic option for production
- 9:23
grade uh
- 9:27
agents that are that are good quality
- 9:29
and end up balancing uh three of these u
- 9:33
this is what has worked for us u which
- 9:37
is the open source models u there are
- 9:40
obviously a lot of them in terms of like
- 9:42
the variety and and variations you can
- 9:44
pick I'm specifically talking about the
- 9:46
two we work with u quen 3.5 and gemma
- 9:50
four. These are uh you know kind of
- 9:53
cutting edge open source models right uh
- 9:56
out in the market right now and we've
- 9:59
done a lot of benchmarking around this
- 10:01
in how they work. It it can be scary to
- 10:04
think like okay I have the models now I
- 10:08
have to host them you know run them on
- 10:10
my own GPUs and so on and so forth but
- 10:12
if you are consistently targeting under
- 10:15
300 ms u this we've seen this to be a a
- 10:19
great option to balance between latency
- 10:21
cost and intelligence now some more deep
- 10:24
dive here if you're doing only English
- 10:27
uh quen 3.5 or GMA both work fine but if
- 10:30
you're doing multilingual uh right
- 10:32
international audiences different
- 10:33
languages uh Gemma 4 is a much better
- 10:36
model for that uh we've seen uh token
- 10:42
fertility evals essentially what that
- 10:43
means is if if I were to dejargonize
- 10:45
that is like how many tokens does it
- 10:47
take to generate one word in that
- 10:49
language okay so Gemma is much much
- 10:52
better at least 2.5 to 3x better than
- 10:55
quen 3.5 from that perspective so your
- 10:58
time to words is much faster on Gemma or
- 11:02
everything else equal right on a on a
- 11:04
multilingual basis. Now what sizes do
- 11:07
you pick at the LLM layer? Uh the
- 11:10
mixture of expert usually works fine. Uh
- 11:13
the three or four billion mixture of
- 11:15
expert usually works fine. The the
- 11:17
problem with mixture of expert is like
- 11:18
if anyone goes down wants to go down the
- 11:21
direction of fine-tuning that can be a
- 11:23
challenge uh because fine-tuning mixture
- 11:24
of experts models are not easy. Uh you
- 11:27
can end up breaking the model uh a a lot
- 11:30
of times. So, so that's one challenge we
- 11:32
see with Make sure experts, but usually
- 11:33
out of the box, it gets you 90% closer
- 11:37
to where you want to be like even
- 11:39
without any fine-tuning or or or custom
- 11:42
work done on the model. Uh so that's the
- 11:44
advantage of mixer experts. Uh now, if
- 11:47
you want to fine-tune and and you you
- 11:49
want to go deeper and say like look, I'm
- 11:51
working for a specific domain,
- 11:52
healthcare, what have you, right? uh and
- 11:55
I want to make sure I I'm able to
- 11:56
fine-tune my model. You want to start at
- 11:58
least with uh the 8 billion 12 billion
- 12:00
at least uh from where we are today.
- 12:02
Maybe maybe six months from now a 4
- 12:04
billion 4 billion model beats the 8
- 12:07
billion model uh hands down. But for
- 12:09
today uh what we've seen is you minimum
- 12:13
need a 8 billion or 12 billion model. Uh
- 12:15
cuz you're looking for two things in
- 12:17
these models. One obviously fast tokens
- 12:19
but uh good instruction following. Okay.
- 12:23
And the second thing is like very high
- 12:25
uh success ratio in tool calling because
- 12:28
if you can do these two things well then
- 12:30
you are on to like 70 80% there from not
- 12:33
even having to fine-tune it fine-tune
- 12:35
any model like models will work out of
- 12:37
the box right u so so that's uh been our
- 12:40
recipe we've actually uh we run two
- 12:43
flavors one a fine tune model
- 12:45
for specific industries and then for uh
- 12:49
you know most generic use cases uh MOE
- 12:53
model just works out of the box. Uh
- 12:55
there are a few more tips and tricks
- 12:56
we'll talk about in the upcoming slides
- 12:58
where we see failure models, but but
- 13:00
that's where we stand from a from a
- 13:02
latency LLM standpoint. Um all right,
- 13:06
I'm running tight on time, so I'm going
- 13:08
to fast track this. U now there are a
- 13:11
couple of other flavors in this. Uh
- 13:12
people build agents with a a mixture of
- 13:15
models. What they do is you know for u
- 13:18
the the talking part of it they have a
- 13:20
conversational model which is a much
- 13:21
lower smaller model and then you know
- 13:24
maybe even a three billion model and
- 13:26
then for tool calling they have a much
- 13:27
larger model so that they have a
- 13:29
improved tool calling success ratio
- 13:31
there. Uh
- 13:35
sorry the second one is uh assume your
- 13:39
transcriptions are going to be brittle
- 13:41
like that's that's uh something you want
- 13:44
to sort of uh live by when you're
- 13:47
building AI agents even if you have the
- 13:49
best transcription engine out there and
- 13:51
I I I'll show you why right like the the
- 13:54
the state-of-the-art transcription
- 13:56
engines out out in the market u you know
- 13:59
sort of get you to four to 6% word error
- 14:03
rate right and this is on known eval
- 14:05
sets on real world noisy calls with you
- 14:10
know sort of uh accents like people
- 14:12
having different sort of accents uh
- 14:14
domain vocabulary and so on and so forth
- 14:17
like those usually end up in the double
- 14:18
digits from a word erate perspective
- 14:21
right uh now you obviously you can
- 14:22
fine-tune you know pick up an open
- 14:24
source model and fine-tune uh but we see
- 14:27
typically like what breaks here often
- 14:29
and there are patterns s here in terms
- 14:31
of what breaks. So, proper nouns,
- 14:33
jarens, uh phone numbers like random
- 14:37
missing digits with phone numbers, uh
- 14:39
wrong substitutions. I I'll walk through
- 14:40
some examples of like how you solve for
- 14:42
these addresses when you're trying to
- 14:44
collect a long address. Uh you know, the
- 14:47
the transcription engine could just end
- 14:49
up missing some parts of it.
- 14:52
Code switch languages. I I'll just take
- 14:54
a example of a language I speak because
- 14:56
that's was easy for me to put on the
- 14:58
slide. uh where you know like if you
- 15:01
were to sort of take English but written
- 15:04
in a different script uh that's what's
- 15:05
used for Hindi right like this is
- 15:08
English written in that script right
- 15:10
whereas like the actual English version
- 15:11
of this is hello how are you so if if
- 15:14
I'm addressing an audience in a
- 15:15
different country where I have code
- 15:17
switched languages and I start getting
- 15:19
my English in a different uh sort of
- 15:21
script everything starts breaking from
- 15:24
the transcription engine to the LLM
- 15:26
layer and then beyond because your LLM
- 15:28
starts then producing output in that
- 15:29
sort of script a lot of times and then
- 15:32
your TTS messes up. Okay. So, so this is
- 15:35
uh very important to be careful about
- 15:37
and if you want to build your agent
- 15:39
independent of the transcription engine,
- 15:41
you need to build a layer that
- 15:43
normalizes all of this, right? We'll
- 15:44
talk about solutions in a minute. And
- 15:46
there is the other case which is Hindi
- 15:48
in Latin or or or you know Roman, right?
- 15:51
Which is like this is Hindi but it reads
- 15:54
English which again messes up everything
- 15:56
uh you know downstream. Those are just
- 15:58
examples. This applies to, you know,
- 15:59
Arabic, Mandarin, uh, Japanese, what
- 16:02
have you. Uh, pretty much any language.
- 16:04
So, what actually moves the needle with
- 16:07
a at the transcription layer? Uh, for
- 16:10
prop proper nouns, we recommend uh you
- 16:13
using not just keyword boosting. I think
- 16:15
a lot of transcription engine engines
- 16:17
provide you keyword boosting where you
- 16:18
can put in specific words into their
- 16:20
engine, but doing dynamic keyword
- 16:22
boosting. What that means is don't keep
- 16:24
the keyword for the entire state of the
- 16:26
call. just add that dynamically when you
- 16:29
think you need that as an answer so that
- 16:32
you get the highest accuracy. Meaning at
- 16:34
different states of the call, the
- 16:35
transcription engine will have different
- 16:38
uh keywords boosted during different
- 16:40
phases, right? Uh and that's what we've
- 16:42
seen works best because if you just
- 16:44
pollute your context of the
- 16:45
transcription engine with tons of
- 16:47
keywords, it'll start hallucinating
- 16:49
again, right? So, so that's what we see
- 16:50
typically working best. Uh
- 16:54
yeah, post-process post-process your
- 16:56
transcripts with an LLM, right? Cuz your
- 16:58
LLM has domain context. Your
- 17:00
transcription engine does not. So a lot
- 17:02
of words that it would say uh I'll give
- 17:04
you some examples may not make sense.
- 17:06
This is transcription like a phone
- 17:08
number from a transcription engine.
- 17:10
Right? Like what do you think that E is?
- 17:13
Right? If you give it to an LM, it knows
- 17:15
that's a three. Similarly, like what
- 17:17
that one is, it's a digit one. So, so
- 17:20
your transcription engine a lot of times
- 17:21
could mess that up, but when you
- 17:23
postprocess it with the LLM layer, it'll
- 17:26
instantly correct that from a collection
- 17:28
standpoint. I mean, uh, and and the last
- 17:30
one, like I said, uh, transliteration is
- 17:33
your ST output that's sort of u, you
- 17:36
know, multilingual also gets normalized
- 17:39
using either an NLM you first
- 17:42
transliterated or, you know, use some
- 17:44
kind of a neural uh, transliteration
- 17:47
engine. There are a lot of them open
- 17:48
source. You can just pick one of them,
- 17:50
right? Uh that would do all of that work
- 17:52
for you. Send cleaned transcripts
- 17:54
consistently independent of the
- 17:56
transcription engine to your LLM.
- 18:00
All right. The third one we typically
- 18:01
see is collecting data. This is where I
- 18:04
think 50 to 60% of AI agents mess up
- 18:06
pretty badly. Uh and like we like to
- 18:10
think of it as
- 18:12
a UX problem. Uh but just for voice. So
- 18:16
think data models uh and not a
- 18:19
transcript coming into an LLM and and
- 18:21
trying to figure out what the transcript
- 18:22
said. So let's take some inspiration
- 18:24
from uh I'm assuming most of us are
- 18:28
developers here um you know take
- 18:29
inspiration from Python's data classes
- 18:31
pantic zod from Typescript or form
- 18:35
fields in the UI right like if you start
- 18:37
thinking of it from that problem
- 18:39
statement we have seen accuracy grow up
- 18:41
from grow from 30% to like 95% from a
- 18:45
data collection standpoint when you
- 18:47
start thinking in that manner. So like
- 18:49
decide your shape before you ask, right?
- 18:52
Like instead of keeping it open-ended,
- 18:54
can you keep it constrained? So can can
- 18:57
a phone number be a phone number type
- 18:59
field? The moment you do that, right,
- 19:01
you know like how many digits it needs
- 19:04
to have. You can do validation on on top
- 19:06
of that, right? And then what sort of
- 19:08
allowed values can even be there. So in
- 19:11
the previous example we saw if an E
- 19:13
comes in in middle of a phone number and
- 19:15
you know it's a phone number you
- 19:17
instantly know like either you smart
- 19:19
guess that to three and confirm that
- 19:20
with a user or you know that's an error
- 19:23
and then you validated that and asked
- 19:24
the user to repeat again right so so
- 19:27
that's I think one of the common
- 19:28
patterns we've seen here from from a a
- 19:32
collection pattern name I think is the
- 19:34
is the interesting one I've just picked
- 19:36
a you know a a hard to pronounce name
- 19:40
like There's no way a human is going to
- 19:41
get this right and and no way a
- 19:43
transcription engine will get this
- 19:44
right. How many ever times you do this
- 19:46
right? So the moment you start thinking
- 19:48
of this as fields and then have rules
- 19:50
and then confirmation mechanisms on on
- 19:53
spelling this uh you know sort of uh
- 19:55
letter by letter only then you kind of
- 19:58
get it right otherwise it's going to
- 19:59
mess up pretty badly in terms of how you
- 20:00
collect this on a voice call and and
- 20:03
that's just an example of you know what
- 20:06
u I'm talking about in terms of the the
- 20:08
data collection piece of it.
- 20:11
Another place where it goes badly
- 20:14
dramatically is relative uh values. Date
- 20:17
being one of the examples. If somebody
- 20:19
says next week uh Wednesday 8, 8 could
- 20:23
mean 8:00 a.m. 8:00 p.m. and then
- 20:25
figuring out what that date actually is.
- 20:27
Again, now becomes a very constrained
- 20:29
problem. If you knew this was a datetime
- 20:31
field and I I'm collecting a datetime
- 20:33
field and then you take the current date
- 20:35
and then figure out what this value
- 20:36
would be bases that, right? So, so
- 20:38
that's how you want to make sure like uh
- 20:40
you do this with a combination of the
- 20:42
LLM with the tool calling and the tool
- 20:44
calling is doing a lot of this heavy
- 20:45
lifting for you from a from a field
- 20:48
standpoint.
- 20:51
Yeah. And then you make you you run like
- 20:53
this from a unit test perspective. So
- 20:56
all of your u evals need to start
- 20:59
treating these fields as unit tests. And
- 21:02
as long as your unit tests uh sort of
- 21:05
validate and pass, you know, your agent
- 21:07
is going to be uh sort of reliable and
- 21:09
repeatable. You don't, you know, run uh
- 21:11
hundreds of end to end agent test cases
- 21:13
just to find out, you know, one field
- 21:15
collection is broken. You do your eval
- 21:18
at a field level and a unit test uh
- 21:20
level.
- 21:25
And then yeah, like I said, I think u
- 21:27
you this this mindset makes everything
- 21:29
more structured instead of hoping I'll
- 21:32
put a ton of prompt, keep changing, you
- 21:34
know, the prompt by a few uh characters
- 21:37
every time and somehow my prompt
- 21:38
engineering is going to make LLM much
- 21:41
more instruction tuned and sort of
- 21:43
magically start following some of these
- 21:44
things. So in fact u like I said right
- 21:47
like we have seen us get to 95 97%
- 21:50
accuracy without having to fine-tune a
- 21:52
model right and then and the trick is
- 21:54
basically like just breaking down your
- 21:56
context of what the agent is doing at
- 21:58
that point with specific u states of
- 22:01
what the agent is going through.
- 22:05
All right u I'm just going to quickly u
- 22:07
skip through this from a
- 22:10
time standpoint. I just see I got three
- 22:12
more minutes. Um hopefully that's a bug
- 22:15
but but we'll leave it at that. Okay. Um
- 22:19
so so this is the fourth area where we
- 22:21
see issues coming in. Most folks take
- 22:24
the LLM output and then we send it to a
- 22:27
TTS. Obviously I think there are a lot
- 22:29
of good TTS's in the market that take
- 22:30
care of a lot of heavy lifting but a lot
- 22:33
of times it it messes up. Uh what we
- 22:36
recommend and what we've seen is you
- 22:38
usually want to have a normalization
- 22:40
layer between your LLM and what is fed
- 22:43
to a TTS. You don't send your LLM output
- 22:47
directly to a TTS, right? And and we'll
- 22:49
just walk through some examples. The
- 22:52
basics which is strip emojis uh markdown
- 22:56
before before any synthesis into the
- 22:58
TTS. Most orchestration pipelines do
- 23:00
this like you know a live kit or a
- 23:02
pipecat would do that for you if you
- 23:03
just set a few flags. So I but but just
- 23:06
make sure if you're not using them or
- 23:08
buildings from scratch that you've set
- 23:10
this explicitly because you don't want
- 23:11
an emoji showing up on on on something
- 23:14
read out or you know markdown showing up
- 23:16
there.
- 23:18
Okay. I think I think some more common
- 23:19
ones uh custom uh dictionaries most TTS
- 23:23
engines provide this to you like how to
- 23:25
pronounce custom words whether it's you
- 23:28
know proper nouns brands uh acronyms and
- 23:32
so on and so forth. So set those in uh
- 23:34
when you go from your LLM to your TTS
- 23:36
output because if you don't, you're
- 23:37
going to mess that up. And I I'll I'll
- 23:39
show you an example of like how we test
- 23:40
that. Uh the the other one is like most
- 23:44
engines also give you speed. So if you
- 23:46
know you're pronouncing an entity, slow
- 23:48
down. Have your agent slow down. So at
- 23:51
point 8x or 7x so that it it's able to
- 23:54
like inunciate on that specific entity
- 23:57
and and doesn't mess up how it's
- 23:59
pronouncing an email or a phone number
- 24:01
or a name letter by letter
- 24:05
and yeah just normalize all the messy
- 24:08
stuff right like emails currency dates
- 24:10
don't leave it to the TTS to do it uh
- 24:13
most of them do it but don't leave it to
- 24:15
the TTS to do it like build your
- 24:17
normalization layer at your end so that
- 24:20
tomorrow you think you need to switch
- 24:21
TTS or you know for whatever reason the
- 24:24
first one's down and you want to use
- 24:25
another TTS you're able to sort of not
- 24:28
rely natively on the TTS's engine but
- 24:31
you are building this in-house uh for
- 24:34
for this to be managed
- 24:36
and then yeah u I think I don't have my
- 24:39
batch here but I I don't have my last
- 24:41
name on that so my first test is if it
- 24:43
cannot pronounce my last name or my
- 24:45
company's name it's already dropping the
- 24:47
ball so my last name is uh Balas
- 24:49
Subramanion and if you cannot pronounce
- 24:51
that using a voice AI agent uh like
- 24:55
that's a check for me. I I know like uh
- 24:58
you know the agent will mess up a lot of
- 25:00
words that uh you know need to be
- 25:03
spelled out day by day. The second one
- 25:06
is our company name Po. So a lot of
- 25:08
engines pronounce pronounce it pivo or
- 25:11
uh pleo and and so on and so forth. But
- 25:13
but I think specifically being able to
- 25:15
control this in your pipeline is super
- 25:17
critical. And then if you're building a
- 25:20
if you're building a customerf facing
- 25:22
product then then um you know sort of
- 25:25
give this option to your customers. All
- 25:26
right I'm just going to skim through the
- 25:28
the the last two slides. U I'm I'm
- 25:31
running badly over time. Uturn
- 25:33
detection. I think this is it own
- 25:35
separate topic but I'm just going to
- 25:36
quickly pull up all the points so you
- 25:38
guys can skim through that and if if you
- 25:40
need a chat u after this we can we can
- 25:43
talk about this. Right. Uh
- 25:47
I'm just going to leave that for like
- 25:48
five seconds and then and then we can
- 25:50
chat about this offline. I'm quite over
- 25:52
time. And then the the the last one is
- 25:55
uh bargin and and back channeling. I
- 25:57
think there's a lot of talk around
- 25:58
speech to speech models that do some of
- 26:00
this, but we've been able to see how we
- 26:02
could do all of this in speech to speech
- 26:04
pipelines. You really don't need a
- 26:05
speech to speech model to do all of this
- 26:07
up. Uh again, I'll just I just put put
- 26:09
this up on the slide and and sort of
- 26:12
close at that. Um
- 26:15
all right I don't think we have time for
- 26:17
questions we can take them offline if
- 26:18
you have any time but uh hopefully this
- 26:20
was helpful and gave you some insights
- 26:22
on uh what we are seeing in productions
- 26:24
uh with billions of calls at scale. All
- 26:26
right thanks