I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI
Read the talk
1 Trillion Phone Calls/yr, 10% Error Rate: The Crisis in Voice AI
Sumanyu Sharma explains how convincing speech can conceal failed actions, why shared prompts spread mistakes, and how replay tests, cross-call analysis, live experiments, and red teaming fit into a continuous reliability loop.
From a talk by Sumanyu Sharma
At a glance
Ideas worth remembering
Evaluate action-taking agents against the actual workflow outcome. A convincing booking claim can coexist with a missing appointment or a skipped prerequisite.
Prioritize frequency and severity together. Shared prompts and architectures can spread one defect across many calls; systematic, high-impact failures deserve P0 attention.
Start with manual listening, scale known checks, and add cross-call analysis. More scored conversations do not automatically reveal problems outside the rubric.
Replay real failures repeatedly, then vary the interaction while preserving intent. Check regressions and use live A/B tests for hypotheses about actual human responses.
Ordinary reliability tests and adversarial tests address different risks. Sharma recommends monitoring callers as well as agents and ongoing red teaming where failure costs are high.
From listening to incidents to taking actions
Sumanyu Sharma, founder and CEO of Hamming, begins with an unusual reference point for voice-agent safety. At Citizen, his team listened to thousands of hours of police radio and sent millions of alerts across San Francisco, New York, LA, Chicago, Baltimore, and other cities. The incidents ranged from grim reports to someone stealing bags of ice cream or taking a ladder while a man was still on the side of a house. Interpreting audio at scale already meant dealing with consequences outside the recording.
Voice agents make him nervous because they are moving from demonstrations into production. Since he began working on voice reliability in early 2024, infrastructure and orchestration have improved enough to make a convincing experience much faster to build. His informal description—perhaps 60% good in a short time—captures the gap between getting started and handling the long tail.
Teams are experimenting with hybrid architectures that combine voice-to-voice capabilities and cascading stacks, seeking reliability while keeping latency low. The talk does not specify how those components divide the work. The consequential change is clearer: connections to calendars, customer relationship management systems, electronic health records, and reservation systems let agents take actions. A good conversation now needs to produce the right change in another system.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A booking confirmation can coexist with an empty schedule
The first failure example concerns a customer seeking trade-in information. Confusing answers leave the customer unsure whether they are talking to AI or a person, and the business needs to call back to repair the situation. Natural speech has made the interaction plausible without making its information dependable. The cost includes the original confusion and the human work needed to restore trust.
The appointment example makes the distinction concrete. Sharma believed a voice agent had booked a physician appointment. He arrived, the front desk found that he was absent from the schedule, and he was turned away. He lost two hours. The observable failure was a mismatch between the outcome he understood from the call and the appointment record; the account does not establish which internal step failed. 3:37
Changing the patient or purpose changes the severity without changing the defect. A missing booking for a parent or grandparent, especially for a procedure rather than a checkup, could cost much more than wasted time. This is why conversational confidence is an inadequate success criterion: the caller acts on the promise, while the downstream system determines whether the service actually exists.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Volume multiplies errors; shared changes multiply exposure
Sharma contrasts what he describes as declining crime with growing voice-agent deployment. He estimates at least a trillion phone calls each year and forecasts that conversational agents will handle a majority within five years. The arithmetic is conditional: one trillion interactions at a 1% error rate would produce 10 billion incidents annually. That calculation explains the stakes of volume; it does not establish the adoption forecast.
Across the 10,000 agents Hamming monitors, Sharma reports an error rate closer to 10%. The recording does not define the error denominator, sampling method, or measurement window, so this is an observation about Hamming’s monitored population rather than an industry-wide rate. Its examples span several mechanisms:
- Skipped prerequisites: An agent says it found the right policy while omitting eligibility or verification steps.
- Unauthorized actions: An agent applies a discount it was not supposed to offer.
- Understanding and information errors: An agent mishears the caller or supplies incorrect information.
- Uncompleted actions: An agent claims it booked an appointment that does not exist.
Frequency alone cannot rank these failures. Repetition and misunderstandings can be annoying; mishandling a hypothetical drive-through order for a vegan burger from someone with peanut allergies can create a safety risk. The relevant question is what the caller needs the system to preserve, and what happens when that requirement is lost.
The crime comparison then turns from frequency to topology. A local robbery or vehicle theft affects the people involved in that event. A voice deployment shares prompts and architecture across calls. One change can therefore alter the experience of millions of downstream users. Centralized control makes an improvement easy to distribute, but it also distributes a defect. Making incidents visible becomes part of controlling that reach.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Build a loop that discovers failures as well as scores them
The proposed response is a debugging loop borrowed from friends who worked on growth at Facebook: identify problems, prioritize by frequency and severity, understand a fix, execute it, check improvement and regressions, and keep monitoring production. A change only earns its place after the team examines what it improved and what else it disturbed.
How does production experience become the next test? The loop below makes the feedback path visible. Monitoring feeds new problems into investigation; verification checks a candidate change before the team resumes watching its behavior. The return path matters because a successful fix does not end the possibility of failure.
Monitoring has two separate ambitions: cover more conversations and discover behavior beyond the problems already named. Latency, interruptions, and automatic speech recognition issues are examples of known problems a team can track over time. Emerging patterns may become apparent only when many conversations are considered together.
The methods contribute different kinds of understanding:
- Manual listening: Start with individual calls to build context and intuition about what actually happens. Sharma explicitly recommends keeping this step, even though it cannot scale to all traffic.
- Rubrics and automated scoring: Turn observations into five or ten checks for greetings, closing, validation, core logic, and other known requirements. Evaluation tools can then apply metrics, model-based judging, and deterministic or stochastic scoring to more calls.
- Cross-conversation analysis: Look for recurring behavior across calls that the existing rubric may miss. Increasing scoring coverage does not automatically expand the set of problems the scores can recognize.
Hamming invests heavily in cross-conversation analysis, as do some of the strongest teams it works with. The talk does not describe a discovery algorithm. Its practical distinction is the unit of analysis: a call can receive acceptable scores while a collection of calls reveals a pattern worth investigating.
Find failures in the conversational experience.
Executing a fix sits inside a larger loop. Verification checks both improvement and regressions; production monitoring supplies the next problems.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Prioritize recurrence and consequence together
An isolated, low-impact issue warrants less attention than a systematic, high-impact failure. But systematic annoyances still deserve work: repetition becomes unpleasant at scale and can matter when comparing competing agents. The P0 target is a recurring failure with serious consequences. Sharma’s example is a financial-services agent meant to freeze credit cards that fails to perform the freeze. It defeats the purpose of the workflow even if the conversation sounds fine.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Turn the missing appointment into a family of tests
“Please fix my agent” is the easy part of the loop. Attempting a change takes less effort than establishing that it works. The missing appointment returns here as a concrete test case: take the failed conversation and replay it five, ten, twenty, or fifty times. Repeated runs estimate how consistently the changed system passes that case, rather than accepting one favorable conversation. 9:56
The test needs to preserve the original failure’s meaning: an apparent booking must correspond to a scheduled appointment. Then broaden the interaction while keeping that intent. Change wording, conversational patterns, accents, and style; add another intent to the call. The observable improvement sought is that the booking task succeeds across related interactions, without introducing failures elsewhere. Sharma proposes this verification procedure rather than demonstrating a repaired version of his appointment call.
What does varying the call add beyond replaying it? The diagram separates consistency on one recorded failure from coverage of different ways to express the same goal. Both feed verification, but they answer different questions: whether the original case repeatedly passes, and whether the improvement survives ordinary variation.
Some questions require actual recipients. For outbound calls, Sharma considers the first five seconds especially important: vocal quality and opening words shape the human response. His recommendation is live A/B testing for those hypotheses because simulations alone cannot supply the relevant result. Synthetic tests exercise controlled interactions; the live experiment observes how people respond to the choices being compared.
Verification therefore has several jobs: reproduce the failure, test variations around its intent, assess human responses where necessary, and check regressions. Continued production monitoring closes the loop by revealing what the pre-deployment checks did not anticipate.
The caller believes a booking exists, but it is absent from the schedule.
Exact replay probes consistency. Intent-preserving variations probe whether the change works across different interactions. Neither removes the need for regression checks.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
When the caller wants the agent to fail
Until this point, the caller earnestly wants a problem solved. Adversarial callers change the task: they may try to induce disclosure of protected health information or personally identifiable information. Sharma raises increasingly capable automated callers as a threat scenario. The concrete concern is persuasion aimed at an agent or a human who can access sensitive information.
The pressure to ship useful agents increases what an attacker can reach. More data gives an agent more information to reveal; more tools give it more actions to misuse. Fast deployment compounds that exposure. Human-sounding voices also create a separate risk for people who trust a persuasive caller. Capability and naturalness improve legitimate service while giving adversarial interactions more to work with.
Hamming released its red-teaming product in April to test this concern. Sharma reports that the team can break approximately one in five agents across its tests in financial services, healthcare, consumer settings, and elsewhere. Reported outcomes include bypassing verification and using prompt injection to obtain data the testers should not receive. The recording does not define a broken agent or give the sampling and attack procedure, so the fraction describes those tests rather than general attack prevalence. 13:01
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Test before launch, watch both sides, keep attacking
The defensive program has complementary parts:
- Pre-deployment testing: Exercise the agent before ordinary callers encounter it. Text-to-text and voice-to-voice testing are both options; Sharma acknowledges tradeoffs without detailing them here.
- Production monitoring: Combine per-call scoring, manual listening, and cross-call analysis. Watch what the agent says and does, and also whether callers are trying to induce prohibited behavior.
- Continuous red teaming: Sharma recommends 24/7 adversarial testing, especially when bad interactions can be costly. This adds deliberate attempts to exploit the agent alongside observation of real traffic.
Watching both sides matters because the same undesirable agent behavior can arise during legitimate use or after deliberate manipulation. Monitoring the caller’s behavior helps a team recognize that difference. Continuous red teaming then asks whether someone intentionally searching for a weakness can find a path that ordinary-call testing misses.
The ending preserves the positive case: voice agents could make service feel more human than clunky interactive voice response trees, chatbots, or waiting on hold. Sharma expects more incidents as deployment grows and attackers exploit vulnerabilities. That is a forecast, grounded in his concern about exposure rather than a measured future outcome. His missing appointment supplies the personal stake: people will act on these systems’ promises.
The closing invitation is to work on this problem and examine deployed agents’ architecture and evaluations. Both deserve attention: architecture determines what the agent can do, while evaluations determine which failures the team can recognize. Better calls remain the goal; the reliability work follows the agent from its first tests through everyday use and deliberate attack.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
A practical implementation of the talk’s reliability loop. The supplied product page explains scenario generation, outcome checks, adversarial probes, and converting production failures into regression tests; its example dashboards are labeled illustrative.
Related talks
- 200 Million Patient Interactions Later: What the Generic Voice Stack Misses
A related healthcare voice talk for readers interested in the setting behind the appointment example.
- AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
A related session focused on AI red teaming, the defensive practice introduced in the technical ending.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
My name is Suman Yu and I'm the founder
- 0:15
and CEO of Hamming. And before working
- 0:19
on voice agent reliability and safety, I
- 0:22
worked at a company called Citizen
- 0:25
out of New York. Anybody here use
- 0:26
Citizen app? Awesome. Thank you. Uh, and
- 0:31
at Citizen,
- 0:33
we listened to crime, thousands of hours
- 0:36
of police radio station data and sent
- 0:40
millions of alerts to users in San
- 0:42
Francisco,
- 0:44
New York, LA, uh, Chicago, Baltimore,
- 0:48
and so on.
- 0:50
Some obviously gory and pretty sad. Uh,
- 0:54
but others more funny like a person
- 0:56
stealing bags of ice cream from Safeway
- 1:00
or report of a man hanging off the side
- 1:02
of the house after a woman stole his
- 1:04
ladder.
- 1:07
If I actually take a look at the citizen
- 1:09
app right now for those who are
- 1:10
customers or users, I can see that there
- 1:13
is a man yelling at person. There's
- 1:16
indecent exposure. This is real. This is
- 1:18
real time. This is, you know,
- 1:21
couple hours ago. These are real-time
- 1:23
alerts that we're sending.
- 1:27
Now, voice agents scare me more because
- 1:29
they're finally graduating from demos
- 1:31
and PC's to production. We should be
- 1:33
super excited, but I'm nervous. I'm
- 1:35
personally nervous. Uh, they're talking
- 1:37
to users at a scale that would make Gary
- 1:39
Tan and Polygram proud.
- 1:43
When I got started in voice agent
- 1:45
reliability in early 2024, voice was
- 1:48
just starting to work. It was not quite
- 1:50
good yet, but it was just starting to
- 1:52
work. You would have to pay me a lot of
- 1:54
money for me to stop using, you know,
- 1:55
Aqua voice, Super Whisper, uh, Whisper
- 1:58
Flow, and so on. These products are just
- 2:00
getting super, super good. And a big
- 2:02
reason is because the underlying
- 2:03
infrastructure is getting better, and
- 2:05
the orchestration layer is getting
- 2:06
meaningfully better. It's getting much
- 2:09
faster to build products and voice
- 2:10
experiences that maybe are 60% good in a
- 2:14
pretty short period of time, but the
- 2:16
long tail is still Hey, Gorov. the long
- 2:18
tail is still uh wise away.
- 2:21
I think speech speech models are getting
- 2:23
better. Um teams are experiment
- 2:25
experimenting with hybrid architectures
- 2:28
of combining more voicetovoice
- 2:32
modalities and also cascading stacks to
- 2:34
make the experience reliable but still
- 2:36
pretty low latency.
- 2:39
Things are obviously getting better.
- 2:41
Agents are being connected to calendars,
- 2:42
CRM, HRs, reservation systems, and so
- 2:45
on. Voice agents can now take actions.
- 2:50
However, reliability is still the number
- 2:52
one problem holding back most voice
- 2:54
agent deployments at scale. This is
- 2:55
still the number one problem. This is an
- 2:58
example I found on Twitter pretty
- 2:59
randomly, you know, two weeks ago and a
- 3:02
person is trying to get information for
- 3:04
a tradein and gets absolutely confused
- 3:06
with the information that they're
- 3:08
receiving. Alex now has to correct for
- 3:10
this loss of trust but trying to, you
- 3:13
know, call the person and see what see
- 3:15
what happened and fix the situation. Let
- 3:17
me see if audio works here.
- 3:20
>> Screwed up with another customer. We're
- 3:22
getting it fixed, but I got to call him
- 3:24
and see if I can work it out.
- 3:25
>> I'm like, dude, half the time I'm like,
- 3:27
I don't know if I'm talking to AI. I
- 3:28
don't know if I'm talking to a person.
- 3:29
It was just confusing, but we got there.
- 3:32
>> It probably is AI and human. And
- 3:37
>> so I think voices sound very confident.
- 3:39
They sound very natural, but the
- 3:40
information provided is often, you know,
- 3:42
not correct. That's the biggest problem
- 3:44
here.
- 3:46
This example is more personal. I had
- 3:48
booked an appointment with a physician a
- 3:49
couple weeks ago or I thought I did. I
- 3:52
showed up to the appointment and turns
- 3:54
out I was not actually on the schedule.
- 3:57
So the front desk, you me turned me
- 3:58
away. I wasted 2 hours. For me, this was
- 4:01
a waste of time. But what if this was
- 4:03
actually your parent?
- 4:06
What if this was your grandparent?
- 4:08
What if this appointment was for a
- 4:10
procedure instead of a regular checkup?
- 4:13
The costs for these different
- 4:15
permutations of the same failure mode
- 4:16
can actually be super super high.
- 4:20
Now, let's compare crime to voice
- 4:22
agents. Um, I think observation number
- 4:24
one is crime is actually decreasing over
- 4:26
time. This is a good thing and I hope it
- 4:30
crosses the x- axis at some point you
- 4:32
know in the future.
- 4:35
Voice on the other hand is generally
- 4:37
taking off right we're seeing a pretty
- 4:38
fast takeoff of voice agents being
- 4:40
deployed in production. There's at least
- 4:42
a trillion calls that are done every
- 4:44
single year and majority of these will
- 4:46
be done by conversational voice agents
- 4:48
over the next you know five years. If
- 4:50
you assume a 1% error rate that is still
- 4:53
10 billion incidents per year. That's a
- 4:55
lot.
- 4:58
In practice, we currently monitor 10,000
- 5:00
agents and the error rate is closer to
- 5:03
10% in practice. These range from agents
- 5:06
saying they found the right policy when
- 5:08
they actually skipped the eligibility or
- 5:10
verification steps or applying discounts
- 5:12
when they were not really supposed to,
- 5:14
misharing what the person said,
- 5:16
providing incorrect information, or
- 5:18
claiming they booked an appointment when
- 5:19
they actually did not, just like it
- 5:21
happened for me.
- 5:23
Now, not every single call has an
- 5:25
equally, you know, bad cost. Uh, some
- 5:29
range, you know, in the crime land, some
- 5:30
range from trash fires, which are kind
- 5:33
of funny, annoying, not really hurting
- 5:35
somebody. For a voice equivalent, that
- 5:38
would be annoyances like repetition, um,
- 5:40
or just sort of not quite understanding
- 5:42
what the user is saying. all the way to
- 5:44
safety risks like mass shootings or in
- 5:47
the voice agent equivalent, it would be
- 5:49
um a drive-thru that's deploying um
- 5:52
voice agents at scale like a Taco Bell
- 5:54
or McDonald's and a person orders a
- 5:57
vegan burger with peanut allergies.
- 5:59
If one of those two situations are not
- 6:02
handled correctly, that is definitely a
- 6:04
safety concern at scale.
- 6:07
The other big difference between crime
- 6:10
and and voice agent deployments is is
- 6:13
crime generally tends to be pretty hyper
- 6:16
local,
- 6:18
tends to be very decentralized, right?
- 6:19
Things like robbery or motor vehicle
- 6:22
theft or lararseny. They're impacting a
- 6:25
finite set of individuals that are
- 6:27
involved in that um situation.
- 6:32
On the other hand, voice agents are much
- 6:34
more centralized. a single prompt change
- 6:37
or an architecture change can have
- 6:38
pretty massive implications downstream
- 6:41
for all of the millions of you know
- 6:43
users that are um in the crossfire. So
- 6:46
the blast radius is is quite quite
- 6:48
massive. So the natural question is how
- 6:51
do you make these incidents much more
- 6:52
visible and obvious? That's the kind of
- 6:54
obvious question here.
- 6:57
I'll borrow a framework from a couple of
- 6:58
my friends who were OG growth folks at
- 7:01
Facebook. So step one is to identify
- 7:04
okay what are all the challenges and
- 7:05
problems that um exist in your
- 7:08
conversation experience. Step two is to
- 7:10
prioritize an impact size. There's a
- 7:12
frequency and severity analysis that's
- 7:14
pretty important. Step three is to
- 7:16
understand okay how do we actually fix
- 7:18
this? Step four execute. Step five okay
- 7:21
did my change actually work and did it
- 7:23
cause any regressions somewhere else.
- 7:25
And lastly we continue to monitor in
- 7:27
production.
- 7:30
On the y- axis, I think it's important
- 7:32
to highlight there are known problems
- 7:34
that already exist. Things like turnover
- 7:36
latency, interruptions, um maybe some
- 7:39
ASR problems you're, you know, aware of.
- 7:41
And these are known problems that exist
- 7:44
that the team should track over time. On
- 7:47
the other axis is actually emerging
- 7:49
behavior or patterns that are only
- 7:51
obvious across lots of conversations. Um
- 7:55
on the x- axis, you have coverage just
- 7:57
like insurance. Are you analyzing few
- 8:00
conversations? Are you analyzing many,
- 8:02
many conversations? Most teams will
- 8:04
typically start by listening to calls
- 8:06
manually.
- 8:08
And I think that's the best place to
- 8:09
start. I don't think you should skip
- 8:11
that step. There's a lot of depth and
- 8:13
insights to get by actually listening to
- 8:15
specific conversations and building that
- 8:17
texture that that comes from that
- 8:19
intuition. However, it's obviously not
- 8:21
scalable. So most teams end up having a
- 8:24
spreadsheet of I don't know five or 10
- 8:27
different rubrics around greetings,
- 8:29
closing, validation,
- 8:31
um, core logic and so on. To scale that
- 8:35
up even further, you then end up
- 8:36
investing in some eval product, right?
- 8:39
You might run some element as a judge
- 8:41
and compute classic metrics and also
- 8:44
more more deterministic and stoastic
- 8:47
scoring logic. Um, but there you're
- 8:49
still stuck with checking for
- 8:51
consistency of known problems, but
- 8:53
you're not really discovering novel
- 8:54
insights that are actually happening
- 8:56
across conversations. We're spending a
- 8:58
ton of time on performing cross
- 9:01
conversation analysis, not a pattern on
- 9:03
a single call, but across conversations.
- 9:05
And some of the best teams that we work
- 9:07
with are are doing the same.
- 9:10
Now, to prioritize an impact size, I
- 9:12
think there's problems that are one-off
- 9:14
that are low impact. I mean, who cares?
- 9:17
uh even low impact and systematic
- 9:19
problems in the crime world that would
- 9:21
be a trash fire in a voice aation world
- 9:23
it could be some repetitions the team is
- 9:25
experiencing they're still annoying at
- 9:27
scale and if you are doing a bake off
- 9:29
it's still worth solving for them I
- 9:31
would not ignore these class of problems
- 9:33
oneoff and high impact well hope it
- 9:35
doesn't chronic and I think systematic
- 9:38
and high impact are obviously the P 0
- 9:40
you know target areas um for the team to
- 9:42
solve an example of that
- 9:45
would be in a fins serve capacity
- 9:47
There's a voice agent that um helps
- 9:49
users freeze their credit cards. And if
- 9:51
it doesn't do that, well, that's a
- 9:53
massive fail.
- 9:56
All right. So, understand and execute.
- 9:58
I'm pretty sure everyone's doing this.
- 10:00
Please fix my agent. Uh I think fixing
- 10:02
or rather attempting to make a fix is
- 10:05
the simplest and the lowest effort
- 10:08
component of this debugging pipeline and
- 10:11
loop. Um the next step is all right, I
- 10:14
made a change to my system. How do I
- 10:16
actually know this thing works um for
- 10:19
real? A great way that's naive is to
- 10:23
take a real call, for example, in my
- 10:25
case, I booked an appointment and it
- 10:27
didn't get scheduled and replay that
- 10:29
exact conversation and run that maybe 5,
- 10:32
10, 20, 50 times and see, okay, what is
- 10:34
my probability of passing this type of
- 10:37
issue? A better way is to keep the same
- 10:40
intent but change the wordings, change
- 10:44
the patterns, change the accents, change
- 10:46
the style, add one more intent to the
- 10:48
mix. And that gives teams much more, you
- 10:51
know, better coverage to feel confident
- 10:53
that yes, I actually made a change and
- 10:56
my changes are net positive instead of
- 10:58
net negative.
- 11:01
There are certain fixes and I guess
- 11:05
hypothesis that are very difficult to
- 11:07
test in a pre-eployment synthetic
- 11:09
setting. And so AB testing ends up
- 11:11
being, you know, pretty pretty critical
- 11:13
for those circumstances. For example, if
- 11:15
you have an outbound agent, the first 5
- 11:17
seconds of a conversation tends to be
- 11:19
the most important. And so the vocal
- 11:22
quality um and the specific words you
- 11:25
end up using, they matter the most. And
- 11:27
so AB testing that is the only way in in
- 11:29
kind of real life setting to to get
- 11:32
results. You can't really do it through
- 11:34
simulations alone.
- 11:37
And so there we have the loop. Identify,
- 11:39
prioritize, impact size, understand the
- 11:42
fix, execute, check, make sure it didn't
- 11:45
break anything, and then continue
- 11:47
monitoring.
- 11:51
So I think making voice agents useful is
- 11:54
already hard as it is. even when dealing
- 11:58
with earnest users on the other line,
- 12:01
right? These are people who who just
- 12:02
want their problem solved. They're not
- 12:03
trying to mess with you. These are like
- 12:04
legit normal people.
- 12:08
Now, what happens when mythos learns how
- 12:10
to dial?
- 12:13
So, if it can extract trade secrets and
- 12:16
uh you know, from the NSA, it can
- 12:18
certainly, you know, seduce you into
- 12:22
revealing PHI and PII data as well.
- 12:26
And I think both voice agents and humans
- 12:29
will [clears throat] be targeted here.
- 12:31
Voice agents because there's a pressure
- 12:33
to make these more capable. Give them
- 12:37
access to more data. Give them access to
- 12:39
more tools.
- 12:41
Deploy them quickly.
- 12:45
The more the capability, the bigger the
- 12:46
surface area. This is this is pretty
- 12:48
pretty common sense. And the more the
- 12:50
voice agents become natural and human
- 12:54
sounding, the more humans will be
- 12:56
tricked along the way as well for those
- 12:58
who are weaponizing.
- 13:01
Uh we ship a we shipped a red tipping
- 13:03
product um back in April just to test
- 13:05
out this hypothesis for how many agents
- 13:07
can we actually break from a adversarial
- 13:10
capacity and we can probably break one
- 13:12
in five agents at this point. We've
- 13:14
tested this across financial services,
- 13:16
healthcare, um consumer and so on. We've
- 13:19
bypassed verification. Uh we've
- 13:22
definitely had agents, you know, we've
- 13:24
been able to promject uh several agents
- 13:27
and and gotten data we should not have.
- 13:29
So this is not theoretical. This is
- 13:31
actually a real a real concern.
- 13:37
I think the only real defense against
- 13:39
the dark arts is
- 13:41
step one to invest deeply in
- 13:43
pre-eployment testing. This could be
- 13:45
textto text. This could be voice to
- 13:48
voice. There's pros and cons to both.
- 13:50
Happy to chat offline if folks are
- 13:51
interested. And this is just making sure
- 13:54
you're not self-owning, you know, when
- 13:56
you're talking to real people who just
- 13:57
want to get their problem solved. Step
- 13:59
two is to have a great monitoring system
- 14:02
of all kinds. And I've highlighted, you
- 14:04
know, different flavors of monitoring
- 14:06
per call scoring, manual kind of evals,
- 14:10
you know, listening to conversations and
- 14:11
cross call analysis. And this is helpful
- 14:14
both for monitoring what the agent is
- 14:16
saying and behaving and how it's
- 14:17
actually doing, but also the users. Are
- 14:19
the users being adversarial? Are they
- 14:21
being annoying? Are are they trying to
- 14:22
trick the agent into doing things it's
- 14:24
not supposed to be doing?
- 14:26
And I think our new recommendation now
- 14:28
is to run 24/7 red teaming um for your
- 14:31
agents, especially if you believe the
- 14:34
cost of bad interactions can be can be
- 14:36
rather large.
- 14:38
So, I think voice agents um have this
- 14:40
awesome potential of of making the world
- 14:43
feel much more human compared to
- 14:48
interacting with clunky IVR trees or
- 14:50
chat bots or worse um being stuck on a
- 14:53
on a hold.
- 14:55
And when we think about crime, we often
- 14:57
think of crime happening to somebody
- 14:59
else. You know, crime does not happen to
- 15:01
you typically with voice agents,
- 15:04
especially bad actors. As these agents
- 15:07
are deployed and as bad actors start to
- 15:09
exploit a lot of the vulnerabilities,
- 15:11
the number of incidents is about to kind
- 15:13
of go way way up. And so the reason I
- 15:16
fear voice agents more than crime is
- 15:18
that one of these incidents is going to
- 15:20
impact you. It already did for me.
- 15:24
Awesome. So it's time for me to shill.
- 15:25
Well, we burn a lot of tokens. If you
- 15:28
are interested in working in this space,
- 15:31
please come and talk to us. And if you
- 15:33
are deploying voice agents and want to
- 15:35
validate whether your architectured or
- 15:39
your eval are set up correctly, please
- 15:40
come and talk to us. We'll be outside.
- 15:42
And here's here's my number. Here's my
- 15:44
WhatsApp. Thanks everyone.
- 16:02
>> [music]