I Monitored Crime Audio. Voice Agents Scare Me More. — Sumanyu Sharma, Hamming AI

Read the talk

1 Trillion Phone Calls/yr, 10% Error Rate: The Crisis in Voice AI

Sumanyu Sharma explains how convincing speech can conceal failed actions, why shared prompts spread mistakes, and how replay tests, cross-call analysis, live experiments, and red teaming fit into a continuous reliability loop.

From a talk by Sumanyu Sharma

At a glance

Ideas worth remembering

  • Evaluate action-taking agents against the actual workflow outcome. A convincing booking claim can coexist with a missing appointment or a skipped prerequisite.

  • Prioritize frequency and severity together. Shared prompts and architectures can spread one defect across many calls; systematic, high-impact failures deserve P0 attention.

  • Start with manual listening, scale known checks, and add cross-call analysis. More scored conversations do not automatically reveal problems outside the rubric.

  • Replay real failures repeatedly, then vary the interaction while preserving intent. Check regressions and use live A/B tests for hypotheses about actual human responses.

  • Ordinary reliability tests and adversarial tests address different risks. Sharma recommends monitoring callers as well as agents and ongoing red teaming where failure costs are high.

From listening to incidents to taking actions

Sumanyu Sharma, founder and CEO of Hamming, begins with an unusual reference point for voice-agent safety. At Citizen, his team listened to thousands of hours of police radio and sent millions of alerts across San Francisco, New York, LA, Chicago, Baltimore, and other cities. The incidents ranged from grim reports to someone stealing bags of ice cream or taking a ladder while a man was still on the side of a house. Interpreting audio at scale already meant dealing with consequences outside the recording.

Recording frame at 71 seconds
Recording frame at 71 seconds

Voice agents make him nervous because they are moving from demonstrations into production. Since he began working on voice reliability in early 2024, infrastructure and orchestration have improved enough to make a convincing experience much faster to build. His informal description—perhaps 60% good in a short time—captures the gap between getting started and handling the long tail.

Teams are experimenting with hybrid architectures that combine voice-to-voice capabilities and cascading stacks, seeking reliability while keeping latency low. The talk does not specify how those components divide the work. The consequential change is clearer: connections to calendars, customer relationship management systems, electronic health records, and reservation systems let agents take actions. A good conversation now needs to produce the right change in another system.

0:120:29
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

A booking confirmation can coexist with an empty schedule

The first failure example concerns a customer seeking trade-in information. Confusing answers leave the customer unsure whether they are talking to AI or a person, and the business needs to call back to repair the situation. Natural speech has made the interaction plausible without making its information dependable. The cost includes the original confusion and the human work needed to restore trust.

The appointment example makes the distinction concrete. Sharma believed a voice agent had booked a physician appointment. He arrived, the front desk found that he was absent from the schedule, and he was turned away. He lost two hours. The observable failure was a mismatch between the outcome he understood from the call and the appointment record; the account does not establish which internal step failed. 3:37

Changing the patient or purpose changes the severity without changing the defect. A missing booking for a parent or grandparent, especially for a procedure rather than a checkup, could cost much more than wasted time. This is why conversational confidence is an inadequate success criterion: the caller acts on the promise, while the downstream system determines whether the service actually exists.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:50 · section reference included

Volume multiplies errors; shared changes multiply exposure

Sharma contrasts what he describes as declining crime with growing voice-agent deployment. He estimates at least a trillion phone calls each year and forecasts that conversational agents will handle a majority within five years. The arithmetic is conditional: one trillion interactions at a 1% error rate would produce 10 billion incidents annually. That calculation explains the stakes of volume; it does not establish the adoption forecast.

Recording frame at 357 seconds
Recording frame at 357 seconds

Across the 10,000 agents Hamming monitors, Sharma reports an error rate closer to 10%. The recording does not define the error denominator, sampling method, or measurement window, so this is an observation about Hamming’s monitored population rather than an industry-wide rate. Its examples span several mechanisms:

  • Skipped prerequisites: An agent says it found the right policy while omitting eligibility or verification steps.
  • Unauthorized actions: An agent applies a discount it was not supposed to offer.
  • Understanding and information errors: An agent mishears the caller or supplies incorrect information.
  • Uncompleted actions: An agent claims it booked an appointment that does not exist.

Frequency alone cannot rank these failures. Repetition and misunderstandings can be annoying; mishandling a hypothetical drive-through order for a vegan burger from someone with peanut allergies can create a safety risk. The relevant question is what the caller needs the system to preserve, and what happens when that requirement is lost.

The crime comparison then turns from frequency to topology. A local robbery or vehicle theft affects the people involved in that event. A voice deployment shares prompts and architecture across calls. One change can therefore alter the experience of millions of downstream users. Centralized control makes an improvement easy to distribute, but it also distributes a defect. Making incidents visible becomes part of controlling that reach.

4:204:35
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:20 · section reference included

Build a loop that discovers failures as well as scores them

The proposed response is a debugging loop borrowed from friends who worked on growth at Facebook: identify problems, prioritize by frequency and severity, understand a fix, execute it, check improvement and regressions, and keep monitoring production. A change only earns its place after the team examines what it improved and what else it disturbed.

How does production experience become the next test? The loop below makes the feedback path visible. Monitoring feeds new problems into investigation; verification checks a candidate change before the team resumes watching its behavior. The return path matters because a successful fix does not end the possibility of failure.

Monitoring has two separate ambitions: cover more conversations and discover behavior beyond the problems already named. Latency, interruptions, and automatic speech recognition issues are examples of known problems a team can track over time. Emerging patterns may become apparent only when many conversations are considered together.

The methods contribute different kinds of understanding:

  • Manual listening: Start with individual calls to build context and intuition about what actually happens. Sharma explicitly recommends keeping this step, even though it cannot scale to all traffic.
  • Rubrics and automated scoring: Turn observations into five or ten checks for greetings, closing, validation, core logic, and other known requirements. Evaluation tools can then apply metrics, model-based judging, and deterministic or stochastic scoring to more calls.
  • Cross-conversation analysis: Look for recurring behavior across calls that the existing rubric may miss. Increasing scoring coverage does not automatically expand the set of problems the scores can recognize.

Hamming invests heavily in cross-conversation analysis, as do some of the strongest teams it works with. The talk does not describe a discovery algorithm. Its practical distinction is the unit of analysis: a call can receive acceptable scores while a collection of calls reveals a pattern worth investigating.

How it fits togetherProduction failures feed the next reliability check

Find failures in the conversational experience.

Executing a fix sits inside a larger loop. Verification checks both improvement and regressions; production monitoring supplies the next problems.

6:327:02
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:02 · section reference included

Prioritize recurrence and consequence together

An isolated, low-impact issue warrants less attention than a systematic, high-impact failure. But systematic annoyances still deserve work: repetition becomes unpleasant at scale and can matter when comparing competing agents. The P0 target is a recurring failure with serious consequences. Sharma’s example is a financial-services agent meant to freeze credit cards that fails to perform the freeze. It defeats the purpose of the workflow even if the conversation sounds fine.

9:109:40
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:10 · section reference included

Turn the missing appointment into a family of tests

“Please fix my agent” is the easy part of the loop. Attempting a change takes less effort than establishing that it works. The missing appointment returns here as a concrete test case: take the failed conversation and replay it five, ten, twenty, or fifty times. Repeated runs estimate how consistently the changed system passes that case, rather than accepting one favorable conversation. 9:56

Recording frame at 632 seconds
Recording frame at 632 seconds

The test needs to preserve the original failure’s meaning: an apparent booking must correspond to a scheduled appointment. Then broaden the interaction while keeping that intent. Change wording, conversational patterns, accents, and style; add another intent to the call. The observable improvement sought is that the booking task succeeds across related interactions, without introducing failures elsewhere. Sharma proposes this verification procedure rather than demonstrating a repaired version of his appointment call.

What does varying the call add beyond replaying it? The diagram separates consistency on one recorded failure from coverage of different ways to express the same goal. Both feed verification, but they answer different questions: whether the original case repeatedly passes, and whether the improvement survives ordinary variation.

Some questions require actual recipients. For outbound calls, Sharma considers the first five seconds especially important: vocal quality and opening words shape the human response. His recommendation is live A/B testing for those hypotheses because simulations alone cannot supply the relevant result. Synthetic tests exercise controlled interactions; the live experiment observes how people respond to the choices being compared.

Verification therefore has several jobs: reproduce the failure, test variations around its intent, assess human responses where necessary, and check regressions. Continued production monitoring closes the loop by revealing what the pre-deployment checks did not anticipate.

How it fits togetherOne booking failure becomes two kinds of coverage

The caller believes a booking exists, but it is absent from the schedule.

Exact replay probes consistency. Intent-preserving variations probe whether the change works across different interactions. Neither removes the need for regression checks.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:56 · section reference included

When the caller wants the agent to fail

Until this point, the caller earnestly wants a problem solved. Adversarial callers change the task: they may try to induce disclosure of protected health information or personally identifiable information. Sharma raises increasingly capable automated callers as a threat scenario. The concrete concern is persuasion aimed at an agent or a human who can access sensitive information.

Recording frame at 767 seconds
Recording frame at 767 seconds

The pressure to ship useful agents increases what an attacker can reach. More data gives an agent more information to reveal; more tools give it more actions to misuse. Fast deployment compounds that exposure. Human-sounding voices also create a separate risk for people who trust a persuasive caller. Capability and naturalness improve legitimate service while giving adversarial interactions more to work with.

Hamming released its red-teaming product in April to test this concern. Sharma reports that the team can break approximately one in five agents across its tests in financial services, healthcare, consumer settings, and elsewhere. Reported outcomes include bypassing verification and using prompt injection to obtain data the testers should not receive. The recording does not define a broken agent or give the sampling and attack procedure, so the fraction describes those tests rather than general attack prevalence. 13:01

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:52 · section reference included

Test before launch, watch both sides, keep attacking

The defensive program has complementary parts:

  • Pre-deployment testing: Exercise the agent before ordinary callers encounter it. Text-to-text and voice-to-voice testing are both options; Sharma acknowledges tradeoffs without detailing them here.
  • Production monitoring: Combine per-call scoring, manual listening, and cross-call analysis. Watch what the agent says and does, and also whether callers are trying to induce prohibited behavior.
  • Continuous red teaming: Sharma recommends 24/7 adversarial testing, especially when bad interactions can be costly. This adds deliberate attempts to exploit the agent alongside observation of real traffic.
Recording frame at 852 seconds
Recording frame at 852 seconds

Watching both sides matters because the same undesirable agent behavior can arise during legitimate use or after deliberate manipulation. Monitoring the caller’s behavior helps a team recognize that difference. Continuous red teaming then asks whether someone intentionally searching for a weakness can find a path that ordinary-call testing misses.

The ending preserves the positive case: voice agents could make service feel more human than clunky interactive voice response trees, chatbots, or waiting on hold. Sharma expects more incidents as deployment grows and attackers exploit vulnerabilities. That is a forecast, grounded in his concern about exposure rather than a measured future outcome. His missing appointment supplies the personal stake: people will act on these systems’ promises.

The closing invitation is to work on this problem and examine deployed agents’ architecture and evaluations. Both deserve attention: architecture determines what the agent can do, while evaluations determine which failures the team can recognize. Better calls remain the goal; the reliability work follows the agent from its first tests through everyday use and deliberate attack.

13:3714:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:37 · section reference included

Resources

From the talk

  • A practical implementation of the talk’s reliability loop. The supplied product page explains scenario generation, outcome checks, adversarial probes, and converting production failures into regression tests; its example dashboards are labeled illustrative.

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    My name is Suman Yu and I'm the founder

  3. 0:15

    and CEO of Hamming. And before working

  4. 0:19

    on voice agent reliability and safety, I

  5. 0:22

    worked at a company called Citizen

  6. 0:25

    out of New York. Anybody here use

  7. 0:26

    Citizen app? Awesome. Thank you. Uh, and

  8. 0:31

    at Citizen,

  9. 0:33

    we listened to crime, thousands of hours

  10. 0:36

    of police radio station data and sent

  11. 0:40

    millions of alerts to users in San

  12. 0:42

    Francisco,

  13. 0:44

    New York, LA, uh, Chicago, Baltimore,

  14. 0:48

    and so on.

  15. 0:50

    Some obviously gory and pretty sad. Uh,

  16. 0:54

    but others more funny like a person

  17. 0:56

    stealing bags of ice cream from Safeway

  18. 1:00

    or report of a man hanging off the side

  19. 1:02

    of the house after a woman stole his

  20. 1:04

    ladder.

  21. 1:07

    If I actually take a look at the citizen

  22. 1:09

    app right now for those who are

  23. 1:10

    customers or users, I can see that there

  24. 1:13

    is a man yelling at person. There's

  25. 1:16

    indecent exposure. This is real. This is

  26. 1:18

    real time. This is, you know,

  27. 1:21

    couple hours ago. These are real-time

  28. 1:23

    alerts that we're sending.

  29. 1:27

    Now, voice agents scare me more because

  30. 1:29

    they're finally graduating from demos

  31. 1:31

    and PC's to production. We should be

  32. 1:33

    super excited, but I'm nervous. I'm

  33. 1:35

    personally nervous. Uh, they're talking

  34. 1:37

    to users at a scale that would make Gary

  35. 1:39

    Tan and Polygram proud.

  36. 1:43

    When I got started in voice agent

  37. 1:45

    reliability in early 2024, voice was

  38. 1:48

    just starting to work. It was not quite

  39. 1:50

    good yet, but it was just starting to

  40. 1:52

    work. You would have to pay me a lot of

  41. 1:54

    money for me to stop using, you know,

  42. 1:55

    Aqua voice, Super Whisper, uh, Whisper

  43. 1:58

    Flow, and so on. These products are just

  44. 2:00

    getting super, super good. And a big

  45. 2:02

    reason is because the underlying

  46. 2:03

    infrastructure is getting better, and

  47. 2:05

    the orchestration layer is getting

  48. 2:06

    meaningfully better. It's getting much

  49. 2:09

    faster to build products and voice

  50. 2:10

    experiences that maybe are 60% good in a

  51. 2:14

    pretty short period of time, but the

  52. 2:16

    long tail is still Hey, Gorov. the long

  53. 2:18

    tail is still uh wise away.

  54. 2:21

    I think speech speech models are getting

  55. 2:23

    better. Um teams are experiment

  56. 2:25

    experimenting with hybrid architectures

  57. 2:28

    of combining more voicetovoice

  58. 2:32

    modalities and also cascading stacks to

  59. 2:34

    make the experience reliable but still

  60. 2:36

    pretty low latency.

  61. 2:39

    Things are obviously getting better.

  62. 2:41

    Agents are being connected to calendars,

  63. 2:42

    CRM, HRs, reservation systems, and so

  64. 2:45

    on. Voice agents can now take actions.

  65. 2:50

    However, reliability is still the number

  66. 2:52

    one problem holding back most voice

  67. 2:54

    agent deployments at scale. This is

  68. 2:55

    still the number one problem. This is an

  69. 2:58

    example I found on Twitter pretty

  70. 2:59

    randomly, you know, two weeks ago and a

  71. 3:02

    person is trying to get information for

  72. 3:04

    a tradein and gets absolutely confused

  73. 3:06

    with the information that they're

  74. 3:08

    receiving. Alex now has to correct for

  75. 3:10

    this loss of trust but trying to, you

  76. 3:13

    know, call the person and see what see

  77. 3:15

    what happened and fix the situation. Let

  78. 3:17

    me see if audio works here.

  79. 3:20

    >> Screwed up with another customer. We're

  80. 3:22

    getting it fixed, but I got to call him

  81. 3:24

    and see if I can work it out.

  82. 3:25

    >> I'm like, dude, half the time I'm like,

  83. 3:27

    I don't know if I'm talking to AI. I

  84. 3:28

    don't know if I'm talking to a person.

  85. 3:29

    It was just confusing, but we got there.

  86. 3:32

    >> It probably is AI and human. And

  87. 3:37

    >> so I think voices sound very confident.

  88. 3:39

    They sound very natural, but the

  89. 3:40

    information provided is often, you know,

  90. 3:42

    not correct. That's the biggest problem

  91. 3:44

    here.

  92. 3:46

    This example is more personal. I had

  93. 3:48

    booked an appointment with a physician a

  94. 3:49

    couple weeks ago or I thought I did. I

  95. 3:52

    showed up to the appointment and turns

  96. 3:54

    out I was not actually on the schedule.

  97. 3:57

    So the front desk, you me turned me

  98. 3:58

    away. I wasted 2 hours. For me, this was

  99. 4:01

    a waste of time. But what if this was

  100. 4:03

    actually your parent?

  101. 4:06

    What if this was your grandparent?

  102. 4:08

    What if this appointment was for a

  103. 4:10

    procedure instead of a regular checkup?

  104. 4:13

    The costs for these different

  105. 4:15

    permutations of the same failure mode

  106. 4:16

    can actually be super super high.

  107. 4:20

    Now, let's compare crime to voice

  108. 4:22

    agents. Um, I think observation number

  109. 4:24

    one is crime is actually decreasing over

  110. 4:26

    time. This is a good thing and I hope it

  111. 4:30

    crosses the x- axis at some point you

  112. 4:32

    know in the future.

  113. 4:35

    Voice on the other hand is generally

  114. 4:37

    taking off right we're seeing a pretty

  115. 4:38

    fast takeoff of voice agents being

  116. 4:40

    deployed in production. There's at least

  117. 4:42

    a trillion calls that are done every

  118. 4:44

    single year and majority of these will

  119. 4:46

    be done by conversational voice agents

  120. 4:48

    over the next you know five years. If

  121. 4:50

    you assume a 1% error rate that is still

  122. 4:53

    10 billion incidents per year. That's a

  123. 4:55

    lot.

  124. 4:58

    In practice, we currently monitor 10,000

  125. 5:00

    agents and the error rate is closer to

  126. 5:03

    10% in practice. These range from agents

  127. 5:06

    saying they found the right policy when

  128. 5:08

    they actually skipped the eligibility or

  129. 5:10

    verification steps or applying discounts

  130. 5:12

    when they were not really supposed to,

  131. 5:14

    misharing what the person said,

  132. 5:16

    providing incorrect information, or

  133. 5:18

    claiming they booked an appointment when

  134. 5:19

    they actually did not, just like it

  135. 5:21

    happened for me.

  136. 5:23

    Now, not every single call has an

  137. 5:25

    equally, you know, bad cost. Uh, some

  138. 5:29

    range, you know, in the crime land, some

  139. 5:30

    range from trash fires, which are kind

  140. 5:33

    of funny, annoying, not really hurting

  141. 5:35

    somebody. For a voice equivalent, that

  142. 5:38

    would be annoyances like repetition, um,

  143. 5:40

    or just sort of not quite understanding

  144. 5:42

    what the user is saying. all the way to

  145. 5:44

    safety risks like mass shootings or in

  146. 5:47

    the voice agent equivalent, it would be

  147. 5:49

    um a drive-thru that's deploying um

  148. 5:52

    voice agents at scale like a Taco Bell

  149. 5:54

    or McDonald's and a person orders a

  150. 5:57

    vegan burger with peanut allergies.

  151. 5:59

    If one of those two situations are not

  152. 6:02

    handled correctly, that is definitely a

  153. 6:04

    safety concern at scale.

  154. 6:07

    The other big difference between crime

  155. 6:10

    and and voice agent deployments is is

  156. 6:13

    crime generally tends to be pretty hyper

  157. 6:16

    local,

  158. 6:18

    tends to be very decentralized, right?

  159. 6:19

    Things like robbery or motor vehicle

  160. 6:22

    theft or lararseny. They're impacting a

  161. 6:25

    finite set of individuals that are

  162. 6:27

    involved in that um situation.

  163. 6:32

    On the other hand, voice agents are much

  164. 6:34

    more centralized. a single prompt change

  165. 6:37

    or an architecture change can have

  166. 6:38

    pretty massive implications downstream

  167. 6:41

    for all of the millions of you know

  168. 6:43

    users that are um in the crossfire. So

  169. 6:46

    the blast radius is is quite quite

  170. 6:48

    massive. So the natural question is how

  171. 6:51

    do you make these incidents much more

  172. 6:52

    visible and obvious? That's the kind of

  173. 6:54

    obvious question here.

  174. 6:57

    I'll borrow a framework from a couple of

  175. 6:58

    my friends who were OG growth folks at

  176. 7:01

    Facebook. So step one is to identify

  177. 7:04

    okay what are all the challenges and

  178. 7:05

    problems that um exist in your

  179. 7:08

    conversation experience. Step two is to

  180. 7:10

    prioritize an impact size. There's a

  181. 7:12

    frequency and severity analysis that's

  182. 7:14

    pretty important. Step three is to

  183. 7:16

    understand okay how do we actually fix

  184. 7:18

    this? Step four execute. Step five okay

  185. 7:21

    did my change actually work and did it

  186. 7:23

    cause any regressions somewhere else.

  187. 7:25

    And lastly we continue to monitor in

  188. 7:27

    production.

  189. 7:30

    On the y- axis, I think it's important

  190. 7:32

    to highlight there are known problems

  191. 7:34

    that already exist. Things like turnover

  192. 7:36

    latency, interruptions, um maybe some

  193. 7:39

    ASR problems you're, you know, aware of.

  194. 7:41

    And these are known problems that exist

  195. 7:44

    that the team should track over time. On

  196. 7:47

    the other axis is actually emerging

  197. 7:49

    behavior or patterns that are only

  198. 7:51

    obvious across lots of conversations. Um

  199. 7:55

    on the x- axis, you have coverage just

  200. 7:57

    like insurance. Are you analyzing few

  201. 8:00

    conversations? Are you analyzing many,

  202. 8:02

    many conversations? Most teams will

  203. 8:04

    typically start by listening to calls

  204. 8:06

    manually.

  205. 8:08

    And I think that's the best place to

  206. 8:09

    start. I don't think you should skip

  207. 8:11

    that step. There's a lot of depth and

  208. 8:13

    insights to get by actually listening to

  209. 8:15

    specific conversations and building that

  210. 8:17

    texture that that comes from that

  211. 8:19

    intuition. However, it's obviously not

  212. 8:21

    scalable. So most teams end up having a

  213. 8:24

    spreadsheet of I don't know five or 10

  214. 8:27

    different rubrics around greetings,

  215. 8:29

    closing, validation,

  216. 8:31

    um, core logic and so on. To scale that

  217. 8:35

    up even further, you then end up

  218. 8:36

    investing in some eval product, right?

  219. 8:39

    You might run some element as a judge

  220. 8:41

    and compute classic metrics and also

  221. 8:44

    more more deterministic and stoastic

  222. 8:47

    scoring logic. Um, but there you're

  223. 8:49

    still stuck with checking for

  224. 8:51

    consistency of known problems, but

  225. 8:53

    you're not really discovering novel

  226. 8:54

    insights that are actually happening

  227. 8:56

    across conversations. We're spending a

  228. 8:58

    ton of time on performing cross

  229. 9:01

    conversation analysis, not a pattern on

  230. 9:03

    a single call, but across conversations.

  231. 9:05

    And some of the best teams that we work

  232. 9:07

    with are are doing the same.

  233. 9:10

    Now, to prioritize an impact size, I

  234. 9:12

    think there's problems that are one-off

  235. 9:14

    that are low impact. I mean, who cares?

  236. 9:17

    uh even low impact and systematic

  237. 9:19

    problems in the crime world that would

  238. 9:21

    be a trash fire in a voice aation world

  239. 9:23

    it could be some repetitions the team is

  240. 9:25

    experiencing they're still annoying at

  241. 9:27

    scale and if you are doing a bake off

  242. 9:29

    it's still worth solving for them I

  243. 9:31

    would not ignore these class of problems

  244. 9:33

    oneoff and high impact well hope it

  245. 9:35

    doesn't chronic and I think systematic

  246. 9:38

    and high impact are obviously the P 0

  247. 9:40

    you know target areas um for the team to

  248. 9:42

    solve an example of that

  249. 9:45

    would be in a fins serve capacity

  250. 9:47

    There's a voice agent that um helps

  251. 9:49

    users freeze their credit cards. And if

  252. 9:51

    it doesn't do that, well, that's a

  253. 9:53

    massive fail.

  254. 9:56

    All right. So, understand and execute.

  255. 9:58

    I'm pretty sure everyone's doing this.

  256. 10:00

    Please fix my agent. Uh I think fixing

  257. 10:02

    or rather attempting to make a fix is

  258. 10:05

    the simplest and the lowest effort

  259. 10:08

    component of this debugging pipeline and

  260. 10:11

    loop. Um the next step is all right, I

  261. 10:14

    made a change to my system. How do I

  262. 10:16

    actually know this thing works um for

  263. 10:19

    real? A great way that's naive is to

  264. 10:23

    take a real call, for example, in my

  265. 10:25

    case, I booked an appointment and it

  266. 10:27

    didn't get scheduled and replay that

  267. 10:29

    exact conversation and run that maybe 5,

  268. 10:32

    10, 20, 50 times and see, okay, what is

  269. 10:34

    my probability of passing this type of

  270. 10:37

    issue? A better way is to keep the same

  271. 10:40

    intent but change the wordings, change

  272. 10:44

    the patterns, change the accents, change

  273. 10:46

    the style, add one more intent to the

  274. 10:48

    mix. And that gives teams much more, you

  275. 10:51

    know, better coverage to feel confident

  276. 10:53

    that yes, I actually made a change and

  277. 10:56

    my changes are net positive instead of

  278. 10:58

    net negative.

  279. 11:01

    There are certain fixes and I guess

  280. 11:05

    hypothesis that are very difficult to

  281. 11:07

    test in a pre-eployment synthetic

  282. 11:09

    setting. And so AB testing ends up

  283. 11:11

    being, you know, pretty pretty critical

  284. 11:13

    for those circumstances. For example, if

  285. 11:15

    you have an outbound agent, the first 5

  286. 11:17

    seconds of a conversation tends to be

  287. 11:19

    the most important. And so the vocal

  288. 11:22

    quality um and the specific words you

  289. 11:25

    end up using, they matter the most. And

  290. 11:27

    so AB testing that is the only way in in

  291. 11:29

    kind of real life setting to to get

  292. 11:32

    results. You can't really do it through

  293. 11:34

    simulations alone.

  294. 11:37

    And so there we have the loop. Identify,

  295. 11:39

    prioritize, impact size, understand the

  296. 11:42

    fix, execute, check, make sure it didn't

  297. 11:45

    break anything, and then continue

  298. 11:47

    monitoring.

  299. 11:51

    So I think making voice agents useful is

  300. 11:54

    already hard as it is. even when dealing

  301. 11:58

    with earnest users on the other line,

  302. 12:01

    right? These are people who who just

  303. 12:02

    want their problem solved. They're not

  304. 12:03

    trying to mess with you. These are like

  305. 12:04

    legit normal people.

  306. 12:08

    Now, what happens when mythos learns how

  307. 12:10

    to dial?

  308. 12:13

    So, if it can extract trade secrets and

  309. 12:16

    uh you know, from the NSA, it can

  310. 12:18

    certainly, you know, seduce you into

  311. 12:22

    revealing PHI and PII data as well.

  312. 12:26

    And I think both voice agents and humans

  313. 12:29

    will [clears throat] be targeted here.

  314. 12:31

    Voice agents because there's a pressure

  315. 12:33

    to make these more capable. Give them

  316. 12:37

    access to more data. Give them access to

  317. 12:39

    more tools.

  318. 12:41

    Deploy them quickly.

  319. 12:45

    The more the capability, the bigger the

  320. 12:46

    surface area. This is this is pretty

  321. 12:48

    pretty common sense. And the more the

  322. 12:50

    voice agents become natural and human

  323. 12:54

    sounding, the more humans will be

  324. 12:56

    tricked along the way as well for those

  325. 12:58

    who are weaponizing.

  326. 13:01

    Uh we ship a we shipped a red tipping

  327. 13:03

    product um back in April just to test

  328. 13:05

    out this hypothesis for how many agents

  329. 13:07

    can we actually break from a adversarial

  330. 13:10

    capacity and we can probably break one

  331. 13:12

    in five agents at this point. We've

  332. 13:14

    tested this across financial services,

  333. 13:16

    healthcare, um consumer and so on. We've

  334. 13:19

    bypassed verification. Uh we've

  335. 13:22

    definitely had agents, you know, we've

  336. 13:24

    been able to promject uh several agents

  337. 13:27

    and and gotten data we should not have.

  338. 13:29

    So this is not theoretical. This is

  339. 13:31

    actually a real a real concern.

  340. 13:37

    I think the only real defense against

  341. 13:39

    the dark arts is

  342. 13:41

    step one to invest deeply in

  343. 13:43

    pre-eployment testing. This could be

  344. 13:45

    textto text. This could be voice to

  345. 13:48

    voice. There's pros and cons to both.

  346. 13:50

    Happy to chat offline if folks are

  347. 13:51

    interested. And this is just making sure

  348. 13:54

    you're not self-owning, you know, when

  349. 13:56

    you're talking to real people who just

  350. 13:57

    want to get their problem solved. Step

  351. 13:59

    two is to have a great monitoring system

  352. 14:02

    of all kinds. And I've highlighted, you

  353. 14:04

    know, different flavors of monitoring

  354. 14:06

    per call scoring, manual kind of evals,

  355. 14:10

    you know, listening to conversations and

  356. 14:11

    cross call analysis. And this is helpful

  357. 14:14

    both for monitoring what the agent is

  358. 14:16

    saying and behaving and how it's

  359. 14:17

    actually doing, but also the users. Are

  360. 14:19

    the users being adversarial? Are they

  361. 14:21

    being annoying? Are are they trying to

  362. 14:22

    trick the agent into doing things it's

  363. 14:24

    not supposed to be doing?

  364. 14:26

    And I think our new recommendation now

  365. 14:28

    is to run 24/7 red teaming um for your

  366. 14:31

    agents, especially if you believe the

  367. 14:34

    cost of bad interactions can be can be

  368. 14:36

    rather large.

  369. 14:38

    So, I think voice agents um have this

  370. 14:40

    awesome potential of of making the world

  371. 14:43

    feel much more human compared to

  372. 14:48

    interacting with clunky IVR trees or

  373. 14:50

    chat bots or worse um being stuck on a

  374. 14:53

    on a hold.

  375. 14:55

    And when we think about crime, we often

  376. 14:57

    think of crime happening to somebody

  377. 14:59

    else. You know, crime does not happen to

  378. 15:01

    you typically with voice agents,

  379. 15:04

    especially bad actors. As these agents

  380. 15:07

    are deployed and as bad actors start to

  381. 15:09

    exploit a lot of the vulnerabilities,

  382. 15:11

    the number of incidents is about to kind

  383. 15:13

    of go way way up. And so the reason I

  384. 15:16

    fear voice agents more than crime is

  385. 15:18

    that one of these incidents is going to

  386. 15:20

    impact you. It already did for me.

  387. 15:24

    Awesome. So it's time for me to shill.

  388. 15:25

    Well, we burn a lot of tokens. If you

  389. 15:28

    are interested in working in this space,

  390. 15:31

    please come and talk to us. And if you

  391. 15:33

    are deploying voice agents and want to

  392. 15:35

    validate whether your architectured or

  393. 15:39

    your eval are set up correctly, please

  394. 15:40

    come and talk to us. We'll be outside.

  395. 15:42

    And here's here's my number. Here's my

  396. 15:44

    WhatsApp. Thanks everyone.

  397. 16:02

    >> [music]