"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow

Read the talk

“My Name Is… My Name Is…”: A Linguistic Map for Voice Agents

Midam Kim maps a failed support call across recognition, pronunciation, timing, word choice, and shared understanding—showing why a voice agent must manage a conversation as an interdependent activity unfolding over time.

From a talk by Midam Kim

At a glance

Ideas worth remembering

  • Diagnose voice-agent failures separately across listening and speaking, then across sounds, words, interaction, and mental model.

  • Recognition, pronunciation, timing, language, and intent tracking are interdependent; improving one component does not guarantee a successful conversation.

  • A failed exchange needs a changed repair strategy. Repeating the same request without clarification deepens frustration and makes human escalation more likely.

  • Spoken signals disappear, but the user’s mental model accumulates. Design the sequence of turns, not just the quality of each isolated output.

  • Voice orchestration must adapt during the call to different people, changing emotion, unfamiliar inputs, and evolving context.

  • The linguistic map is a diagnostic framework rather than a universal fix, and it must remain useful as users adapt and language changes.

A small recognition error becomes a failed call

Midam Kim, an ML engineer at ServiceNow and a researcher of speech communication in the wild, starts with a support call that deteriorates one failure at a time. The voice agent asks her to spell her first name. She supplies M-I-D-A-M slowly; it confirms M-I-D-A-N. She corrects the final letter, and the agent thanks her—then still pronounces her name incorrectly. The confirmation language sounds successful, but the spoken result shows that the correction did not survive the trip through the system.

Shows the next stage of the failed call as Kim searches for and slowly reads an unfamiliar account number before being cut off.
Shows the next stage of the failed call as Kim searches for and slowly reads an unfamiliar account number before being cut off.

The next request compounds the problem. The agent asks for an account number without making clear which identifier it means. Kim searches for the unfamiliar string and begins reading it carefully. Because she needs extra time, the agent cuts her off, reports that it cannot find the record, and asks her to repeat the same information. Nothing about the second attempt has changed: no narrower question, no partial confirmation, no alternative input method. She asks for a person.

This is one observable failure—a caller abandons the automated path—but it has several causes. A letter was misrecognized, a name was mispronounced, the requested identifier was unclear, turn detection fired too early, and the repair strategy ignored what had just happened. Treating the episode as one generic “voice AI bug” would hide the engineering decisions needed to fix it. Kim’s proposed diagnostic lens is linguistics: separate the mechanisms, then inspect how they interact across the whole call.

0:190:33
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Conversation is joint work, not one-way inference

Human conversation is a joint activity. One participant produces sounds and words; the other listens and responds; the exchange moves back and forth through interaction. Meanwhile, both participants continuously update their mental models: what the other person means, what has already been established, and what should happen next. People bring those expectations to a voice agent even if its implementation is organized as a sequence of models and services.

Revisits the name failure through separate listening and speaking channels, distinguishing ASR recognition from TTS pronunciation.
Revisits the name failure through separate listening and speaking channels, distinguishing ASR recognition from TTS pronunciation.

Replaying the call through that lens separates listening from speaking. Confusing M with N belongs to the agent’s listening side: the recognized signal does not preserve the supplied spelling. Incorrectly saying the resulting name belongs to its speaking side: the text-to-speech system applies English-centric pronunciation rules that do not fit the name. Fixing recognition alone would therefore leave the pronunciation problem intact.

The account-number exchange crosses even more layers. Kim adapts to an unclear request by searching for the identifier and reading it slowly. The agent fails to recognize the relevant unit, decides that her turn is over, and speaks over her. When it requests an unchanged repetition, it also fails at conversational repair. A human might confirm a prefix, ask for one character at a time, explain which identifier is needed, or offer another channel. The agent instead resets the burden onto the caller, making escalation the rational choice.

3:344:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:34 · section reference included

Eight cells locate different kinds of failure

What question does the linguistic map answer? It shows where a failure occurs by crossing two channels—listening and speaking—with four levels: sounds, words, interaction, and mental model. The resulting eight cells distinguish responsibilities that are often collapsed into “the voice stack.” The relationship made visible is symmetry: receiving speech and producing speech each require acoustic competence, understandable language, appropriate timing, and an adequate model of the user’s goal.

Covers the transition from listening-side intention to speaking-side pronunciation, vocabulary, and timing, helping illustrate the eight-cell map.
Covers the transition from listening-side intention to speaking-side pronunciation, vocabulary, and timing, helping illustrate the eight-cell map.

On the listening side:

  • Sounds — recognition: Does the agent recognize the user’s speech correctly?
  • Words — understanding: Does it understand the words the user supplied?
  • Interaction — turn detection: Does it wait until the user has finished?
  • Mental model — intention: Does it understand what the user is trying to accomplish?

On the speaking side:

  • Sounds — pronunciation: Does the synthesized speech say names and other content appropriately?
  • Words — vocabulary: Does the agent choose language the user can understand?
  • Interaction — turn taking: Does it speak at the right time?
  • Mental model — useful information: Does it provide what the user actually needs next?
Compare the ideasThe linguistic map for a voice agent

Recognize the user’s speech.

Listening and speaking mirror one another across four levels. The cells locate failures, but they still operate as one conversation.

6:347:04
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:34 · section reference included

A better component can still produce a worse conversation

The grid is diagnostic, not a claim that its cells can be optimized independently. Sound-level work has to account for words: recognizing individual phonetic material is not enough if the assembled identifier is wrong. Sounds and words must then align with interaction: even accurate transcription cannot help if endpointing cuts the user off before the unit is complete. Task completion at the mental-model layer depends on all of them working together.

Captures the key contrast between vanishing spoken signals and the user’s persistent, accumulating mental model.
Captures the key contrast between vanishing spoken signals and the user’s persistent, accumulating mental model.

The name example makes the dependency concrete. First, recognition changes the final letter from M to N. After correction, the speaking path still applies unsuitable pronunciation rules. The user then enters the account-number task already carrying evidence that the agent may not handle her input reliably. Premature turn detection adds another failure, and the unchanged repetition request confirms that the agent has not learned from the exchange. Each event updates the interpretation of the next one.

What survives this sequence after the audio is gone? The timeline below separates transient signals from persistent understanding. Unlike chat, which leaves a visible text history, spoken sounds vanish as air vibrations. What accumulates for the caller is a mental model: whether the agent listens, whether corrections matter, whether waiting is safe, and whether another attempt is worth making. That accumulating interpretation is the relationship the visual makes visible—and the experience the system ultimately has to improve.

How it fits togetherSpeech vanishes while the user’s mental model accumulates

The acoustic signal is temporary.

Each transient exchange leaves a persistent interpretation that shapes the next turn.

8:238:53
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:23 · section reference included

Engineering work must cover models, timing, language, and context

The map turns into several parallel areas of work rather than one preferred model choice:

  • Recognition: Select and configure automatic speech recognition, then add appropriate post-processing.
  • Speech generation: Select and configure text-to-speech, with preprocessing for what the system must say.
  • Shared vocabulary: Curate words and expressions that both the agent and its users can understand.
  • Interaction: Manage turn detection, latency, and turn taking together.
  • Conversation state: Retain context across all these layers, and detect and handle changes in the user’s emotional state.
Introduces EVA Bench as an end-to-end diagnostic option without claiming evaluation results not given in the talk.
Introduces EVA Bench as an end-to-end diagnostic option without claiming evaluation results not given in the talk.

The orchestration must also change during the call. A child and an adult may have different timing, vocabulary, and clarification needs. The same person may slow down while reading an unfamiliar identifier, become frustrated after an interruption, or stop trusting confirmations after a failed correction. Static settings chosen before the call cannot fully account for those changes; the agent has to manage them along the timeline, including when the user is unhappy.

This is why “voice is natural” can be misleading as an engineering description. Speech feels natural to people because humans perform a great deal of linguistic coordination automatically. A voice agent has to recreate enough of that coordination across recognition, generation, timing, context, and repair. When those mechanisms disagree, the failure is experienced as one conversation rather than a set of isolated component errors.

Kim connects this work to user frustration, task failure, live-agent escalation, abandoned calls, and silent failures. The talk does not report measured reductions for those outcomes. She also points to ServiceNow’s EVA Bench as an end-to-end diagnostic benchmark, but does not describe its evaluation procedure or results here; its role in the talk is as a possible way to assess a voice agent’s current state.

10:2510:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

10:25 · section reference included

The framework diagnoses; the remedy belongs to the system

The closing distinction is practical: voice AI is a joint activity between the user and the agent, not merely a pipeline that transforms audio into text and back again. Needs must be served across multiple layers in real time. There is no universal configuration that repairs every agent; the map helps a team locate its own failure, understand adjacent dependencies, and decide what to change. Kim’s next steps are direct: evaluate the system, learn linguistics, and involve linguists in its design.

Supports the closing argument that speakers adapt to one another and that voice agents should be designed for adaptation.
Supports the closing argument that speakers adapt to one another and that voice agents should be designed for adaptation.

The longer-term challenge is adaptation. During any conversation, people learn one another’s accent, vocabulary, rhythm, and preferred ways of explaining things. Users will likewise adapt to a voice agent over the course of a call. A system should ask whether that learning can make the next exchange easier rather than forcing the user to rediscover the same limitations each time.

Language itself also moves. Pronunciations, vocabulary, and expectations can change over a year—or even six months. A voice agent that performs acceptably today therefore needs mechanisms for reevaluation and revision. The useful map is not a finished checklist. It is a way to keep finding where a changing agent, a changing user, and a changing language have fallen out of alignment.

12:4312:59
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:43 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> Okay, hello everyone.

  3. 0:16

    So,

  4. 0:19

    my name is Midam Kim. I am an ML

  5. 0:22

    engineer from ServiceNow and I'll be

  6. 0:25

    talking about a linguistic framework for

  7. 0:28

    voice AI.

  8. 0:33

    So,

  9. 0:35

    quick background of me so you know where

  10. 0:37

    I'm coming from.

  11. 0:39

    Like I said, I'm an ML engineer at

  12. 0:41

    ServiceNow, but I'm also a researcher,

  13. 0:43

    lifelong researcher, of speech

  14. 0:45

    communication in the wild.

  15. 0:47

    So, my motto is doing linguistics and

  16. 0:51

    what I'm going to be doing today is to

  17. 0:54

    hand you that lens of linguistics.

  18. 1:00

    So, have you experienced voice AI

  19. 1:03

    failures?

  20. 1:05

    Yeah, like everyone.

  21. 1:06

    >> [laughter]

  22. 1:10

    >> So, I'm going to introduce an example

  23. 1:12

    that I experienced myself.

  24. 1:15

    So, the bot asked me, "Could you please

  25. 1:18

    spell your first name?"

  26. 1:20

    And then I slowly start to spell my

  27. 1:23

    name.

  28. 1:24

    Yes, it is m i d a m.

  29. 1:29

    And the bot says, "Confirming with you,

  30. 1:31

    is it m i d a n?"

  31. 1:35

    And then I say, "No, it is m i d a m."

  32. 1:41

    Um

  33. 1:42

    and the bot says,

  34. 1:44

    "Thank you for your correction. Happy to

  35. 1:46

    help you today, Madam."

  36. 1:48

    And I then I

  37. 1:50

    get slightly annoyed, more annoyed,

  38. 1:51

    because my name is Midam, not Madam.

  39. 1:55

    And then it asked me about, "Now, what

  40. 1:57

    is your account number?

  41. 1:59

    And then, I start start getting

  42. 2:01

    confused. What is that account number

  43. 2:03

    thing?

  44. 2:04

    And then,

  45. 2:06

    I try to find uh information about that.

  46. 2:10

    So,

  47. 2:11

    which one? Um it must be and I start

  48. 2:16

    uh

  49. 2:17

    slowly start spelling the account

  50. 2:20

    number. So, it is A X 4 5 1.

  51. 2:25

    And then, I take time because I'm not

  52. 2:28

    used to reading this strange number.

  53. 2:32

    And then, the bot cuts me off.

  54. 2:34

    And then, it says, I couldn't find your

  55. 2:36

    record.

  56. 2:37

    And then, without even trying, it asked

  57. 2:40

    me to repeat that again. Can you please

  58. 2:42

    repeat that? And then, I get super

  59. 2:44

    annoyed and then, I can say, can I talk

  60. 2:46

    to a person?

  61. 2:48

    I just don't want to deal with you

  62. 2:49

    anymore.

  63. 2:50

    So, this is a very typical pattern of

  64. 2:53

    voice AI, unfortunately, at this point.

  65. 2:56

    So, I just want to navigate how we can

  66. 2:59

    solve this problem

  67. 3:01

    with linguistics.

  68. 3:06

    So, voice AI is booming.

  69. 3:08

    But users are still often preferring

  70. 3:10

    human agents over voice agents.

  71. 3:13

    How can we mitigate this issue?

  72. 3:17

    But in the first place, what are the

  73. 3:19

    actual problems?

  74. 3:21

    So, I think we can think about a

  75. 3:23

    fundamental frame framework to

  76. 3:25

    understand this into an architecture of

  77. 3:28

    voice AI,

  78. 3:29

    which is called linguistics.

  79. 3:34

    So, as all of us already know,

  80. 3:38

    human communication is a joint activity,

  81. 3:41

    like the thing that we're doing right

  82. 3:42

    now.

  83. 3:43

    So, I give you my sounds and words.

  84. 3:47

    You hear them.

  85. 3:49

    And then, if it is a conversation,

  86. 3:51

    you're going to give me your sounds and

  87. 3:53

    your words.

  88. 3:55

    And then this is going back and forth

  89. 3:58

    through interaction.

  90. 4:01

    And then in this process, we're

  91. 4:03

    continuously

  92. 4:04

    processing and updating our mental

  93. 4:07

    models.

  94. 4:09

    So that's a joint activity

  95. 4:12

    for human communication.

  96. 4:14

    And I would like to say

  97. 4:17

    in the voice AI human communication,

  98. 4:20

    it also has to be a joint activity like

  99. 4:23

    this.

  100. 4:24

    Because that's the only thing that we

  101. 4:27

    know about human communication as a

  102. 4:29

    human being. We have been evolving

  103. 4:31

    thousands of years as communicators, and

  104. 4:34

    this is what we know. So we expect the

  105. 4:36

    same thing to bots.

  106. 4:41

    So let me go over the failure scene of

  107. 4:44

    my call with the voice agent

  108. 4:47

    in this framework.

  109. 4:49

    So you see there's listen

  110. 4:52

    and speak for each party.

  111. 4:57

    So I start spelling my first name.

  112. 5:00

    And then the bot did not hear that the

  113. 5:04

    difference between M and N correctly, so

  114. 5:07

    it's an

  115. 5:08

    SCT failure in the listening level.

  116. 5:12

    And then the TTS applies only

  117. 5:15

    English-centric reading rules to my

  118. 5:17

    name, M I D A M, would read it as Midam

  119. 5:21

    in the

  120. 5:22

    American English version.

  121. 5:25

    So I'm confused, but at this time I'm

  122. 5:27

    kind of generous because that happens a

  123. 5:29

    lot even with human beings. So I'm okay.

  124. 5:34

    But then when it brought

  125. 5:35

    brought up account number thing

  126. 5:38

    because I don't know what that is,

  127. 5:40

    I'm confused again.

  128. 5:42

    But I'm adaptive, I can find I can look

  129. 5:45

    for it.

  130. 5:46

    So I found the number, start reading it,

  131. 5:48

    but

  132. 5:50

    the STT did not recognize the word unit

  133. 5:53

    correctly, so

  134. 5:54

    it cuts me off, and uh

  135. 5:58

    uh finally, it's uh eventually talked

  136. 6:02

    over me.

  137. 6:03

    So, I get

  138. 6:05

    really irritated.

  139. 6:08

    And then, when it asked me for the

  140. 6:09

    repetition of the same information, and

  141. 6:13

    then, it is clear that the spot is not

  142. 6:15

    tracking the mental model with me.

  143. 6:18

    And then, very rudely, it's uh does not

  144. 6:22

    even try interactive clarification,

  145. 6:24

    which is a common strategy by human

  146. 6:26

    beings.

  147. 6:27

    So, I don't want to deal with this

  148. 6:29

    anymore, so I say, "Can I talk to a

  149. 6:30

    person?"

  150. 6:34

    So,

  151. 6:35

    let's go over the uh the framework

  152. 6:37

    again. So, the these are the linguistic

  153. 6:39

    components that are expected and well

  154. 6:41

    maintained in human-to-human voi- uh

  155. 6:45

    uh conversation.

  156. 6:47

    So, there are listening channels, a

  157. 6:49

    listening channel and speaking channel,

  158. 6:50

    and there are different components like

  159. 6:52

    sounds, words, interaction, and mental

  160. 6:54

    model.

  161. 6:55

    So, the first component is, does the bot

  162. 6:58

    recognize the user's speech well?

  163. 7:01

    And all of these technical terms

  164. 7:04

    uh will fall under this.

  165. 7:07

    And then, there was there's going to be

  166. 7:08

    this second component, which is words in

  167. 7:11

    the listening channel. So, does the bot

  168. 7:13

    understand the user's words?

  169. 7:17

    And then, the third one is, does the bot

  170. 7:19

    wait until the right timing to for its

  171. 7:22

    turn? It's about It's going to be about

  172. 7:24

    uh listening channel interaction.

  173. 7:28

    And then, uh the last part is mental

  174. 7:31

    model. So, does the bot understand the

  175. 7:33

    user's intention

  176. 7:35

    in the listening part?

  177. 7:38

    And then, we can also go to the speaking

  178. 7:39

    channel, so it's going to be about

  179. 7:41

    pronunciation for the sound.

  180. 7:43

    And also there's about understand the

  181. 7:46

    the words users are

  182. 7:49

    uh there's about choose the words the

  183. 7:51

    user can understand.

  184. 7:53

    And in the interaction part, there's

  185. 7:55

    about speak with the right timing.

  186. 7:58

    And lastly, there's about speak with the

  187. 8:01

    information the user actually need.

  188. 8:05

    So, there are a lot of engineering or

  189. 8:08

    linguistic or cognitive science terms

  190. 8:10

    that are in here that that are here. Um

  191. 8:14

    you can see now see that all of those

  192. 8:17

    have their right spots in this

  193. 8:18

    linguistic framework.

  194. 8:23

    And importantly, these components are

  195. 8:24

    interdependent,

  196. 8:26

    not separate or uh independent from each

  197. 8:29

    other. They're interdependent and

  198. 8:31

    they're aligned. So, when you want to do

  199. 8:34

    good things about sounds,

  200. 8:37

    you have to think about words level.

  201. 8:39

    And then when you want to do good things

  202. 8:41

    about these sounds and words,

  203. 8:43

    you also have to uh account for

  204. 8:46

    interaction, so turn taking or turn

  205. 8:48

    detection.

  206. 8:50

    And then finally, you want to uh have

  207. 8:53

    good uh task completion, which is the

  208. 8:56

    goal of these mental model uh layer.

  209. 8:59

    Then you have to have all of these.

  210. 9:02

    Without all of those, without any of

  211. 9:04

    those, any of those components, your

  212. 9:06

    voice agent will fail.

  213. 9:09

    And then finally,

  214. 9:11

    uh it has to be well aligned. All of

  215. 9:13

    these have to be well aligned.

  216. 9:17

    And additionally, you have to keep your

  217. 9:20

    mind keep in mind that

  218. 9:22

    this is happening on the timeline.

  219. 9:26

    What I mean by that is it is silently

  220. 9:29

    tracked. Unlike in chat, in chat you see

  221. 9:33

    the history of what was said

  222. 9:35

    uh as text.

  223. 9:37

    But in voice agent experience,

  224. 9:40

    uh, you say something, and the bot says

  225. 9:42

    something, you go back and forth,

  226. 9:45

    and then see, all these waveforms, the

  227. 9:49

    air via the vibration in the, uh, in the

  228. 9:51

    air, they're all gone.

  229. 9:53

    And only the user's mental model is the

  230. 9:56

    thing that's left, and that matters.

  231. 10:00

    So, sounds, words, interactions vanish

  232. 10:03

    the moment they're spoken,

  233. 10:04

    but the mental model proceeds and grows

  234. 10:07

    over the timeline.

  235. 10:09

    So, this is what you have to

  236. 10:12

    target

  237. 10:14

    for user satisfaction.

  238. 10:16

    And then, what can we do

  239. 10:19

    for the bot to meet the standard of the

  240. 10:22

    user?

  241. 10:25

    So, what we can do, uh, would include,

  242. 10:28

    of course, choosing good ASR models or

  243. 10:31

    configurations and do some

  244. 10:32

    post-processing,

  245. 10:34

    uh, choosing good TTS models,

  246. 10:36

    configurations, and pre-processing,

  247. 10:38

    and, uh, carefully curate the vocabulary

  248. 10:42

    that can be shared between the bot and

  249. 10:44

    the user,

  250. 10:45

    and do good job of a turn-to-turn

  251. 10:48

    detection, latency, and turn-taking.

  252. 10:52

    Um, and very importantly, we have to, it

  253. 10:56

    would be great if we can do good emotion

  254. 10:57

    detection and handling, and context

  255. 11:00

    retention, and by context, what I mean

  256. 11:02

    is context about all of these.

  257. 11:07

    And importantly,

  258. 11:09

    uh, it has to be dynamic because things

  259. 11:12

    are always changing, uh, throughout over

  260. 11:14

    the course of the call. So, we would

  261. 11:17

    have to do this management dynamically

  262. 11:19

    along the timeline

  263. 11:21

    for different kinds of people.

  264. 11:23

    So, kids or different kinds of people

  265. 11:26

    like these will have different

  266. 11:28

    expectations that we have to satisfy.

  267. 11:32

    Uh, not just when they're happy, but

  268. 11:34

    also when they're not happy.

  269. 11:36

    So, only then you can pursue a dynamic

  270. 11:39

    and truly scalable orchestration of

  271. 11:41

    voice AI.

  272. 11:43

    So, it's a very difficult job to do.

  273. 11:48

    We always say that voice is the most

  274. 11:50

    natural way of communication, but it is

  275. 11:53

    not actually not easy. Behind the scene,

  276. 11:55

    it is thanks to this linguistic

  277. 11:57

    orchestration.

  278. 11:59

    When your bot is not good at it,

  279. 12:01

    it's a catastrophic failure.

  280. 12:06

    Um, so paying attention to this

  281. 12:09

    linguistic framework would have lots of

  282. 12:12

    business implications because then you

  283. 12:14

    can uh

  284. 12:17

    decrease all of these user frustration,

  285. 12:19

    task failures, live agent escalation, or

  286. 12:22

    abandoned calls, or silent failures.

  287. 12:28

    So, in ServiceNow, we have made a a good

  288. 12:32

    uh benchmark end-to-end benchmark called

  289. 12:34

    Eva bench. So, you can try that to

  290. 12:36

    diagnose your voice agent's uh status.

  291. 12:43

    Um, key takeaways.

  292. 12:45

    So, voice AI is a joint activity between

  293. 12:50

    the bot and the user, not just a

  294. 12:52

    pipeline.

  295. 12:54

    And we must serve users' needs in

  296. 12:55

    multiple layers real time.

  297. 12:59

    It's not that I have given you a fix

  298. 13:01

    today because there's nothing like that.

  299. 13:04

    It just uh the fix is in you and your

  300. 13:07

    system.

  301. 13:09

    But, what I have given you is today is

  302. 13:14

    the linguistic framework you can try to

  303. 13:16

    diagnose your system

  304. 13:18

    and to build your system upon.

  305. 13:21

    You can try Eva, but also you can learn

  306. 13:24

    linguistics and hire linguists.

  307. 13:27

    Um, another thing I want to remind you

  308. 13:29

    of is that business implications are

  309. 13:32

    linguistic implications and vice versa

  310. 13:35

    in this voice AI scene. Because voice is

  311. 13:39

    fundamentally a linguistic and very

  312. 13:41

    human and cognitive experience.

  313. 13:46

    I would like to ask you a longer term

  314. 13:48

    question.

  315. 13:50

    Speakers adapt. So, I

  316. 13:54

    I'm pretty sure that in this talk in my

  317. 13:58

    talk with you guys today, you have

  318. 14:00

    learned something about me, about my

  319. 14:02

    speaking style, what kind of accents I

  320. 14:04

    speak, what kind of words I'm using. So,

  321. 14:07

    next time I see you guys in person, you

  322. 14:10

    would find it more comfortable to talk

  323. 14:12

    to me because you have paid attention to

  324. 14:14

    me.

  325. 14:15

    Right? So, speakers are always adapting.

  326. 14:17

    So, the user will be adapting to your

  327. 14:20

    voice agent throughout the call. So, is

  328. 14:24

    your system ready for them to

  329. 14:27

    use you better, use it your voice agent

  330. 14:29

    better the next time?

  331. 14:31

    And

  332. 14:32

    language is always change. So, is your

  333. 14:35

    voice agent ready for language change in

  334. 14:38

    1 year or 6 months even?

  335. 14:44

    So, thank you.

  336. 14:47

    >> [applause]