"My name is... my name is...": A Linguistic Map for Voice Agents — Midam Kim, ServiceNow
Read the talk
“My Name Is… My Name Is…”: A Linguistic Map for Voice Agents
Midam Kim maps a failed support call across recognition, pronunciation, timing, word choice, and shared understanding—showing why a voice agent must manage a conversation as an interdependent activity unfolding over time.
From a talk by Midam Kim
At a glance
Ideas worth remembering
Diagnose voice-agent failures separately across listening and speaking, then across sounds, words, interaction, and mental model.
Recognition, pronunciation, timing, language, and intent tracking are interdependent; improving one component does not guarantee a successful conversation.
A failed exchange needs a changed repair strategy. Repeating the same request without clarification deepens frustration and makes human escalation more likely.
Spoken signals disappear, but the user’s mental model accumulates. Design the sequence of turns, not just the quality of each isolated output.
Voice orchestration must adapt during the call to different people, changing emotion, unfamiliar inputs, and evolving context.
The linguistic map is a diagnostic framework rather than a universal fix, and it must remain useful as users adapt and language changes.
A small recognition error becomes a failed call
Midam Kim, an ML engineer at ServiceNow and a researcher of speech communication in the wild, starts with a support call that deteriorates one failure at a time. The voice agent asks her to spell her first name. She supplies M-I-D-A-M slowly; it confirms M-I-D-A-N. She corrects the final letter, and the agent thanks her—then still pronounces her name incorrectly. The confirmation language sounds successful, but the spoken result shows that the correction did not survive the trip through the system.
The next request compounds the problem. The agent asks for an account number without making clear which identifier it means. Kim searches for the unfamiliar string and begins reading it carefully. Because she needs extra time, the agent cuts her off, reports that it cannot find the record, and asks her to repeat the same information. Nothing about the second attempt has changed: no narrower question, no partial confirmation, no alternative input method. She asks for a person.
This is one observable failure—a caller abandons the automated path—but it has several causes. A letter was misrecognized, a name was mispronounced, the requested identifier was unclear, turn detection fired too early, and the repair strategy ignored what had just happened. Treating the episode as one generic “voice AI bug” would hide the engineering decisions needed to fix it. Kim’s proposed diagnostic lens is linguistics: separate the mechanisms, then inspect how they interact across the whole call.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Conversation is joint work, not one-way inference
Human conversation is a joint activity. One participant produces sounds and words; the other listens and responds; the exchange moves back and forth through interaction. Meanwhile, both participants continuously update their mental models: what the other person means, what has already been established, and what should happen next. People bring those expectations to a voice agent even if its implementation is organized as a sequence of models and services.
Replaying the call through that lens separates listening from speaking. Confusing M with N belongs to the agent’s listening side: the recognized signal does not preserve the supplied spelling. Incorrectly saying the resulting name belongs to its speaking side: the text-to-speech system applies English-centric pronunciation rules that do not fit the name. Fixing recognition alone would therefore leave the pronunciation problem intact.
The account-number exchange crosses even more layers. Kim adapts to an unclear request by searching for the identifier and reading it slowly. The agent fails to recognize the relevant unit, decides that her turn is over, and speaks over her. When it requests an unchanged repetition, it also fails at conversational repair. A human might confirm a prefix, ask for one character at a time, explain which identifier is needed, or offer another channel. The agent instead resets the burden onto the caller, making escalation the rational choice.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Eight cells locate different kinds of failure
What question does the linguistic map answer? It shows where a failure occurs by crossing two channels—listening and speaking—with four levels: sounds, words, interaction, and mental model. The resulting eight cells distinguish responsibilities that are often collapsed into “the voice stack.” The relationship made visible is symmetry: receiving speech and producing speech each require acoustic competence, understandable language, appropriate timing, and an adequate model of the user’s goal.
On the listening side:
- Sounds — recognition: Does the agent recognize the user’s speech correctly?
- Words — understanding: Does it understand the words the user supplied?
- Interaction — turn detection: Does it wait until the user has finished?
- Mental model — intention: Does it understand what the user is trying to accomplish?
On the speaking side:
- Sounds — pronunciation: Does the synthesized speech say names and other content appropriately?
- Words — vocabulary: Does the agent choose language the user can understand?
- Interaction — turn taking: Does it speak at the right time?
- Mental model — useful information: Does it provide what the user actually needs next?
Recognize the user’s speech.
Listening and speaking mirror one another across four levels. The cells locate failures, but they still operate as one conversation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A better component can still produce a worse conversation
The grid is diagnostic, not a claim that its cells can be optimized independently. Sound-level work has to account for words: recognizing individual phonetic material is not enough if the assembled identifier is wrong. Sounds and words must then align with interaction: even accurate transcription cannot help if endpointing cuts the user off before the unit is complete. Task completion at the mental-model layer depends on all of them working together.
The name example makes the dependency concrete. First, recognition changes the final letter from M to N. After correction, the speaking path still applies unsuitable pronunciation rules. The user then enters the account-number task already carrying evidence that the agent may not handle her input reliably. Premature turn detection adds another failure, and the unchanged repetition request confirms that the agent has not learned from the exchange. Each event updates the interpretation of the next one.
What survives this sequence after the audio is gone? The timeline below separates transient signals from persistent understanding. Unlike chat, which leaves a visible text history, spoken sounds vanish as air vibrations. What accumulates for the caller is a mental model: whether the agent listens, whether corrections matter, whether waiting is safe, and whether another attempt is worth making. That accumulating interpretation is the relationship the visual makes visible—and the experience the system ultimately has to improve.
The acoustic signal is temporary.
Each transient exchange leaves a persistent interpretation that shapes the next turn.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Engineering work must cover models, timing, language, and context
The map turns into several parallel areas of work rather than one preferred model choice:
- Recognition: Select and configure automatic speech recognition, then add appropriate post-processing.
- Speech generation: Select and configure text-to-speech, with preprocessing for what the system must say.
- Shared vocabulary: Curate words and expressions that both the agent and its users can understand.
- Interaction: Manage turn detection, latency, and turn taking together.
- Conversation state: Retain context across all these layers, and detect and handle changes in the user’s emotional state.
The orchestration must also change during the call. A child and an adult may have different timing, vocabulary, and clarification needs. The same person may slow down while reading an unfamiliar identifier, become frustrated after an interruption, or stop trusting confirmations after a failed correction. Static settings chosen before the call cannot fully account for those changes; the agent has to manage them along the timeline, including when the user is unhappy.
This is why “voice is natural” can be misleading as an engineering description. Speech feels natural to people because humans perform a great deal of linguistic coordination automatically. A voice agent has to recreate enough of that coordination across recognition, generation, timing, context, and repair. When those mechanisms disagree, the failure is experienced as one conversation rather than a set of isolated component errors.
Kim connects this work to user frustration, task failure, live-agent escalation, abandoned calls, and silent failures. The talk does not report measured reductions for those outcomes. She also points to ServiceNow’s EVA Bench as an end-to-end diagnostic benchmark, but does not describe its evaluation procedure or results here; its role in the talk is as a possible way to assess a voice agent’s current state.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The framework diagnoses; the remedy belongs to the system
The closing distinction is practical: voice AI is a joint activity between the user and the agent, not merely a pipeline that transforms audio into text and back again. Needs must be served across multiple layers in real time. There is no universal configuration that repairs every agent; the map helps a team locate its own failure, understand adjacent dependencies, and decide what to change. Kim’s next steps are direct: evaluate the system, learn linguistics, and involve linguists in its design.
The longer-term challenge is adaptation. During any conversation, people learn one another’s accent, vocabulary, rhythm, and preferred ways of explaining things. Users will likewise adapt to a voice agent over the course of a call. A system should ask whether that learning can make the next exchange easier rather than forcing the user to rediscover the same limitations each time.
Language itself also moves. Pronunciations, vocabulary, and expectations can change over a year—or even six months. A voice agent that performs acceptably today therefore needs mechanisms for reevaluation and revision. The useful map is not a finished checklist. It is a way to keep finding where a changing agent, a changing user, and a changing language have fallen out of alignment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Watch the recording alongside its timestamped transcript and chapter navigation, including the original call example and the eight-cell linguistic framework.
Related talks
- Designing Voice Agents for Real Conversations
Extends the interaction layer with a technical comparison of voice activity detection, provider endpointing, and semantic turn classification.
- Why ChatGPT Keeps Interrupting You
Develops the premature-interruption problem through turn-taking, conversational context, and full-duplex approaches.
- Beyond Transcription: Building Voice AI That Actually Understands Conversations
Explains why transcripts alone omit speaker identity, overlap, and other interaction structure that voice systems need to understand a conversation.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Okay, hello everyone.
- 0:16
So,
- 0:19
my name is Midam Kim. I am an ML
- 0:22
engineer from ServiceNow and I'll be
- 0:25
talking about a linguistic framework for
- 0:28
voice AI.
- 0:33
So,
- 0:35
quick background of me so you know where
- 0:37
I'm coming from.
- 0:39
Like I said, I'm an ML engineer at
- 0:41
ServiceNow, but I'm also a researcher,
- 0:43
lifelong researcher, of speech
- 0:45
communication in the wild.
- 0:47
So, my motto is doing linguistics and
- 0:51
what I'm going to be doing today is to
- 0:54
hand you that lens of linguistics.
- 1:00
So, have you experienced voice AI
- 1:03
failures?
- 1:05
Yeah, like everyone.
- 1:06
>> [laughter]
- 1:10
>> So, I'm going to introduce an example
- 1:12
that I experienced myself.
- 1:15
So, the bot asked me, "Could you please
- 1:18
spell your first name?"
- 1:20
And then I slowly start to spell my
- 1:23
name.
- 1:24
Yes, it is m i d a m.
- 1:29
And the bot says, "Confirming with you,
- 1:31
is it m i d a n?"
- 1:35
And then I say, "No, it is m i d a m."
- 1:41
Um
- 1:42
and the bot says,
- 1:44
"Thank you for your correction. Happy to
- 1:46
help you today, Madam."
- 1:48
And I then I
- 1:50
get slightly annoyed, more annoyed,
- 1:51
because my name is Midam, not Madam.
- 1:55
And then it asked me about, "Now, what
- 1:57
is your account number?
- 1:59
And then, I start start getting
- 2:01
confused. What is that account number
- 2:03
thing?
- 2:04
And then,
- 2:06
I try to find uh information about that.
- 2:10
So,
- 2:11
which one? Um it must be and I start
- 2:16
uh
- 2:17
slowly start spelling the account
- 2:20
number. So, it is A X 4 5 1.
- 2:25
And then, I take time because I'm not
- 2:28
used to reading this strange number.
- 2:32
And then, the bot cuts me off.
- 2:34
And then, it says, I couldn't find your
- 2:36
record.
- 2:37
And then, without even trying, it asked
- 2:40
me to repeat that again. Can you please
- 2:42
repeat that? And then, I get super
- 2:44
annoyed and then, I can say, can I talk
- 2:46
to a person?
- 2:48
I just don't want to deal with you
- 2:49
anymore.
- 2:50
So, this is a very typical pattern of
- 2:53
voice AI, unfortunately, at this point.
- 2:56
So, I just want to navigate how we can
- 2:59
solve this problem
- 3:01
with linguistics.
- 3:06
So, voice AI is booming.
- 3:08
But users are still often preferring
- 3:10
human agents over voice agents.
- 3:13
How can we mitigate this issue?
- 3:17
But in the first place, what are the
- 3:19
actual problems?
- 3:21
So, I think we can think about a
- 3:23
fundamental frame framework to
- 3:25
understand this into an architecture of
- 3:28
voice AI,
- 3:29
which is called linguistics.
- 3:34
So, as all of us already know,
- 3:38
human communication is a joint activity,
- 3:41
like the thing that we're doing right
- 3:42
now.
- 3:43
So, I give you my sounds and words.
- 3:47
You hear them.
- 3:49
And then, if it is a conversation,
- 3:51
you're going to give me your sounds and
- 3:53
your words.
- 3:55
And then this is going back and forth
- 3:58
through interaction.
- 4:01
And then in this process, we're
- 4:03
continuously
- 4:04
processing and updating our mental
- 4:07
models.
- 4:09
So that's a joint activity
- 4:12
for human communication.
- 4:14
And I would like to say
- 4:17
in the voice AI human communication,
- 4:20
it also has to be a joint activity like
- 4:23
this.
- 4:24
Because that's the only thing that we
- 4:27
know about human communication as a
- 4:29
human being. We have been evolving
- 4:31
thousands of years as communicators, and
- 4:34
this is what we know. So we expect the
- 4:36
same thing to bots.
- 4:41
So let me go over the failure scene of
- 4:44
my call with the voice agent
- 4:47
in this framework.
- 4:49
So you see there's listen
- 4:52
and speak for each party.
- 4:57
So I start spelling my first name.
- 5:00
And then the bot did not hear that the
- 5:04
difference between M and N correctly, so
- 5:07
it's an
- 5:08
SCT failure in the listening level.
- 5:12
And then the TTS applies only
- 5:15
English-centric reading rules to my
- 5:17
name, M I D A M, would read it as Midam
- 5:21
in the
- 5:22
American English version.
- 5:25
So I'm confused, but at this time I'm
- 5:27
kind of generous because that happens a
- 5:29
lot even with human beings. So I'm okay.
- 5:34
But then when it brought
- 5:35
brought up account number thing
- 5:38
because I don't know what that is,
- 5:40
I'm confused again.
- 5:42
But I'm adaptive, I can find I can look
- 5:45
for it.
- 5:46
So I found the number, start reading it,
- 5:48
but
- 5:50
the STT did not recognize the word unit
- 5:53
correctly, so
- 5:54
it cuts me off, and uh
- 5:58
uh finally, it's uh eventually talked
- 6:02
over me.
- 6:03
So, I get
- 6:05
really irritated.
- 6:08
And then, when it asked me for the
- 6:09
repetition of the same information, and
- 6:13
then, it is clear that the spot is not
- 6:15
tracking the mental model with me.
- 6:18
And then, very rudely, it's uh does not
- 6:22
even try interactive clarification,
- 6:24
which is a common strategy by human
- 6:26
beings.
- 6:27
So, I don't want to deal with this
- 6:29
anymore, so I say, "Can I talk to a
- 6:30
person?"
- 6:34
So,
- 6:35
let's go over the uh the framework
- 6:37
again. So, the these are the linguistic
- 6:39
components that are expected and well
- 6:41
maintained in human-to-human voi- uh
- 6:45
uh conversation.
- 6:47
So, there are listening channels, a
- 6:49
listening channel and speaking channel,
- 6:50
and there are different components like
- 6:52
sounds, words, interaction, and mental
- 6:54
model.
- 6:55
So, the first component is, does the bot
- 6:58
recognize the user's speech well?
- 7:01
And all of these technical terms
- 7:04
uh will fall under this.
- 7:07
And then, there was there's going to be
- 7:08
this second component, which is words in
- 7:11
the listening channel. So, does the bot
- 7:13
understand the user's words?
- 7:17
And then, the third one is, does the bot
- 7:19
wait until the right timing to for its
- 7:22
turn? It's about It's going to be about
- 7:24
uh listening channel interaction.
- 7:28
And then, uh the last part is mental
- 7:31
model. So, does the bot understand the
- 7:33
user's intention
- 7:35
in the listening part?
- 7:38
And then, we can also go to the speaking
- 7:39
channel, so it's going to be about
- 7:41
pronunciation for the sound.
- 7:43
And also there's about understand the
- 7:46
the words users are
- 7:49
uh there's about choose the words the
- 7:51
user can understand.
- 7:53
And in the interaction part, there's
- 7:55
about speak with the right timing.
- 7:58
And lastly, there's about speak with the
- 8:01
information the user actually need.
- 8:05
So, there are a lot of engineering or
- 8:08
linguistic or cognitive science terms
- 8:10
that are in here that that are here. Um
- 8:14
you can see now see that all of those
- 8:17
have their right spots in this
- 8:18
linguistic framework.
- 8:23
And importantly, these components are
- 8:24
interdependent,
- 8:26
not separate or uh independent from each
- 8:29
other. They're interdependent and
- 8:31
they're aligned. So, when you want to do
- 8:34
good things about sounds,
- 8:37
you have to think about words level.
- 8:39
And then when you want to do good things
- 8:41
about these sounds and words,
- 8:43
you also have to uh account for
- 8:46
interaction, so turn taking or turn
- 8:48
detection.
- 8:50
And then finally, you want to uh have
- 8:53
good uh task completion, which is the
- 8:56
goal of these mental model uh layer.
- 8:59
Then you have to have all of these.
- 9:02
Without all of those, without any of
- 9:04
those, any of those components, your
- 9:06
voice agent will fail.
- 9:09
And then finally,
- 9:11
uh it has to be well aligned. All of
- 9:13
these have to be well aligned.
- 9:17
And additionally, you have to keep your
- 9:20
mind keep in mind that
- 9:22
this is happening on the timeline.
- 9:26
What I mean by that is it is silently
- 9:29
tracked. Unlike in chat, in chat you see
- 9:33
the history of what was said
- 9:35
uh as text.
- 9:37
But in voice agent experience,
- 9:40
uh, you say something, and the bot says
- 9:42
something, you go back and forth,
- 9:45
and then see, all these waveforms, the
- 9:49
air via the vibration in the, uh, in the
- 9:51
air, they're all gone.
- 9:53
And only the user's mental model is the
- 9:56
thing that's left, and that matters.
- 10:00
So, sounds, words, interactions vanish
- 10:03
the moment they're spoken,
- 10:04
but the mental model proceeds and grows
- 10:07
over the timeline.
- 10:09
So, this is what you have to
- 10:12
target
- 10:14
for user satisfaction.
- 10:16
And then, what can we do
- 10:19
for the bot to meet the standard of the
- 10:22
user?
- 10:25
So, what we can do, uh, would include,
- 10:28
of course, choosing good ASR models or
- 10:31
configurations and do some
- 10:32
post-processing,
- 10:34
uh, choosing good TTS models,
- 10:36
configurations, and pre-processing,
- 10:38
and, uh, carefully curate the vocabulary
- 10:42
that can be shared between the bot and
- 10:44
the user,
- 10:45
and do good job of a turn-to-turn
- 10:48
detection, latency, and turn-taking.
- 10:52
Um, and very importantly, we have to, it
- 10:56
would be great if we can do good emotion
- 10:57
detection and handling, and context
- 11:00
retention, and by context, what I mean
- 11:02
is context about all of these.
- 11:07
And importantly,
- 11:09
uh, it has to be dynamic because things
- 11:12
are always changing, uh, throughout over
- 11:14
the course of the call. So, we would
- 11:17
have to do this management dynamically
- 11:19
along the timeline
- 11:21
for different kinds of people.
- 11:23
So, kids or different kinds of people
- 11:26
like these will have different
- 11:28
expectations that we have to satisfy.
- 11:32
Uh, not just when they're happy, but
- 11:34
also when they're not happy.
- 11:36
So, only then you can pursue a dynamic
- 11:39
and truly scalable orchestration of
- 11:41
voice AI.
- 11:43
So, it's a very difficult job to do.
- 11:48
We always say that voice is the most
- 11:50
natural way of communication, but it is
- 11:53
not actually not easy. Behind the scene,
- 11:55
it is thanks to this linguistic
- 11:57
orchestration.
- 11:59
When your bot is not good at it,
- 12:01
it's a catastrophic failure.
- 12:06
Um, so paying attention to this
- 12:09
linguistic framework would have lots of
- 12:12
business implications because then you
- 12:14
can uh
- 12:17
decrease all of these user frustration,
- 12:19
task failures, live agent escalation, or
- 12:22
abandoned calls, or silent failures.
- 12:28
So, in ServiceNow, we have made a a good
- 12:32
uh benchmark end-to-end benchmark called
- 12:34
Eva bench. So, you can try that to
- 12:36
diagnose your voice agent's uh status.
- 12:43
Um, key takeaways.
- 12:45
So, voice AI is a joint activity between
- 12:50
the bot and the user, not just a
- 12:52
pipeline.
- 12:54
And we must serve users' needs in
- 12:55
multiple layers real time.
- 12:59
It's not that I have given you a fix
- 13:01
today because there's nothing like that.
- 13:04
It just uh the fix is in you and your
- 13:07
system.
- 13:09
But, what I have given you is today is
- 13:14
the linguistic framework you can try to
- 13:16
diagnose your system
- 13:18
and to build your system upon.
- 13:21
You can try Eva, but also you can learn
- 13:24
linguistics and hire linguists.
- 13:27
Um, another thing I want to remind you
- 13:29
of is that business implications are
- 13:32
linguistic implications and vice versa
- 13:35
in this voice AI scene. Because voice is
- 13:39
fundamentally a linguistic and very
- 13:41
human and cognitive experience.
- 13:46
I would like to ask you a longer term
- 13:48
question.
- 13:50
Speakers adapt. So, I
- 13:54
I'm pretty sure that in this talk in my
- 13:58
talk with you guys today, you have
- 14:00
learned something about me, about my
- 14:02
speaking style, what kind of accents I
- 14:04
speak, what kind of words I'm using. So,
- 14:07
next time I see you guys in person, you
- 14:10
would find it more comfortable to talk
- 14:12
to me because you have paid attention to
- 14:14
me.
- 14:15
Right? So, speakers are always adapting.
- 14:17
So, the user will be adapting to your
- 14:20
voice agent throughout the call. So, is
- 14:24
your system ready for them to
- 14:27
use you better, use it your voice agent
- 14:29
better the next time?
- 14:31
And
- 14:32
language is always change. So, is your
- 14:35
voice agent ready for language change in
- 14:38
1 year or 6 months even?
- 14:44
So, thank you.
- 14:47
>> [applause]