Act, Confirm, or Stop? Smarter behavior for AI assistants, wearables & robots — Amit Desai, Roku
Read the talk
Act, Confirm, or Stop? Smarter Behavior for AI Assistants, Wearables & Robots
Amit Desai shows how an assistant can reduce the pain of errors without improving recognition accuracy: assign a recovery cost to each possible outcome, then optimize when the system should act, ask for confirmation, or decline.
From a talk by Amit Desai
At a glance
Ideas worth remembering
Interpretation accuracy and behavior under uncertainty are separate controls. The example keeps accuracy fixed at 79 percent while changing which hypotheses the assistant executes.
Choose thresholds by minimizing weighted recovery cost, not by selecting a confidence percentage that merely feels safe.
Wrong actions, refusals, affirmations, and corrections impose different costs. Treating them as equivalent hides the product decision that matters most to users.
Adding confirmation creates a useful middle region, but it helps only when its delay and correction costs are included in the objective.
Outcome costs must reflect the actual interface and action. Visual choices can make confirmation cheaper, while disruptive or irreversible actions make incorrect execution more expensive.
Errors get more expensive when AI acts
Amit Desai, credited with Roku for this recording, starts from a tension familiar to anyone building voice interfaces. Talking is natural, but speech systems remain error-prone. That becomes more consequential as assistants move from answering questions to taking digital or physical actions. Playing the wrong song is irritating; a robot throwing away your watch is materially worse.
This creates two independent controls over the experience. The familiar control is accuracy: improve wake-word detection, speech recognition, language understanding, intent classification, entity extraction, or another part of the interpretation stack. The second is the system decision under uncertainty: given a hypothesis that might be wrong, should the assistant act, confirm it, or stop? Better decisions can reduce user pain even when the underlying interpretation remains unchanged.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Freeze accuracy at 79 percent
The worked example is a simple music speaker. Collect 1,000 spoken requests, compare each requested song with what the system selected, and label the results. In the illustrative dataset, 790 hypotheses are correct and 210 are wrong: 79 percent accuracy. Conventional model work would try to shrink those 210 errors one percentage point at a time.
Instead, keep all 1,000 interpretations exactly as they are. The baseline assistant always acts, so every hypothesis immediately becomes playback. Desai adds one alternative: reject the hypothesis and ask the user to repeat. That sounds obviously safer, but it creates a harder quantitative question—exactly when should the system stop?
For the simplified exercise, every labeled hypothesis receives one reasonably calibrated confidence score between zero and one. A threshold T implements the policy: stop when confidence C < T; otherwise play. A cascaded production system may expose several confidence signals rather than one clean score, so this is a teaching model rather than the proposed final architecture.
A threshold such as 65 percent feels plausible: perhaps the system should act only when it is “confident enough.” But confidence alone does not identify the best decision. Moving the threshold prevents some wrong plays while also rejecting some requests the system understood correctly. The right boundary depends on what those outcomes cost the user.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Count recovery effort, then minimize the OUCH
The concrete failure is a request for “Kiss” by Prince that instead starts a Chris Brown song. The observable problem is not merely an incorrect label. Music begins playing; the user listens long enough to detect the error, talks over the speaker to stop it, and makes the request again. A refusal—“Sorry, I didn’t understand”—still delays success, but it avoids the work of undoing an incorrect action.
Desai turns that difference into a recovery-cost heuristic: estimate the additional time required to reach the intended result. He assigns 10 seconds to wrong playback and four seconds to stopping and repeating. These values are illustrative rather than measured universal costs; changing them, or changing the confidence distribution, changes the optimum.
For any candidate threshold, the cost is therefore 10 × wrong acts + 4 × stops. Correct immediate playback contributes no recovery cost. Raising the threshold generally reduces wrong actions but creates more refusals, including refusals of correct hypotheses. The weighted sum makes that tradeoff explicit instead of hiding it inside an intuitive confidence cutoff.
Desai names the objective the Outcome User Cost Heuristic, or OUCH—a language nerd’s cost function for pain. With the original always-act policy, 210 wrong plays cost 2,100 points across 1,000 requests, or 2.1 per turn. A guessed 65 percent stop threshold lowers that to 1,904, or about 1.9 per turn. Searching the demonstrated curve finds a much better threshold at 43 percent, reducing the average to approximately 1.27 OUCH points per turn.
Nothing about the 790 correct and 210 incorrect interpretations has changed. The improvement comes entirely from routing uncertainty into a less costly behavior. Accuracy describes whether the hypothesis is right; the decision policy determines how much damage a possibly wrong hypothesis can cause.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Add a confirmation region between stopping and acting
The next behavior is confirmation. Instead of immediately playing or refusing, the speaker can state its interpretation—such as “Play ‘Kiss’ by Prince?”—and let the user affirm or correct it. This splits confidence into three regions governed by two thresholds: stop at low confidence, confirm in the middle, and act at high confidence.
What does that three-way policy make visible? Low-confidence hypotheses are rejected before they cause damage, middling hypotheses pay the smaller cost of a question, and high-confidence hypotheses preserve the fast path. The two boundaries jointly determine how many correct and incorrect requests land in each behavior.
The recognizer produces one proposed song and a simplified calibrated confidence score.
Two optimized boundaries divide the example’s confidence range into stop, confirm, and act regions. The policy changes behavior, not whether the original hypothesis was correct.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Confirmation helps only when its own cost is modeled
Confirmation is not free. Even a correct read-back delays playback, while an incorrect proposal forces the user to reject it and restate the request. The example assigns two seconds to affirmation and six seconds to correction, producing the extended objective 10 × wrong acts + 4 × stops + 2 × affirmations + 6 × corrections. Because two thresholds must now be selected together, the cost landscape becomes a two-dimensional heat map.
For the demonstrated data and costs, the optimized boundaries are 41 and 49 percent. The recorded display distinguishes a current 30/60 configuration costing 1,464 in total from the optimized 41/49 configuration costing 1,260. The recap expresses the progression as 2.100 OUCH points per turn for always acting, 1.904 for the guessed stop threshold, 1.274 for optimized act-or-stop behavior, and 1.260 after optimized confirmation. Confirmation improves the result again, although much less dramatically than introducing and optimizing the stop behavior.
The thresholds are not reusable constants. If a wrong action costs 20 instead of 10, the optimum moves because preventing incorrect execution becomes more valuable. It also moves when the confidence distributions change. OUCH is therefore a way to derive a policy from a system’s observed outcomes and product-specific costs, not a recommendation that every assistant act above 49 percent confidence.
A production assistant would not necessarily use one or two offline scalar thresholds. Desai expects a learned real-time decision model that can consider richer signals. The simplified exercise supplies the objective: choose behavior under uncertainty according to expected user cost while keeping the interpretation model and the behavioral policy conceptually separate.
A television shows why the interface changes the equation. For “open a channel,” confirmation can present several visual choices instead of speaking one candidate. Selecting the right option with a remote may be cheaper than a spoken correction, so confirmation’s cost falls. But launching the wrong channel may eject the viewer from their current state, making an incorrect act more expensive.
The ending returns to higher-stakes actions: making a phone call, sending an email, or moving an object. As consequences rise, accuracy improvements alone become an increasingly narrow response to uncertainty. The assistant also needs a policy for when to proceed, seek permission, or decline. If that interface remains painful, it becomes the choke point even while the models behind it improve. Desai’s memorable instruction is the pun that carries the whole mechanism: minimize the OUCH of the experience.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Watch the recording alongside its chapter markers, corrected transcript, numeric clarification, and discussion prompts.
Related talks
- Building Effective Voice Agents
Extends the discussion from decision policies into production voice architecture, including latency, determinism, tool use, handoffs, and evaluation.
- Designing Voice Agents for Real Conversations
Explains another behavior-under-uncertainty problem: deciding whether a pause means the user has finished speaking.
- Bounded Autonomy: Between Free Will and Determinism
Develops the adjacent idea that agent autonomy should be constrained by context, curated knowledge, and short real-world feedback loops.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
Hi everyone. How's it going? Hey
- 0:15
Patricia, how are you?
- 0:17
>> Uh so last presentation of the day, so
- 0:20
let's make it count. Um,
- 0:23
all right. Let's, uh, let me start with
- 0:25
a little bit of background on myself.
- 0:28
And, um, my background, I'm a voice
- 0:33
subject matter expert. I've been working
- 0:35
in voice AI for a long time across
- 0:37
different surfaces, devices, and um,
- 0:42
both at Alexa, at at Roku, at my own
- 0:46
startups, you know, in the app store.
- 0:49
And my perspective is a little different
- 0:52
from a lot of other voice AI
- 0:55
practitioners. I think it's a
- 0:58
combination of um a deep um voice user
- 1:02
interface expertise and intu intuition
- 1:05
mixed in with new technical approaches
- 1:09
uh that I think can produce really
- 1:11
magical experiences. So I think it's
- 1:12
both sides and I think that's especially
- 1:15
true in this new area that we're in with
- 1:17
frontier tech where the human interface
- 1:20
is basically being redefined. So let me
- 1:24
start with uh I'll just blast through
- 1:26
the first couple of slides then get to
- 1:28
the premise. I think everybody knows
- 1:30
that voice has incredible potential.
- 1:32
There's the power of voice I think
- 1:35
across everywhere. It's the most natural
- 1:37
interface. Humans love talking. And uh
- 1:40
the problem is the other half is the
- 1:43
pain of voice. So it's the power and the
- 1:44
pain. Voice is errorprone. And I think
- 1:49
those errors are going to continue for a
- 1:52
while. And I think the cost or
- 1:54
consequence of those errors is going to
- 1:56
grow, especially as we go fromational
- 1:59
AI bots to embodied AI where rather than
- 2:03
just giving answers that might be
- 2:05
erroneous,
- 2:07
we're going to have AI systems take
- 2:09
physical actions or digital actions
- 2:12
where, you know, if the robot throws
- 2:15
your watch out with the trash, it's a
- 2:17
lot worse than playing the wrong song.
- 2:19
So I do think that a new approach is
- 2:23
definitely needed and here's the TLDDR
- 2:26
of the premise we're going to walk
- 2:28
through today. Um there are two ways to
- 2:31
improve customer or user satisfaction of
- 2:35
a voice AI assistant and that is by
- 2:38
increasing accuracy which people know
- 2:39
about I mean technically accuracy and
- 2:43
the other is a different knob that we
- 2:46
have that we are not using adequately
- 2:48
and I'll call that a system decision
- 2:51
which we will define which is orthogonal
- 2:54
which is different from accuracy and I
- 2:57
believe This approach which I have used
- 3:01
in several different environments and
- 3:04
seen some success I think is a promising
- 3:07
area that we should consider developing.
- 3:10
Um let me walk through this with a
- 3:13
simple smart speaker example and we'll
- 3:16
go step by step with this approach but
- 3:19
it is a scalable approach that I think
- 3:22
uh can apply across different surfaces
- 3:24
and devices. So let's get started. So
- 3:27
suppose we all you know are making a
- 3:30
smart speaker coincidentally called uh
- 3:33
Alexa and Alexa is very simple. It just
- 3:37
allows you to you know ask for music and
- 3:40
it'll play a song and of course it will
- 3:43
play either the song you wanted or a
- 3:45
different song. So it'll be right or
- 3:47
it'll be wrong. This isn't that
- 3:49
different from what you've seen out
- 3:51
there. Um now let's to first talk about
- 3:54
accuracy. Accuracy. Let's say we define
- 3:57
it as we you know take a thousand spoken
- 3:59
requests. We observe the input and the
- 4:02
output. We label it and we look at this.
- 4:05
This is the map of a thousand points and
- 4:08
79% of the time 790 dots here were
- 4:12
actually the correct song. This is let's
- 4:14
say human annotated 20% 21% wrong song.
- 4:19
So that's the accuracy. Now, like I
- 4:22
said, knob one is to spend a lot of time
- 4:25
working on improving the accuracy, you
- 4:28
know, um, percentage point by percentage
- 4:30
point at any layer in the stack.
- 4:33
There's, if it's a cascaded system, you
- 4:35
know, there's a perhaps a wakeword layer
- 4:38
and a speech ASR layer and a NLU layer
- 4:41
which might have intent classification,
- 4:44
entity extraction, a lot of different
- 4:45
layers, VAD, etc. And any of those can
- 4:47
contribute to errors. So we spent time
- 4:50
we might be able to reduce that 210 to a
- 4:52
smaller number that is I think a known
- 4:56
area that we're tackling but I think
- 4:59
knob 2 which is what I was talking about
- 5:01
is what we'll go through here which is
- 5:03
keeping the accuracy exactly the same.
- 5:06
So 79% what could we do in conditions of
- 5:10
uncertainty to improve user satisfaction
- 5:13
apparent and I I think we can do a lot.
- 5:15
So let's start first with the original
- 5:17
system is just acting like I said user
- 5:20
says something system plays a song it's
- 5:21
either the right song or the wrong song
- 5:24
immediately I think just common sense
- 5:26
tells us that we could introduce at
- 5:28
least one system behavior to stop or
- 5:30
rather to reject the hypothesis and do
- 5:33
nothing. So uh there is now one more
- 5:36
option to decide the system may decide
- 5:38
and say sorry I didn't get that or sorry
- 5:41
could you repeat that? Uh the challenge
- 5:43
of course is how how when do we decide
- 5:47
to stop and I mean quantitatively. Um
- 5:51
here's one approach to kind of
- 5:53
visualizing this because if we don't
- 5:55
we'll just take probably some swag like
- 5:57
some guesstimate and I'll prove that if
- 5:59
we just took a guesstimate we would end
- 6:02
up with a worse situation than a more
- 6:04
rigorous approach. So let's just assume
- 6:06
I took those thousand data points and
- 6:08
like I said they've been annotated and
- 6:10
we assign a confidence score a single
- 6:13
confidence score to the hypothesis that
- 6:16
was generated by the system you know
- 6:18
between zero and one and let's say it's
- 6:19
reasonably calibrated. This is a
- 6:21
simplification of if it's a cascaded
- 6:23
system there are multiple layers and
- 6:24
multiple you know confidence scores but
- 6:26
let's just assume that for now. Whoops.
- 6:28
So we're going to have 790 points 200
- 6:31
that are correct 210 wrong. Each one has
- 6:34
a confidence score and we're going to
- 6:36
plot it, you know, plot the
- 6:37
distributions. Uh on the x-axis, I've
- 6:40
just converted from 0ero to one to
- 6:42
percentages. And the question is how do
- 6:46
we choose a threshold t such that
- 6:49
whatever that percentage is um to the
- 6:52
left of it meaning if when the system um
- 6:55
forms a hypothesis if the confidence
- 6:58
score c is less than that t stop and say
- 7:01
sorry otherwise play question is how do
- 7:04
we choose a t so far everything I'm
- 7:06
saying is fairly common sensical but
- 7:08
this is where um intuition will fail us
- 7:12
we might say something like okay I don't
- 7:14
know let's do 65%. It seems you know gut
- 7:17
feeling like okay it's kind of confident
- 7:19
that's probably when we should speak. Um
- 7:22
now here's where we start coming out
- 7:24
with some sophistication.
- 7:26
Any tea we choose is producing bad
- 7:29
outcomes. Bad in the in two fields. One
- 7:33
is obviously on the left side anytime
- 7:35
you stop it's bad. The user doesn't want
- 7:38
it to stop. He wants to they want to
- 7:40
hear their song. The other bad is if you
- 7:43
do play a wrong song, of course that's
- 7:46
bad as well. So these are two two kinds
- 7:48
of bad outcomes. But here's the
- 7:51
important part. Now I've like elaborated
- 7:53
on the um tree diagram on the right hand
- 7:56
side. The bad outcomes are not equally
- 8:00
bad. They're not the same thing from a
- 8:02
user perspective. And obviously let's
- 8:05
let's think about it. If the wrong song
- 8:07
plays, you said play kiss and it starts
- 8:11
playing kiss by Chris Brown instead of
- 8:14
the one by Prince. That's going to be um
- 8:18
the highest user cost. Now I'm defining
- 8:21
user cost from the user's perspective.
- 8:23
First I have to like hear music and
- 8:25
realize that is not Prince. Then I have
- 8:27
to shout over my Alexa and um you know
- 8:31
get it to stop and then I have to
- 8:33
re-request. All of that is a lot of
- 8:35
effort. that is definitely a worse
- 8:37
outcome than the system stopping and
- 8:39
saying sorry I didn't understand that
- 8:42
however we should go further and try to
- 8:45
quantify that relative badness and there
- 8:47
many ways to do it and I think this is
- 8:49
an area to be explored for now let's
- 8:51
just consider this a heristic of if that
- 8:55
outcome happens how many more seconds
- 8:58
additional seconds will it take for the
- 8:59
user to get back to success which is to
- 9:01
play the song they wanted kiss by Prince
- 9:04
and I'm I just put down some numbers
- 9:06
here. Let's say in the case of a bad
- 9:08
song, it's 10 seconds if you add up all
- 9:10
the things I got to do. And if it's a I
- 9:14
didn't understand you, it's 4 seconds
- 9:15
because that's how long it would take
- 9:16
you to respe and and the extra latency.
- 9:19
And now here's where we can start
- 9:23
utilizing that. If we go back to our
- 9:25
distribution curve on trying to find out
- 9:27
where is T. Now we've basically turned
- 9:31
this into a problem of minimizing a cost
- 9:34
function. It's a user cost function. It
- 9:36
is the number of bad acts wherever that
- 9:38
whatever the t causes times 10 because
- 9:40
that was a unit cost we gave plus the
- 9:43
number of stops times four because
- 9:45
that's the the unit cost we gave. By the
- 9:48
way, one thing I should have elaborated
- 9:51
because I work in voice and we like
- 9:52
language and we like puns. So this whole
- 9:55
thing is called an outcome user cost
- 9:57
heruristic. So that spells the word ouch
- 10:00
and that is some expression of pain.
- 10:03
Yes, we are you know language nerds. So
- 10:06
these kinds of things amuse us. Um so
- 10:08
now let's consider that is the cost
- 10:10
function is to minimize the ouch. And
- 10:12
now um that let's see if uh I'm going to
- 10:16
bring up a tool. Let's see if this
- 10:18
works.
- 10:20
Where I have actually gotten or with one
- 10:23
of my coding assistants gotten uh an
- 10:26
interactive
- 10:29
um graph where we have actually plotted
- 10:32
those thousand points and as we vary the
- 10:36
threshold t you can see that the total
- 10:40
user cost here which is that function of
- 10:43
you know x * y + a * b actually changes.
- 10:46
So let's in the very beginning when we
- 10:49
said the system was just playing
- 10:52
the the cost across those thousand
- 10:54
points was 2100 or divided by a,000 is
- 10:57
2.1 ouch points per turn. Then we said
- 11:01
okay let's insert a stop behavior and
- 11:04
let's like wing it and say 65%. That's
- 11:07
when I want the threshold. If we brought
- 11:10
this up to 65 yeah that's better. Now
- 11:13
it's 1904 or 1.9 per turn, but it's not
- 11:16
optimal. As it turns out, if we do
- 11:19
actually um ask for the AI to solve the
- 11:23
uh the problem across this curve, it
- 11:25
turns out 43%. So I'll drag it now to 43
- 11:30
is in fact
- 11:34
the optimal
- 11:36
optimal point of t. This minimizes the
- 11:39
cost function. You can see it's the
- 11:41
lowest point on this graph down here to
- 11:43
1
- 11:44
27. So effectively we haven't changed
- 11:48
the accuracy at all. The system is not
- 11:50
any smarter in that sense. But with some
- 11:52
clever system behavior, conversational
- 11:55
behavior is what we'd call it and some
- 11:56
optimization and a cost function called
- 11:59
ouch. Um we have from the user's
- 12:02
perspective produced a more satisfactory
- 12:06
assistant. And this is not a trivial you
- 12:08
know accomplishment. Okay. Now, let me
- 12:10
go back to this. [clears throat] Let me
- 12:12
see if I can get this. Oh, great. Okay,
- 12:16
let's continue this. Let's continue this
- 12:19
with by now adding one more behavior.
- 12:22
Let's call it the confirm behavior. So,
- 12:23
there was play obviously, then stop,
- 12:26
confirm. Confirm is basically the system
- 12:28
after you said something saying uh kiss
- 12:32
play kiss by Prince or maybe play kiss
- 12:35
by Chris Brown. And uh you know the user
- 12:38
can either confirm like affirm it or
- 12:40
they can correct it. It is a different
- 12:42
kind of behavior and again this is kind
- 12:44
of how humans behave. Um that's
- 12:46
obviously the inspiration. Now if we go
- 12:49
back to our problem of optimization,
- 12:52
we have a third obviously um option
- 12:55
which is to confirm. And so this would
- 12:58
translate to two thresholds
- 13:00
um two thresholds which are separating
- 13:03
the distribution into three spaces of
- 13:07
stop, confirm and uh act. And the
- 13:12
question is now where are these T's? and
- 13:15
we have now given up on guesstimating
- 13:16
because we know it doesn't work. So
- 13:18
we're going to be a lot smarter and go
- 13:21
back to the concept of user outcome cost
- 13:25
and then you know use it go look for
- 13:27
some optimization in that graph. So
- 13:29
let's uh define what are the what are
- 13:32
all the possible bad outcomes that t1
- 13:34
and t2 um make for. So good you can see
- 13:38
my cursor. So uh of course any stops are
- 13:42
still bad. Then in the middle are
- 13:45
confirmations. Confirmations are bad
- 13:47
because they slow the user down. There
- 13:49
is a confirmation outcome called confirm
- 13:52
yes where they just affirmed it by
- 13:54
saying yeah or no where they had to
- 13:56
correct it. And going back to our
- 13:59
formula these outcomes are not equally
- 14:02
bad. And in fact, nobody will, I think,
- 14:06
argue here from a user's perspective.
- 14:08
Affirming, just saying yes is obviously
- 14:11
less painful than saying no and then
- 14:13
having to restate whatever it is that
- 14:15
you wanted in the first place. So now we
- 14:17
I've assigned values of two or six. And
- 14:19
again, I said it was a heristic. This
- 14:21
would be roughly the amount of time it
- 14:23
would take for the extra for the user to
- 14:25
get to the song they want. Saying
- 14:27
listening and then saying yes is like
- 14:28
two seconds. Um and then now
- 14:34
uh we restate the cost function for this
- 14:37
you know added behavior as this number
- 14:40
of you know bad type one times unit cost
- 14:43
bad type plus bad type two times unit
- 14:45
cost etc. And now we try to minimize
- 14:49
this user cost function and minimize the
- 14:52
ouch.
- 14:53
Yes, I'm going to keep doing that pun.
- 14:56
Um let's go back. So this is now the
- 15:00
interactive graph but
- 15:03
with
- 15:06
um the cost values the unit costs here
- 15:08
10264
- 15:10
and uh you know we're just going to ask
- 15:13
the AI to tell us here's the heat map
- 15:16
because it's now two dimensions saying
- 15:18
that the optimal values are 41 for the
- 15:22
the T1 and 49 for the T2 and if we
- 15:27
employed that then we would go to 1464.
- 15:31
Uh, by the way, whatever numbers I put
- 15:33
in here, like let's say I thought wrong
- 15:36
act was 20. It's really irritating and
- 15:39
painful and takes way longer to actually
- 15:42
correct it when you hear a wrong song.
- 15:44
That would change you know all these
- 15:45
numbers uh and the optim optimal point.
- 15:48
So again it is about how what is the
- 15:50
relative badness of these outcomes also
- 15:52
of course the distribution curves
- 15:54
naturally. Uh let's go back here. Okay.
- 15:58
So, um I'm gonna
- 16:02
speed up a little bit. Uh let's go back
- 16:06
here.
- 16:07
Presentation mode. Okay. So, what have
- 16:11
we shown that if we did the super naive
- 16:13
approach, it's 2.1 act and stop 1.9 then
- 16:19
1.27 then 1.26. We are able to bring
- 16:22
this with every added layer of
- 16:25
sophistication, adding more behaviors,
- 16:27
being smart about outcome, uh, user cost
- 16:30
and optimizing. Um, we have made a
- 16:33
tremendous difference without changing
- 16:35
the accuracy at all. Um, this was a
- 16:38
super simplified example. In real
- 16:40
systems, you're not going to have
- 16:42
obviously some offline decision
- 16:44
threshold or two. It's going to be a
- 16:46
real time, you know, learned decision
- 16:48
model. But the principle is the same.
- 16:50
And I believe this is uh scalable across
- 16:54
all voice AI surfaces. Obviously this is
- 16:56
a smart speaker but if we go across any
- 17:01
of these surfaces you will find the
- 17:02
equivalence. If we um we will find the
- 17:07
analogies with some differences but the
- 17:09
spirit and the I think the the gain will
- 17:13
be similar. So just for example in the
- 17:16
TV AI assistant space if you employ it
- 17:20
here it's going to you're going to have
- 17:22
the same thing when users express
- 17:24
intents like on TV it's you know open a
- 17:27
channel that's one of the most common
- 17:29
obviously um requests on a TV voice
- 17:32
assistant same thing you're going to
- 17:34
find you'll have exactly the same
- 17:35
approach but the difference will be
- 17:38
maybe in the the assignments of the user
- 17:42
outcomes because the UI and the
- 17:43
modalities are different when you have a
- 17:46
TV you have a multimodal interface where
- 17:48
choices can be shown. So instead of you
- 17:51
know asking did you mean ABC you know uh
- 17:55
news live by speech that you will the
- 17:59
system would display choices and not
- 18:01
just one it show ABC News live this that
- 18:03
would be the confirm step and if it's
- 18:05
visual and you can use your remote
- 18:07
control to select something it's less
- 18:10
pain so you would change some of these
- 18:12
values or if in fact launching the
- 18:15
channel would kick you out of your
- 18:16
current state then it would go in the
- 18:18
other direction than cost of you know a
- 18:21
bad act would go much higher. So it's
- 18:24
the same concept but in this new
- 18:26
modalities
- 18:27
um variables can change, values can
- 18:30
change, arguments can change but the
- 18:32
premise still holds and you can improve
- 18:35
from the user's perspective because
- 18:37
we're all about you know making humans
- 18:39
happy. Um you can make them happier and
- 18:44
this as I said in conclusion can be
- 18:46
applied across all surfaces. I did say
- 18:49
at the very beginning, just to recap for
- 18:52
us, that voice is great when it works,
- 18:55
bad when it doesn't. And as we get into
- 18:58
embodied AI, where these AI assistants
- 19:01
are taking actions, physical or even
- 19:04
digital, like making a phone call or
- 19:06
sending an email, it is getting more and
- 19:09
more difficult just to rely on accuracy
- 19:12
to improve user satisfaction. I believe
- 19:15
there's a whole knob the second knob
- 19:17
called smarter conversational behavior
- 19:19
under uncertainty
- 19:21
and um if we actually exploit that we
- 19:25
can uh very much help these AI systems
- 19:30
reach a acceptable user experience
- 19:34
otherwise I think this will continue to
- 19:36
be a bottleneck like a lot of things
- 19:38
will get better but if the voice
- 19:40
interface as experienced by user does
- 19:43
not improve it is going to be a a a
- 19:46
choke point. And um if you just remember
- 19:50
one word or two words from this whole um
- 19:54
presentation, it would be to minimize
- 19:57
the ouch of the experience. Um so thank
- 20:00
you. I'll stick around for questions if
- 20:03
you guys got any. Thanks a lot.
- 20:06
[applause]
- 20:21
>> [music]