Voice Agents Can Just Do Things — Charlie Guo, OpenAI
Read the talk
Voice Agents Can Just Do Things
Charlie Guo of OpenAI reframes voice as an input and attention channel, not a requirement to answer aloud: an agent can converse, invoke application tools, update the interface, or speak only when an event deserves interruption.
From a talk by Charlie Guo
At a glance
Ideas worth remembering
Speech input does not require speech output. A voice request can produce conversation, tool execution, or visible interface feedback.
Existing application verbs—such as API endpoints and React hooks—offer a practical route to voice control, but callable tools still require guardrails and safety checks.
Use event-driven speech selectively. Animation and popups can handle lower-priority events; audio belongs higher in the escalation path because it interrupts attention.
Native audio preserves acoustic and timing context that transcription can discard, while reasoning and tool calls still add latency.
GPT Realtime-2 adds reasoning and parallel tool calls to audio interactions in Guo’s account; preambles explain the resulting wait rather than removing it.
Design voice by deciding what the model perceives, which actions it may take, when it should communicate, and whether the response should be audible or visual.
A voice interface does not have to answer with voice
Charlie Guo, who works on developer experience at OpenAI, opens with a deceptively simple correction: a voice agent does not have to talk back. Speech can be the input while the response appears as an action or a visual change. Once input and output are separated, “voice agent” stops describing one conversational interface and starts describing a wider design space.
That space has three modes: speech-to-speech, where the user talks and the model replies aloud; speech-to-action, where spoken intent leads to tool use; and event-to-speech, where an event causes the system to speak. None is fundamentally new. The Moviefone hotline loosely fits speech-to-speech, while spoken GPS directions fit event-to-speech. What has changed is how much more these patterns can do and how freely products can combine them.
What can happen after audio enters a product? The comparison makes the central design choice visible: the input channel does not determine the output channel, and one product can move among all three paths.
The user talks; the model answers aloud.
Voice may begin the interaction, carry the response, or appear later as an attention mechanism. The modes describe transitions, not mutually exclusive product categories.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Conversation still has clear jobs
Speech-to-speech remains the familiar loop, and its strongest uses depend on information carried by delivery as well as words.
- Practice and coaching: A language-learning system can hear emphasis or emotion and respond with spoken feedback.
- Concierge support: A richer conversational agent can replace the rigid experience of navigating a phone tree. Guo hopes such support can become more pleasant than an ordinary support interaction, though the talk provides no measured comparison.
- Live translation: Fast speech processing can translate content as it is delivered, suggesting events that stream simultaneous dubbing in several languages.
These cases explain why spoken output sometimes matters: pronunciation feedback must be heard, customer support benefits from conversational pacing, and translation needs to preserve the flow of a live exchange. The mistake is making speech the response by default even when the user primarily needs work completed or state changed.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Speech-to-action turns intent into work
The underexplored mode is speech-to-action: the user speaks and the model uses tools. Guo calls voice a possible “capability overhang” because many applications already contain useful operations but do not expose them through speech. Three examples widen the idea from a mundane workflow to control of an entire computer.
- Form filling: Instead of manually entering repeated names and addresses, a user could describe the information aloud. Guo imagines replacing an hour of government paperwork with five minutes of speech, having the system complete 90% of the form, and then reviewing the result. Those quantities describe a desired workflow, not a demonstrated measurement. The review step remains essential because spoken input does not guarantee that every resulting field is correct.
- Creative tools: Someone may know what feels wrong in a piece of music or an image without knowing Ableton or Photoshop well enough to make the edit. Voice can help steer the software when “taste exceeds capability,” although it cannot eliminate the difficulty of articulating an aesthetic.
- General computer use: If a model can operate applications effectively, voice control need not stop at one app or terminal. The longer-term interface could address the computer as a whole; this remains a conditional direction rather than an established claim of complete human-level computer control.
The practical on-ramp is the software developers already maintain. A modern web application exposes verbs through API endpoints and React hooks. Turning selected verbs into model-callable tools gives voice input a path into existing behavior: the model interprets the request, chooses a tool, and the application performs the operation. Exposing a tool does not make it safe by itself; Guo explicitly retains the need for guardrails and safety checks while leaving their implementation outside this talk.
The response can stay visual. A successful tool call might populate fields, highlight affected text, change a button color, add a drop shadow, show a notification, or move a ghost cursor through the interface. These established UI signals often communicate progress more precisely—and less intrusively—than a spoken narration of every click.
How would the government-form example visibly change? The flow below follows one request from speech to populated fields and makes the human review point explicit.
Names, addresses, and other repeated details are supplied through speech.
Voice supplies intent and data; application tools perform the work; the interface shows the result; the user remains responsible for checking it.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Speech belongs near the top of an attention ladder
Event-to-speech reverses the direction: an event arrives, and the model talks to the user. Guo considers this the least settled mode; a clearly AI-native pattern has not yet emerged. Two situations already justify it. In hands-free or screen-free contexts, the user cannot operate the interface or is attending to something else, as when cooking with a recipe app. In proactive outreach, the system needs to deliver information even though the user is not watching the screen.
Proactive speech must be selective. Applications produce endless events, but speaking every log would be intolerable. Guo instead places voice at the top of an escalation ladder: first animate an element, then show a popup, and speak only if the situation still needs the user’s attention. Audio is powerful precisely because it can interrupt; that makes restraint part of the mechanism, not a cosmetic preference.
Accessibility gives both action and speech a purpose beyond convenience. Guo describes developers who lost hand mobility or finger dexterity and later used language models, coding agents, and voice agents to produce far more code. The reported “orders of magnitude” improvement is anecdotal and lacks a defined measurement period, so it should not be generalized as a performance estimate. The concrete benefit is still clear: moving control away from repetitive manual input can let someone continue programming.
The modes become more useful when combined. In a car, a spoken request can start music while a navigation event later triggers a warning about traffic and rerouting. A game character could converse about the world, execute actions for the player, and react aloud to world events. The product is not choosing one voice architecture forever; it is choosing the appropriate transition at each moment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Native audio removes the text bottleneck
Why revisit these old interaction patterns now? Traditional voice agents use a chain: speech is transcribed, text goes to a language model, the model may call tools and writes a text response, and text-to-speech converts that response back into audio. Each stage adds work, and transcription reduces the incoming signal to words.
OpenAI’s Realtime model family instead operates on native audio tokens: audio goes in and audio comes out, with continuous streaming replacing a rigid sequence of conversational turns. Native audio and continuous streaming solve different problems. The former preserves acoustic information for the model; the latter lets the exchange unfold without waiting for a complete turn at every boundary.
A transcript can omit tone, cadence, emotional force, attempted interruption, and background sound. Those cues can affect whether the system should answer, wait, clarify, or recognize urgency. Guo also says OpenAI’s earlier chained voice modes had significantly higher latency than its native approach, but the talk gives no latency values or task conditions, so it does not establish an expected response time for a particular application.
What architectural work disappears in the native path? The comparison shows why fewer representation changes can preserve more of the signal and reduce the number of serial stages before a reply.
The user’s audio begins the chained pipeline.
The chained design converts audio to text and back again. The native design keeps audio as the model’s input and output representation while streaming continuously.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
GPT Realtime-2: reasoning and tools add latency
Guo introduces GPT Realtime-2 as the latest model in OpenAI’s Realtime family and describes its central change as bringing reasoning to audio: it can spend additional reasoning effort before speaking. In his account, that offers a way to improve answers that would suffer from an immediate response, much as a text reasoning model can think before producing its final output.
Guo also attributes tool calling and parallel tool delegation to the model. These capabilities let a voice interaction check external information rather than improvise immediately, but they create a direct tradeoff: more reasoning and more tool work may improve the result while making the user wait.
A preamble fills the otherwise confusing silence. Before reasoning or calling tools, the model can say what it is about to do. The travel-agent example is ordinary and effective: it can explain that it is checking flight prices and ask for a couple of seconds, then perform the calls in the background. The preamble does not reduce the work or guarantee success. It tells the user why the agent has paused and what progress to expect.
Guo further describes GPT Realtime-2 as having longer context, better domain understanding, more natural voices, and greater steerability, including prompting it to wait until it hears an assigned name. He also says it performs well on recent audio benchmarks. These are capability claims in the presentation; no benchmark scores or wake-word reliability rates are supplied.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Start with the role of voice, not the label “voice agent”
The closing question is not “What kind of voice agent should I build?” It is “What role do voice and audio play in this interaction?” That wording forces the design back onto the user’s situation rather than the novelty of the medium.
The answer requires four decisions:
- Perception: What can the model hear or otherwise sense, and which context does it receive?
- Action: Which tools are available, and which operations should it execute safely and correctly?
- Timing: Should it respond now, wait, or continue working in the background?
- Feedback: Does the user need speech, a visual notification, an interface state change, or some combination?
Guo closes with the belief that advanced intelligence will be spoken rather than typed. The more immediately useful conclusion is narrower: voice is one component of an interaction system. Treat it as an input, an action trigger, an accessibility mechanism, or a high-priority alert according to the job at hand—and let the software answer in the medium that communicates the result best.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Watch the recording with chapter links and the complete timestamped transcript for checking examples in their original sequence.
Related talks
- Building voice agents with OpenAI
A hands-on companion covering browser-based realtime voice agents, tool calls, approval, interruption handling, handoffs, and guardrails.
- Building Effective Voice Agents
Extends the architectural comparison between chained and native speech systems and examines production tradeoffs around latency, accuracy, determinism, and telephony.
- Designing Voice Agents for Real Conversations
Explains the turn-taking, voice-activity detection, interruption, and latency choices that determine whether a spoken interaction feels natural.
Read the complete timestamped transcript
- 0:01
[music]
- 0:13
So, my name is Charlie uh and I work on
- 0:16
the developer experience team at OpenAI.
- 0:18
And part of my job is talking to
- 0:21
developers to understand and see, you
- 0:24
know, what and how they're building with
- 0:26
our models. um whether that's text,
- 0:28
image or audio. And lately I've been
- 0:32
thinking about a misconception that I
- 0:35
have seen or maybe it's just a
- 0:37
misunderstanding
- 0:39
and it's the idea that voice agents
- 0:43
have to talk back.
- 0:48
>> And to some of you that might sound, you
- 0:50
know, absurd. It's a voice agent. What
- 0:52
do you mean it's not supposed to talk?
- 0:54
Uh, but I think if there's one thing
- 0:56
that you take away from this
- 0:59
presentation, I would like it to be the
- 1:02
idea that speech is not the only way
- 1:05
that a voice model has to respond.
- 1:12
And I think models are getting
- 1:13
intelligent enough and capable enough
- 1:16
that they're starting to open up some uh
- 1:20
new modes of design. I mean there's
- 1:22
actually three kind of modes that that
- 1:24
you know I kind of see emerging these
- 1:25
days right uh speech to speech speech to
- 1:29
action and event to speech and there's a
- 1:32
couple of things I think worth pointing
- 1:34
out about um these three categories. The
- 1:37
first is that they're not new, right? I
- 1:41
think as we've seen from from previous
- 1:43
talks, even just today, um there's a
- 1:45
long history of building these types of
- 1:48
systems in and around voice. Um if you
- 1:50
squint, you could make the argument that
- 1:52
the movie phone hotline where you called
- 1:55
in to get showtimes was an example of a
- 1:57
speech-to-pech system. Uh and I think
- 1:59
you could pretty reasonably make the
- 2:01
argument that you know GPS navigation in
- 2:03
your car which has existed since I was a
- 2:05
kid is an example of an event to speech
- 2:08
system. So, it's not that they are brand
- 2:11
new, but I think it is that we are able
- 2:12
to do some some much more interesting
- 2:14
things with them uh now that uh we're in
- 2:17
this era, right? And I think, you know,
- 2:19
the other thing I would mention here is
- 2:20
that um they're they're remixable,
- 2:24
right? They're not meant to be mutually
- 2:25
exclusive. Um and I think as we'll see
- 2:27
in a little bit, the best products uh
- 2:29
exist in a way that combines all of
- 2:31
these modes. So, speechto um everybody
- 2:34
knows it. Hopefully, everybody loves it.
- 2:36
the user talks uh and then the model
- 2:38
talks back and I think there are a few
- 2:41
examples that that I can give um for
- 2:43
this type of use case right you've got
- 2:46
things like live practice and coaching
- 2:48
especially around language learning
- 2:50
right I think the ability to um hear a
- 2:53
lot of emphasis or emotion and give
- 2:55
people that feedback is really powerful
- 2:57
um I think you can have you know what I
- 2:58
am sort of cheekily calling concierge
- 3:00
experiences which I think is just
- 3:02
another way to say customer support
- 3:04
plus+ Um the first voice tutorial that
- 3:08
most people you know try to build when
- 3:09
they have access to this technology is
- 3:11
some sort of customer support chatbot
- 3:13
for very good reasons. But I think with
- 3:15
you know when we add a richness and a
- 3:18
depth to the voice models as has been
- 3:19
happening in recent months in recent
- 3:21
years um we can build something that is
- 3:23
like much more enjoyable to use than
- 3:25
like talking your way through a phone
- 3:27
tree. Um, and so, you know, I have the
- 3:30
hope that like soon, if not like, you
- 3:32
know, now, we're capable of building
- 3:34
support experiences with agents that
- 3:36
actually feel much more enjoyable to
- 3:38
talk to than like arguably like the
- 3:40
median human support agent.
- 3:42
Um, and as we just saw if you hear the
- 3:45
last talk, uh, live translation, right?
- 3:47
The models have gotten good enough and
- 3:48
fast enough that we can just dynamically
- 3:51
translate content on the fly, um, with
- 3:53
like little to no latency. It wouldn't
- 3:55
shock me if at next year's keynote um
- 3:57
you know they live streamed it from the
- 3:58
main stage but also uh dubbed it in real
- 4:01
time across multiple languages.
- 4:05
The second category is speech to action.
- 4:07
Uh these are talks and the model uses
- 4:10
tools and I think this is one of the
- 4:12
most underexplored areas that we have.
- 4:15
Um I actually almost titled this talk uh
- 4:17
voice is the next capability overhang
- 4:19
because I think there is just a vast
- 4:21
vast amount of stuff um that we could be
- 4:24
doing in this category that we are not
- 4:25
currently doing. Uh for example um
- 4:29
there's a broad spectrum I don't have
- 4:31
you know there's way too many examples
- 4:32
even fit on this slide but three
- 4:34
categories that that I find particularly
- 4:36
interesting. Uh first is form filling
- 4:38
right um so much of the internet is just
- 4:41
filling out forms. Um, and there is, you
- 4:44
know, today no reason why you shouldn't
- 4:45
be able to just talk. You know, I would
- 4:47
love it if instead of spending an hour
- 4:48
filling out a government document, I
- 4:50
could just talk for five minutes and it
- 4:52
would get 90% of it for me and I would
- 4:54
do a quick check, you know, just to make
- 4:56
sure that everything looked good, right?
- 4:57
That is a vastly superior experience
- 4:59
than like having to type in every single
- 5:01
name and address that I've lived in the
- 5:03
last 5 years and, you know, all of my
- 5:04
previous identities. Um and so I think I
- 5:08
think that one is though it may seem
- 5:09
boring you know affects a a significant
- 5:11
GDP of the internet right the next
- 5:13
category is creative tools uh where I am
- 5:16
privileged enough that I can speak the
- 5:18
language of software and so I can tell
- 5:20
codeex you know here's exactly what I
- 5:22
want you to build and I can articulate
- 5:24
it in a way that um I get much more
- 5:26
leverage than sort of just like cludily
- 5:28
trying to iterate one thing at a time
- 5:29
but I can't do that when it comes to you
- 5:32
know using making music or painting um
- 5:34
and so if I don't have the ability to
- 5:37
articulate um the exact aesthetic that
- 5:39
I'm looking for. Um and if I don't know
- 5:41
how to use Photoshop or Ableton, I'm
- 5:43
left in this state where, you know, my
- 5:45
my taste exceeds my capability. Um and
- 5:47
so I'm really looking forward to
- 5:48
integrating voice into creative tools so
- 5:50
that I can just sort of cludgy go along
- 5:52
and, you know, vibe create, vibe
- 5:55
compose, vibe paint, um and make
- 5:56
something that that's really beautiful
- 5:58
to me. And I think the generalizable um
- 6:02
category here, right, then just starts
- 6:04
to become computer use. And we've
- 6:05
already seen some companies start to do
- 6:07
this. Um, you know, it raises the
- 6:08
question of like, look, if the models
- 6:10
are just getting good enough to do
- 6:11
everything on a computer that a human
- 6:13
can do, like why am I talking to an app?
- 6:17
Why am I talking to a terminal? Why am I
- 6:18
not just talking to the entire computer?
- 6:22
Uh, and so I think that's sort of a
- 6:23
really interesting uh way to start
- 6:25
exploring. But if you're a developer
- 6:27
today, right? Whoops. If you're a
- 6:28
developer today, um, what does that mean
- 6:31
for building your own software, right?
- 6:32
And I think it is like much easier than
- 6:34
you think to start adding audio as an
- 6:36
intelligence layer to the intelligence
- 6:38
layer to the apps that you already have.
- 6:40
Um, if you're building a modern web
- 6:41
application, you already expose so much
- 6:43
of it as like action as nouns and verbs,
- 6:46
right? And if you think about all the
- 6:47
verbs that you have, you have uh API
- 6:49
endpoints, you have, you know, React
- 6:51
hooks. Each of those things can like
- 6:53
pretty relatively easily be converted
- 6:55
into a tool that you expose to a model
- 6:57
and then you can give the user the
- 6:58
ability to just drive your existing
- 7:00
software um with their voice, right? And
- 7:02
yes, you still need guardrails, you
- 7:03
still need safety checks. Like many of
- 7:05
the talks today are going to talk about
- 7:06
securing and you know productizing this,
- 7:08
but um for this I just want you to think
- 7:10
about you know what would it mean to
- 7:11
take your existing software and just
- 7:13
talk to it.
- 7:16
Um and to go back to that misconception,
- 7:18
right? I think there are a lot of um you
- 7:20
know like if you're talking to the
- 7:21
software maybe it can talk back but
- 7:22
we've been developing other ways of
- 7:24
communicating with the user for decades
- 7:26
right we know these things we know we
- 7:28
can show notifications and popups we can
- 7:30
change state like the color of a button
- 7:32
or a drop shadow we can highlight text
- 7:34
um if you've used computer use in the
- 7:35
codeex app you know there's this amazing
- 7:37
like little ghost cursor animation that
- 7:39
goes around and clicks things for you so
- 7:41
we don't have to use words to actually
- 7:42
tell the user what is happening on
- 7:44
screen with their software
- 7:48
Uh and then the last bucket here is
- 7:49
event to speech, right? Um the model
- 7:51
receives an event and talks to the user.
- 7:53
Um and sort of the counterpoint from
- 7:55
speech to action. I think this one is
- 7:56
still very very exploratory, right? Um
- 7:59
you know, if you saw Quinn's talk, I
- 8:00
think there's a lot of uh space here of
- 8:02
like things we can do. Um and to me, we
- 8:04
haven't quite seen what AI native really
- 8:07
looks like in this vein yet. But um of
- 8:09
the things that I've seen, I think
- 8:11
there's a couple of through lines that I
- 8:12
tend to notice, right? The first is
- 8:15
hands-free or screen-free experiences.
- 8:17
There might be times where uh I need to
- 8:19
interact with software, interact with
- 8:20
objects and I can't use my hands or more
- 8:22
importantly my attention is diverted
- 8:24
elsewhere. Um that might be something
- 8:26
you know like uh a recipe app. Maybe I'm
- 8:29
cooking and I need to just say like
- 8:30
what's going on and and have something
- 8:32
else have something happen. Um the other
- 8:35
category is proactive outreach, right?
- 8:36
Where you the model needs to be able to
- 8:38
tell you something or get your attention
- 8:39
in a way um that you might not be
- 8:41
looking at, right? I think every
- 8:42
developer um has uh an endless amount of
- 8:46
notifications and events happening in
- 8:47
their software. But um no developer in
- 8:50
their right mind would sort of say I
- 8:51
should show all of these logs. Nor you
- 8:52
know would they say I should speak all
- 8:54
of these logs. But we can start to
- 8:56
conceive of voice as this like upper
- 8:58
level in this escalatory path of like
- 9:00
okay maybe you animate something and
- 9:01
then maybe you pop something up and then
- 9:02
if that doesn't work maybe you talk to
- 9:04
the user to get their attention.
- 9:07
And underlying both of these categories
- 9:09
and I think this this you know this
- 9:10
whole presentation is this broader theme
- 9:12
of accessibility. Um on a personal
- 9:14
personal note, I know like multiple
- 9:16
developers who um over the course of
- 9:18
their careers lost mobility in their
- 9:21
hands, lost dexterity in their fingers
- 9:23
and for many of them, they thought their
- 9:25
career as a programmer was more or less
- 9:26
over. Um and then came large language
- 9:29
models, right? Then came coding agents
- 9:31
and voice agents and now they generate
- 9:33
orders of magnitude more code than they
- 9:35
like previously did um you know on a
- 9:37
given given day or month. Um, and so I
- 9:39
think there's there's a lot that we can
- 9:40
unlock here uh for the broader world as
- 9:42
well.
- 9:45
Um, to go back to like I said, you know,
- 9:46
I think like when it comes to these
- 9:47
three modalities, you can mix and match
- 9:49
them and we already have some, you know,
- 9:51
rudimentary ways that we're seeing this.
- 9:53
I think there's things like, you know,
- 9:54
all of these pieces for in-car
- 9:56
assistants exist, though nothing has
- 9:58
quite like combined them into this
- 9:59
seamless way. You can talk to like the
- 10:02
CarPlay dashboard. Um, you can tell it,
- 10:04
"Hey, go play some Spotify music for
- 10:06
me." Um, and then it can come back and
- 10:07
tell you, Google Maps can come back and
- 10:09
tell you, hey, like, you know, there's
- 10:10
traffic on this route. We're going to
- 10:11
reroute you. But like we can now start
- 10:13
to think about what does it mean to
- 10:14
combine that into like a single voice
- 10:16
agent across multiple modes. [snorts]
- 10:18
Um, similarly, you know, there's a lot
- 10:20
of experimentation in the game space
- 10:21
with multimodality. Um, you can think
- 10:23
about a real life character where you're
- 10:25
talking to it to, you know, mine
- 10:27
information about the game, about the
- 10:28
world. um you can talk to it to execute
- 10:30
actions on your behalf and then it can
- 10:32
react to like world events right that
- 10:34
are happening and then give that
- 10:35
information to you rather than just like
- 10:37
a simple notification
- 10:41
um and I think the question you know
- 10:42
behind the question here right is like I
- 10:44
mentioned we've had all these things for
- 10:45
a while people have been prototyping
- 10:46
them for a while why focus on them now
- 10:48
why think about building with them now
- 10:50
um and I think that brings me to uh a
- 10:53
little bit of context here right as as
- 10:54
I'm hopefully most of you know
- 10:56
traditionally voice agents are built in
- 10:58
this chain Ed model, right? You uh talk,
- 11:01
you transcribe, you send that to a
- 11:02
language model. It calls tools.
- 11:04
Hopefully, it doesn't take too long to
- 11:05
respond. Um it then generates text
- 11:07
output. You make that into audio and
- 11:09
then you play that back to the user.
- 11:12
And some time ago, um OpenAI, you know,
- 11:14
decided on a different approach, right?
- 11:16
The real-time model family does not do
- 11:18
any transcription behind the scenes. It
- 11:20
is trained on native audio as tokens.
- 11:23
So, you send audio in and you get audio
- 11:26
back out. And the industry I think in
- 11:28
general has been, you know, trending
- 11:29
more in this direction. Um, and not just
- 11:31
making it native audio, but even just
- 11:33
letting go of the turnbased abstraction
- 11:35
that we've had, right? Um, and so making
- 11:37
it that it's just continuous streaming
- 11:38
audio in and out.
- 11:42
And the reason that opening I did this
- 11:43
was, you know, turns out there's a lot
- 11:46
of stuff that you lose when you
- 11:47
transcribe speech and when you
- 11:49
transcribe audio, right? Um there's the
- 11:50
old saying that when humans communicate
- 11:52
face to face, 55% of the information is
- 11:55
in body language, another 38% is in your
- 11:57
tone of voice and the last like 7% is
- 11:59
the actual words you are saying. Um and
- 12:01
so when you transcribe, you lose tone
- 12:03
and cadence and emotional uh you know
- 12:05
impact, you lose like whether they're
- 12:07
trying to interrupt you, you lose
- 12:08
background noise, all of this stuff
- 12:09
which is really important context for
- 12:11
the model to understand.
- 12:13
Um a much more quantitative reason to do
- 12:15
it is that you know the first two uh
- 12:17
voice modes in chat GBT were built with
- 12:20
this chained approach. Um and as you can
- 12:21
see had you know significantly higher
- 12:23
latency than using uh the native
- 12:25
approach with advanced voice mode.
- 12:29
And that brings me to GPT realtime 2. Um
- 12:32
and this is going to be the one part of
- 12:33
the talk where you know I make my
- 12:34
shameless plug. Um real time 2 is the
- 12:37
the latest model in the real time
- 12:39
family. We released it a couple of
- 12:40
months ago. Um, and the the really cool
- 12:43
thing about this model is that it brings
- 12:45
reasoning to the audio medium. Um, and
- 12:48
so much like our text models, it can now
- 12:50
think before it speaks. Um, I'm sure
- 12:52
many of us have seen some demos of voice
- 12:54
models saying things that are a little
- 12:56
bit less than intelligent. Um, and so
- 12:58
you can now, you know, try to ensure
- 13:00
that you give it more reasoning budget
- 13:02
uh to come up with a good answer. Part
- 13:04
of why that's also useful is that we
- 13:05
introduced tool calling a little while
- 13:07
ago. And so the model in addition to
- 13:09
thinking it can also delegate parallel
- 13:10
tool calls. Um you can start to bring
- 13:12
these together. Um though of course that
- 13:14
adds latency, right? And you know that's
- 13:16
why we also added preamles. Um preamles
- 13:19
are a way that you can prompt the model
- 13:20
to uh give the user a heads up if it's
- 13:23
going to be thinking or if it's going to
- 13:24
be calling tools. Um you know if you
- 13:27
think about the scenario of a travel
- 13:29
agent, right? If I called the travel
- 13:30
agent on the phone, uh you would want
- 13:32
the travel agent to say to say, "Hey,
- 13:34
like I'm going to go check flight
- 13:35
prices, right? give me a couple seconds
- 13:37
to do that. Um, and now with an AI
- 13:40
travel agent, um, and preamles, you can
- 13:42
actually have it communicate that to the
- 13:43
user while it's performing actions in
- 13:45
the background. Uh, there's a few other
- 13:47
things here, right? It's got longer
- 13:49
context, better domain understanding,
- 13:51
um, more natural voices, and it's much
- 13:52
more steerable. Um, there's some really
- 13:54
cool features, uh, that it can do when
- 13:56
it comes to like wake words and just
- 13:59
waiting for you to to tell it. You you
- 14:01
can give it a name. Uh, you can say
- 14:02
like, you know, hey, Marin, do you want
- 14:04
to say hi to the room? Um, and if you've
- 14:06
like prompted that into the model, then
- 14:08
it'll, you know, go ahead and and
- 14:09
respond to you, right? Uh, and of
- 14:11
course, you know, uh, obligatory
- 14:13
benchmark slide, uh, it does pretty well
- 14:15
on the the latest audio benchmarks, too.
- 14:18
So, TLDDR, uh, is a pretty good model.
- 14:21
Um,
- 14:23
but I think the the kind of final thing
- 14:25
that you know I want to leave you with
- 14:26
here is um when building voice agents uh
- 14:31
not to
- 14:33
start with the question of like what
- 14:35
kind of voice agent am I trying to
- 14:37
build, right? I think the thing I want
- 14:38
to leave you with is start with the
- 14:40
question of like what is the role of
- 14:43
voice and audio in this interaction? Um
- 14:46
and then how do I move forward from
- 14:47
there, right? Right? And often when I
- 14:49
ask that question, it leads to a bunch
- 14:51
more questions after that. Things like
- 14:53
what can the model perceive? What
- 14:55
context does it have? Right? Um what
- 14:57
tools are available to it and which of
- 14:59
those tools should it be, you know, uh
- 15:01
executing safely and correctly? Um
- 15:04
should it communicate now? Should it
- 15:06
wait? Uh you know, how should it
- 15:08
communicate? Should it be sending visual
- 15:10
notifications or using audio? Um, and so
- 15:13
taken together, yeah, I hope everybody
- 15:15
in here can can start to build some much
- 15:17
richer experiences with voice. Um,
- 15:19
because like others have said, uh, I do
- 15:21
believe that AGI will be spoken, not
- 15:23
typed.
- 15:26
Thank you very much. I'll be at the
- 15:27
OpenAI booth, uh, for any Q&A after. Um,
- 15:30
yeah, have a good event.
- 15:46
>> [music]