Speech-to-Speech Model Research at Google DeepMind — Valeria Wu Fon & Tom Ouyang, Google DeepMind

Read the talk

Speech-to-Speech Model Research at Google DeepMind

Valeria Wu Fon and Tom Ouyang explain how joint audio, video, and text training supports versatile voice agents—and why reasoning, response timing, language use, and visual output must work together.

From a talk by Valeria Wu Fon and Tom Ouyang

At a glance

Ideas worth remembering

  • End-to-end speech recognition simplified audio-to-text modeling, but translation, responses, speaker interpretation, and visual context still required additional system work.

  • Interleaved audio, video, and text pre-training teaches cross-modal relationships that a shared model can draw on for translation, visual questions, and speech generation.

  • Reasoning quality and conversational speed can conflict: more thinking before an answer or tool call delays the first audio response.

  • Useful voice behavior includes selective localization and selective attention: keep familiar borrowed terms when appropriate, and avoid treating every background sound as a conversational interruption.

  • The longer-term goal is a single promptable model that can switch among translation, task execution, brainstorming, and informal conversation while coordinating multimodal input and output.

A transcription model still leaves the conversation to build

Adding voice to an agent can look straightforward: put automatic speech recognition (ASR) before the agent and text-to-speech (TTS) after it. But recognizing words and speaking an answer leave much of a conversation to the surrounding system. Valeria Wu Fon, Gemini’s speech-to-speech product lead, and Tom Ouyang, a speech-to-speech engineer, introduce their work through that gap. The intended applications span everyday questions in Search Live and Gemini Live, as well as enterprise voice agents exposed through cloud and API products. 1:33

In Ouyang’s account, speech recognition before roughly 2018 usually required a chain of specialized components: feature extraction, acoustic modeling, pronunciation modeling, language modeling, and a second pass that rescored candidate transcriptions. Moving toward end-to-end neural systems let a model learn the mapping from acoustic input to text, reducing the amount of domain knowledge needed to assemble that chain.

The scope of the learned task remained narrow. Audio went in; a transcription came out. Translation, replies, descriptions of a speaker’s emotion or pace, and the use of images still required additional system work. Collapsing the recognizer’s internal pipeline therefore did not create a general conversational model. Every new capability could bring another integration problem.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Pre-training teaches relationships between audio, video, and text

Gemini changes where those relationships are learned. Rather than merely attaching audio embeddings to a text model, Ouyang describes pre-training on interleaved multimodal examples. A single training example can begin with a text instruction, continue through a sequence of video and audio, and require an output that depends on both. The modalities meet during the stage that supplies the bulk of the training data. 3:03

Consider the bedtime-story example. The instruction asks for a summary; the following video and audio supply the story. The expected output includes a written summary and timestamps for interesting events. To satisfy that target, the model must turn its understanding of what it hears and sees into text, while retaining where relevant events occurred.

Different training tasks exercise different directions through the same foundation:

  • Video captioning: Audio and video both inform the caption, so the text can reflect information carried by either signal.
  • Speech recognition and synthesis: Examples connect audio to text and text to audio.
  • Agentic tasks: Task-oriented examples join those modality conversions in a unified token embedding space, giving the model a foundation for understanding how audio, video, and text relate.

What does the bedtime-story task require the model to connect? The diagram follows the example from instruction through sensory input to its output target. The summary and event timestamps depend on the audio and video together: training asks the model to express both the story’s meaning and where events happen. This is a picture of the training task, without specifying an encoder design or tokenization scheme.

How it fits togetherThe bedtime-story training example

Ask for a bedtime-story summary.

A text instruction establishes the task; interleaved audio and video supply the story; the output expresses its meaning and event timing.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

2:33 · section reference included

Translation must begin before the utterance is complete

Live translation puts those learned relationships under a timing constraint. A listener can specify English as the desired output while friends speak Spanish, Italian, or Chinese. The model must translate as they speak. An offline translator has the full utterance available before producing a result; the streaming system begins with incomplete input. 4:33

Ouyang reports that the team finds streaming translation quality about as good as offline systems, though the talk does not specify the evaluation conditions or numerical results for that comparison. The task combines several demands: switching among languages without knowing them beforehand, preserving the source speaker’s voice, understanding multiple speakers, tolerating noise, and producing output in real time. He attributes much of that capability to pre-training, making the application almost a prompting task.

Translation is one mode of the shared foundation. With an instruction to translate and incoming audio, the model produces streaming translation. With image and audio input and a question about the image, it answers. An embodied agent can add a face and tools that show information. The proposed advantage is reuse: different prompts call on relationships already learned by one model. 5:33

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:03 · section reference included

More thinking can make a conversation slower

Wu Fon organizes the research around three objectives that must coexist in the same product:

  • Conversational: Low latency, quick responses, and a natural conversational rhythm.
  • Intelligent: Task completion, instruction following, and reasoning that help the model accomplish something useful.
  • Multimodal: Inputs can include video, screen sharing, and PDFs alongside audio, while outputs can go beyond speech.

6:10

Internationalization runs through all three. Wu Fon says she believes the majority of Gemini users are non-English speakers, so the goal includes making these capabilities work across the languages customers use. A system that speaks fluently in one locale still has work to do before it can serve the intended audience.

The objectives interfere with one another. Increasing the thinking budget gives the model more time to reason before answering or calling a tool, and Wu Fon reports that this improves intelligence evaluations. But that work delays the first audio response. The user experiences the delay as a pause in the conversation, which can undermine its natural rhythm. The research challenge is to improve useful reasoning while keeping the exchange responsive; the talk presents this as ongoing work. 7:39

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

6:10 · section reference included

From a multilingual meeting to a sofa question

The first demonstration returns to translation. The live model is described as supporting 70-plus languages, and a Google Meet example shows the interaction: enable speech translation, select the desired language, and let several participants converse. The exchange moves through weather in Shanghai, a family visit to a park, and a birthday dinner at a restaurant in Sweden. 8:39

The timing matters as much as the translated content. Wu Fon highlights that translation begins shortly after a participant starts speaking, reducing the feeling of a rigid, turn-by-turn relay. The model starts catching the new speech directly, so participants do not have to treat every interruption as a disruption to a separate translation cycle.

The next example moves into Search Live, which Wu Fon says uses the same speech-to-speech model as the developer-facing Live API. In the sofa demonstration, real-time video and audio let the user ask a question without first describing the furniture. A tool call then brings up relevant search cards, adding a visual route to more information alongside the spoken exchange. 11:12

Localization also changes the answer itself. Wu Fon describes the response as Spanish localized for Spain, with the furniture term mid-century left in English because that term is commonly used in Spanish. The useful behavior is selective: preserve a familiar borrowed expression while speaking the surrounding answer in the user’s language. Translating every word would miss that usage.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:39 · section reference included

A roadside agent needs accurate identifiers and selective attention

The roadside-assistance demonstration gives the same model a different job. A driver has blown a tire and pulled over. The agent begins by asking for a name and policy number. When the driver cannot supply the policy number, the agent changes the information it requests: a registration plate and postcode can support the lookup instead. That change keeps the task moving despite a missing identifier. 12:14

A truck horn interrupts the situation, and the driver reacts before providing the requested details. The agent subsequently says it has found the record, warns the driver to stay clear of the road, and asks whether the vehicle is a blue MINI Cooper F series. The driver confirms. The observable progression is from an unavailable policy number to alternate identifiers, then a retrieved vehicle description and confirmation. The demonstration stops there; it does not show assistance being dispatched.

This task makes two requirements especially visible:

  • Alphanumeric accuracy: Plates, postcodes, and addresses carry task-critical characters. A plausible-sounding conversation is insufficient if the identifiers used to locate a record are wrong.
  • Proactive audio: The model decides when external input warrants a response or interruption. Background noise or another person talking should not automatically make it stop speaking or cut itself short.

13:21

Here, proactive includes knowing when to leave the conversation alone. The model must distinguish relevant input from sound that happens around the user. Wu Fon connects that requirement to the places voice agents will actually be used: on trains, during walks, and on the go. Quiet-office assumptions would make the model’s conversational timing poorly suited to those environments.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:14 · section reference included

Visual presence adds another output to coordinate

The final demonstration adds a visual presence. Wu Fon introduces a pilot with Citi at Cloud Next that supports customized real-time avatars, ranging from hyperrealistic humans to cartoons. The experience uses the same speech-to-speech model and combines low-latency conversation with multilingual lip syncing. An avatar gives the spoken interaction another output whose timing must remain coordinated with the exchange. 14:21

In the example, a user asks about progress toward a daughter’s college fund. The agent replies about the goal, mentions a possible opportunity, and then acknowledges that another person has joined, offering congratulations on college acceptance. Wu Fon uses the demo to bring together multimodal input and output, tool calling for relevant user information, conversational fluidity, and internationalization. It is presented as an early convergence of those capabilities in a pilot.

The closing ambition is broader than giving an assistant a good voice. Wu Fon’s forecast is that AGI will be spoken. For the team, that means one promptable model must let a user move among translation, taking action, brainstorming, and rambling. Each mode changes what a useful response looks like: translation follows someone else’s speech, an agent advances a task, and an open-ended exchange follows the user’s thought. The intended product must switch among those modes while keeping its understanding, timing, language use, and outputs working together. 15:56

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

13:51 · section reference included

Resources

  • A route into Ouyang’s earlier work on language modeling and multilingual input, including keyboard decoding under latency and memory constraints. These publications provide background on input-system tradeoffs rather than specifications for the speech model.

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> Hi everyone. Thanks for coming to our

  3. 0:14

    talk. We're going to talk a lot about

  4. 0:16

    how we are planning on voiceifying the

  5. 0:18

    agentic future with speech-to-speech

  6. 0:20

    research here at Google DeepMind. My

  7. 0:22

    name is Valeria. I'm the product lead

  8. 0:24

    for the speech-to-speech model in Gemini

  9. 0:25

    and Tom.

  10. 0:27

    >> Yeah, my name is Tom. I work on

  11. 0:28

    speech-to-speech as an engineer and

  12. 0:30

    Gemini. So, yeah, great to be presenting

  13. 0:32

    our work.

  14. 0:33

    >> Cool.

  15. 0:35

    So, before we start diving into the

  16. 0:37

    history of Google research and audio and

  17. 0:39

    the latest research we've been doing, I

  18. 0:41

    just want to take a step back and really

  19. 0:43

    kind of like think about why this is

  20. 0:45

    important to us in the team. And the

  21. 0:47

    main reason is that voice is the most

  22. 0:49

    natural way for humans to interact with

  23. 0:51

    both the physical and the virtual world.

  24. 0:53

    And we're already seeing so many

  25. 0:55

    applications that are starting even

  26. 0:56

    within the own Google products. We see

  27. 0:58

    voice being used to ask questions and

  28. 1:01

    you know, do homework or info seeking or

  29. 1:03

    EDU in search live and Gemini live. All

  30. 1:05

    the way to these same models being

  31. 1:07

    deployed in cloud and the API for

  32. 1:09

    enterprise or voice agent use cases.

  33. 1:11

    Should I use this as a

  34. 1:13

    microphone?

  35. 1:15

    Hello. Okay.

  36. 1:16

    Um

  37. 1:18

    So, because of that, we only think that

  38. 1:19

    the number of applications is going to,

  39. 1:21

    you know, exponentially increase over

  40. 1:23

    the next few years and we believe that

  41. 1:25

    speech-to-speech models are the way to

  42. 1:26

    go when we want to build robust

  43. 1:28

    universal voice agents. And with that,

  44. 1:30

    I'll hand it off to Tom.

  45. 1:33

    >> All right. So, of course, like one way

  46. 1:34

    to create a voice agent is to just add

  47. 1:37

    an ASR like speech-to-text model in the

  48. 1:39

    pipeline and then a text-to-speech model

  49. 1:40

    on the other end, right? So, of course,

  50. 1:42

    that's where a lot of the history of

  51. 1:43

    speech come from. Of course, Google has

  52. 1:45

    been working on this for a long time.

  53. 1:46

    I'm going to give you a little bit of a

  54. 1:47

    historical overview of what speech

  55. 1:49

    modeling, especially speech-to-text

  56. 1:51

    automatic

  57. 1:52

    speech recognition, typically looks

  58. 1:54

    like, right? So, up until around 2018,

  59. 1:56

    this usually involved a lot of different

  60. 1:57

    components. You have um you know of

  61. 1:59

    course feature extraction that's fairly

  62. 2:01

    general and then you have all these

  63. 2:02

    different pieces like acoustic modeling,

  64. 2:04

    pronunciation modeling, language

  65. 2:06

    modeling, a second pass rescoring that

  66. 2:08

    allows you to go from the audio input to

  67. 2:10

    a text transcription. Right? And of

  68. 2:12

    course like around 2018, these moved

  69. 2:14

    more and more towards end-to-end

  70. 2:15

    systems. You don't have to have so much

  71. 2:16

    domain knowledge. You can actually have

  72. 2:18

    mostly the the neural model learn this

  73. 2:21

    pattern and mapping between acoustic

  74. 2:23

    inputs and text. But these aren't really

  75. 2:25

    the end-to-end models that we think

  76. 2:26

    about when we think about LLMs. They're

  77. 2:27

    only doing kind of one thing, which is

  78. 2:29

    speech to a transcription of that

  79. 2:31

    speech. They're not responding, they're

  80. 2:32

    not translating. If you wanted to have

  81. 2:34

    the model tell you about the tone or the

  82. 2:36

    emotion or the speed of the speaker. If

  83. 2:38

    you wanted to bias it towards words or

  84. 2:40

    much less like images, those are all

  85. 2:41

    things you have to build yourself as

  86. 2:43

    part of the system and there's really

  87. 2:45

    really a barrier to how easily you can

  88. 2:46

    scale these systems. So,

  89. 2:49

    fast forward to now, which is kind of of

  90. 2:51

    course the LLM era, right? Like you

  91. 2:52

    know, the first LLMs were mostly text,

  92. 2:54

    but even there I think you could kind of

  93. 2:55

    hack audio embeddings into these text

  94. 2:57

    models and it kind of worked, but now of

  95. 2:59

    course for a long time now, Gemini

  96. 3:00

    models have been very natively

  97. 3:02

    multimodal. So, what does that mean? It

  98. 3:04

    means that when we train these models in

  99. 3:05

    pre-training, which is where the bulk of

  100. 3:06

    the data comes from, these are

  101. 3:08

    multimodal interleaved examples, right?

  102. 3:10

    So, the bottom uh diagram here gives you

  103. 3:13

    kind of one example of what that might

  104. 3:14

    look like. So, this is a task where

  105. 3:16

    you're asking the model to summarize a

  106. 3:18

    bedtime story and there's a text prompt

  107. 3:19

    in the beginning, but then there's this

  108. 3:21

    sequence of video and audio inputs that

  109. 3:24

    the model gets

  110. 3:25

    and then of course like what you expect

  111. 3:27

    the model to do here is produce um both

  112. 3:29

    the summary and also annotate timestamps

  113. 3:31

    for where interesting things happen and

  114. 3:33

    so forth. So, this example is teaching

  115. 3:35

    the model to translate its understanding

  116. 3:37

    of the audio and the video into text.

  117. 3:39

    You might have other examples in

  118. 3:41

    pre-training that ask you to caption a

  119. 3:43

    video. So, you might have a video that

  120. 3:44

    has audio and the model is learning to

  121. 3:47

    bias towards both the video and the

  122. 3:48

    audio signal to caption this well. And

  123. 3:51

    of course like there's limitless like

  124. 3:53

    YouTube videos with captions that you

  125. 3:54

    can train these models on. Other models

  126. 3:55

    might actually try to generate audio

  127. 3:57

    from the video from the text, right? So,

  128. 3:59

    you can have ASR, TTS, or any

  129. 4:01

    combination of these plus all of these

  130. 4:02

    sort of agentic tasks all kind of

  131. 4:05

    learned under one unified token

  132. 4:07

    embedding space. So, this becomes a

  133. 4:09

    foundation for a lot of what we want to

  134. 4:10

    do in audio because we already have a

  135. 4:11

    model that understands audio, video,

  136. 4:14

    text, and how these things relate and

  137. 4:16

    transition from one to the next.

  138. 4:19

    So, very quickly, like one of the

  139. 4:21

    applications that this enables that

  140. 4:22

    we've launched recently is live

  141. 4:24

    translation, right? And this is a kind

  142. 4:25

    of application that kind of only works

  143. 4:27

    when you have all of these capabilities

  144. 4:29

    working within the same model. You have

  145. 4:30

    basically state-of-the-art translation

  146. 4:32

    quality. Even though this model is

  147. 4:33

    translating basically as the user or

  148. 4:35

    speakers are speaking, you kind of ask,

  149. 4:37

    "Hey, I I speak English. There's maybe

  150. 4:39

    friends who are talking in Spanish and

  151. 4:41

    Italian and Chinese." And it's

  152. 4:42

    translating all of them to your language

  153. 4:43

    as they talk. Um and then we're we're

  154. 4:45

    finding is the translation quality for

  155. 4:46

    this like streaming real-time

  156. 4:48

    translation is about as good as you

  157. 4:50

    would get with offline systems, right?

  158. 4:52

    Where you kind of know the full

  159. 4:53

    utterance

  160. 4:54

    um from the very beginning. So, that's

  161. 4:56

    something that has been classically very

  162. 4:58

    hard to do with these cascaded systems,

  163. 5:00

    but with LLMs, it actually just

  164. 5:03

    a lot of it comes out of the

  165. 5:04

    pre-training. So, of course, to do this

  166. 5:06

    task it needs to do multilingual

  167. 5:07

    switching because you could be

  168. 5:08

    translating across different languages.

  169. 5:09

    You don't know what those languages are

  170. 5:10

    beforehand. It needs to preserve the

  171. 5:12

    speaker voice of the source speaker and

  172. 5:15

    be able to understand multiple speakers,

  173. 5:17

    be robust to noise, and of course, like

  174. 5:19

    do all this in real time, right? So,

  175. 5:21

    again, it would be very hard to try to

  176. 5:23

    engineer this, but then with the LLM and

  177. 5:25

    Gemini models, this almost becomes a

  178. 5:28

    prompting task.

  179. 5:29

    And on that, like, I know, in this

  180. 5:31

    diagram we're saying, "Hey, at the top

  181. 5:32

    with these models, if you prompt it to

  182. 5:34

    do the speaking like this streaming

  183. 5:36

    translation task,

  184. 5:38

    and you give it the audio, it will

  185. 5:39

    produce the streaming translation

  186. 5:40

    output, right? The same model, if you

  187. 5:42

    ask it to act like an agent and respond

  188. 5:44

    to maybe image and audio input, maybe

  189. 5:47

    asking questions about that image, it

  190. 5:48

    will give you an answer. And finally,

  191. 5:50

    like very well we'll show examples of

  192. 5:51

    this, you can also have it create this

  193. 5:53

    embodied, you know, virtual agent that

  194. 5:55

    has a face, that has, you know, things

  195. 5:57

    and tools that it can show you, and it

  196. 5:59

    will produce this sort of embodied agent

  197. 6:01

    experience. So, with that, I'm going to

  198. 6:03

    give it to Valeria to talk more about

  199. 6:04

    the North Star and some of the key demo

  200. 6:07

    products that we built.

  201. 6:10

    >> Yeah, so to create this type of kind of

  202. 6:12

    universal, versatile, uh kind of

  203. 6:14

    model/product,

  204. 6:16

    there's like three vectors that we think

  205. 6:17

    about when we do research and product

  206. 6:19

    for these models. Um and they also come

  207. 6:21

    with some challenges, so I'll like walk

  208. 6:23

    you through some of them. So, at the

  209. 6:25

    core of a speech-to-speech model, the

  210. 6:27

    first thing that people usually think

  211. 6:28

    about is that it has very

  212. 6:29

    conversational, right? It's low latency,

  213. 6:31

    it's very conversational, very snappy,

  214. 6:33

    very natural. But, I think within our

  215. 6:35

    team, we really don't only want this

  216. 6:36

    model to sound nice. We also have two

  217. 6:39

    pillars at the top that we also really

  218. 6:40

    care about like pulling all together

  219. 6:42

    into one model, which is intelligence

  220. 6:44

    and it being multimodal. So, when we

  221. 6:46

    talk about intelligence, we talk about,

  222. 6:48

    you know, task completion, instruction

  223. 6:49

    following, reasoning, like capabilities

  224. 6:51

    that the model needs to have natively in

  225. 6:54

    order to complete tasks and to like do

  226. 6:56

    things uh that have high customer

  227. 6:58

    satisfaction, for example. And on the

  228. 7:00

    other side of the Venn diagram, we also

  229. 7:02

    have the idea that these models should

  230. 7:04

    be very multimodal, both in audio in and

  231. 7:06

    audio out, right? So, sometimes a user

  232. 7:08

    doesn't only want to input audio in and

  233. 7:11

    have that be the start of the

  234. 7:12

    conversation. We need video, your screen

  235. 7:14

    sharing, PDFs, whatever you would want

  236. 7:16

    the model to interpret and understand,

  237. 7:18

    we should be able to stream it in and

  238. 7:20

    also produce output out of it. So,

  239. 7:22

    that's kind of like the trifecta of

  240. 7:23

    which we think about speech-to-speech

  241. 7:25

    models. Um and I want to add a caveat

  242. 7:27

    about ITNN. I think actually the

  243. 7:29

    majority of our Gemini users are

  244. 7:30

    non-English speakers. Um so, we put a

  245. 7:32

    big focus on having and making sure that

  246. 7:34

    all these capabilities work within not

  247. 7:37

    only, you know, EN-US, but all the

  248. 7:39

    languages that our customers care about.

  249. 7:41

    Um

  250. 7:41

    of course, this also our North Star also

  251. 7:44

    becomes one of the biggest challenges in

  252. 7:45

    our research because

  253. 7:47

    once you move one of the knobs, it's

  254. 7:49

    very easy for the other knobs to kind of

  255. 7:51

    like mess up, right? Like a very quick

  256. 7:53

    example, oh, how do we increase

  257. 7:55

    intelligence in the model? Well, you can

  258. 7:57

    turn thinking high or like the thinking

  259. 7:59

    is high as possible to have the model

  260. 8:01

    think a lot before calling a tool or

  261. 8:03

    answering a question, which in eval's it

  262. 8:05

    does show that it does improve the

  263. 8:07

    model's intelligence. But when what does

  264. 8:09

    that do to latency, right? And time to

  265. 8:11

    first audio and the naturalness of the

  266. 8:12

    conversation? So, within the Gemini

  267. 8:15

    team, we're really trying to push

  268. 8:16

    forward research initiatives that can

  269. 8:18

    kind of blend in the three of them

  270. 8:20

    without really sacrificing any of those

  271. 8:22

    by a lot. Um,

  272. 8:24

    but in the meantime, we're going to show

  273. 8:25

    you some of the demos that we think are

  274. 8:27

    hinting at how our speech-to-speech

  275. 8:29

    model can combine all these three into

  276. 8:32

    really cool application. So,

  277. 8:34

    the first one is you kind of saw this as

  278. 8:36

    a preview, but our live model, as you

  279. 8:38

    know, uh, powers streaming translation

  280. 8:40

    that supports 70+ languages. So, this is

  281. 8:42

    a little bit around the core of

  282. 8:44

    conversation quality and IT&N efforts

  283. 8:46

    that we have in the team. I'll play a

  284. 8:47

    quick video on how this works on Google

  285. 8:49

    Meets to help two people or maybe

  286. 8:51

    multiple people that are speaking

  287. 8:53

    different languages still have a live

  288. 8:54

    conversation.

  289. 8:57

    Oh.

  290. 8:58

    One sec. Okay.

  291. 9:00

    >> Okay, let's turn on speech translation.

  292. 9:04

    >> So, here the user can simply select the

  293. 9:06

    language that they want the translation

  294. 9:08

    to happen in and then the rest will be

  295. 9:10

    done in real time in multiple speakers

  296. 9:12

    kind of having a conversation back and

  297. 9:13

    forth. So, I'll show you a snippet of

  298. 9:15

    this video of what happens.

  299. 9:20

    >> It's great to see you both. Cassie,

  300. 9:22

    how's the weather in Shanghai?

  301. 9:25

    >> It's nice to meet you. The weather here

  302. 9:27

    is really nice, sunny and bright. I

  303. 9:29

    spent all day Saturday in the park with

  304. 9:31

    my family.

  305. 9:35

    >> That sounds wonderful.

  306. 9:37

    Anna, you mentioned last week that you

  307. 9:39

    were celebrating your birthday. How was

  308. 9:40

    it?

  309. 9:41

    >> Oh, that was fantastic. I had dinner at

  310. 9:45

    my favorite restaurant with some

  311. 9:46

    friends. If you visit Sweden, you must

  312. 9:49

    try this restaurant. It was

  313. 9:52

    >> So, um a lot of things happening. Not

  314. 9:54

    only there's, you know, real-time

  315. 9:56

    translation. Whoops.

  316. 9:58

    How do I

  317. 9:59

    back to Okay.

  318. 10:00

    Um not only there's real-time

  319. 10:02

    translation that is happening, uh but we

  320. 10:04

    see that it's in a multi-speaker

  321. 10:05

    setting, you know, low latency, like

  322. 10:07

    right after the user starts speaking,

  323. 10:09

    the translation kicks off, so that it

  324. 10:10

    doesn't feel like it's really turn by

  325. 10:12

    turn and robotic and you're afraid to

  326. 10:14

    interrupt because, you know, once you

  327. 10:16

    start speaking, the model will start

  328. 10:17

    catching your translation directly. So,

  329. 10:19

    now I'll show some other applications

  330. 10:21

    where we see that we have one single

  331. 10:23

    speech-to-speech model for many

  332. 10:24

    conversational frontiers. So, two

  333. 10:26

    products that I want to highlight here

  334. 10:28

    are the same model that we power search

  335. 10:30

    live for everyday conversations in any

  336. 10:32

    language. It's the same model that we

  337. 10:34

    use to power developer experiences in

  338. 10:36

    our live API. As you'll see in the

  339. 10:38

    videos, both use cases are extremely

  340. 10:40

    different, but because we try and make

  341. 10:42

    this this model or this speech-to-speech

  342. 10:44

    model very promptable and universal, you

  343. 10:47

    can see how it adapts to the different

  344. 10:48

    use cases. So, let me just play the

  345. 10:50

    first video.

  346. 11:12

    So, to highlight in this video, many

  347. 11:14

    things are happening. First, the

  348. 11:15

    multimodal, you know, it's real-time uh

  349. 11:17

    video and audio streaming in, so that

  350. 11:19

    you don't have to actually describe what

  351. 11:20

    the sofa looks like. But then the

  352. 11:22

    multimodal out, you see that there was a

  353. 11:24

    tool called made to pull up the right

  354. 11:25

    search cards if the user wanted to know

  355. 11:27

    more information. And then finally, the

  356. 11:29

    ITN component, you know, we're

  357. 11:31

    localizing the speaker's Spanish to

  358. 11:32

    Spain Spanish. And also, when we do the

  359. 11:36

    kind of responding in Spanish, we're

  360. 11:37

    also smart enough to know that some

  361. 11:39

    words are better to be left in English,

  362. 11:41

    like the mid-century type of word. It

  363. 11:43

    didn't translate it directly to Spanish

  364. 11:45

    cuz it knew that the term was commonly

  365. 11:46

    known in English. Sorry, in Spanish. So,

  366. 11:49

    this is one example of how users are

  367. 11:51

    using search live for like info-seeking

  368. 11:53

    questions. Now, let me show a very

  369. 11:55

    different demo on how we do this for

  370. 11:57

    conversational voice agents.

  371. 12:02

    Oh.

  372. 12:04

    Sorry.

  373. 12:14

    >> Hi, I've blown a tire and had to pull

  374. 12:17

    over to the side of the road. Can you

  375. 12:18

    help?

  376. 12:20

    >> Hello, you've reached Auto Transform

  377. 12:23

    Insurance Roadside Assistance. My name

  378. 12:25

    is Ace. I understand this is stressful,

  379. 12:27

    but we'll take care of everything. To

  380. 12:29

    get started, could I please have your

  381. 12:30

    name and your policy number?

  382. 12:32

    >> Uh my name is Richard Jenkins. I I don't

  383. 12:34

    actually have my policy number to hand.

  384. 12:39

    >> No worries, Richard. I can look you up

  385. 12:41

    with your registration plate and

  386. 12:42

    postcode instead. Could you provide

  387. 12:44

    those for me, please?

  388. 12:46

    >> Whoa, that was a bit close. Um yeah, my

  389. 12:49

    registration plate is BD21

  390. 12:54

    XYA

  391. 12:56

    and uh my postcode is SN48ZX.

  392. 13:03

    >> Policy details BD21

  393. 13:06

    XYA. Thank you. For your safety, please

  394. 13:09

    stay clear of the road. I've found your

  395. 13:11

    details and I see you're in a blue Mini

  396. 13:14

    Cooper F-Series. Is that the vehicle

  397. 13:16

    you're in?

  398. 13:16

    >> Yeah, yeah, that that's the vehicle.

  399. 13:20

    >> Got it.

  400. 13:21

    >> Okay, so I'll pause it here, but as you

  401. 13:23

    can see, very different things are

  402. 13:24

    happening under the hood with the same

  403. 13:25

    model. You know, it's more about

  404. 13:27

    alphanumeric accuracy for complex kind

  405. 13:30

    of like postcode numbers or addresses.

  406. 13:32

    There's like this feature that we have

  407. 13:34

    called proactive audio, which is the

  408. 13:36

    idea that the LLM knows when or when not

  409. 13:38

    to respond to an external input. So, for

  410. 13:40

    example, if you're in a conversation and

  411. 13:42

    someone else is talking or there's

  412. 13:43

    background noise happening in the back,

  413. 13:45

    the model knows to not stop or cut

  414. 13:47

    itself short because there's external

  415. 13:49

    noise happening. Because we realized

  416. 13:51

    that the majority of these conversations

  417. 13:53

    are not happening in an office room with

  418. 13:55

    like no noise, you know, in a sealed

  419. 13:56

    environment. They're happening on the

  420. 13:58

    go, on the train, while you're on a

  421. 13:59

    walk. Um and that's the type of

  422. 14:01

    experiences that we want to facilitate

  423. 14:03

    with the speech-to-speech model.

  424. 14:05

    Um lastly, um another thing we're really

  425. 14:07

    excited about is not only voice out, but

  426. 14:09

    also multimodal out. We believe that,

  427. 14:12

    you know, the true AGI conversational

  428. 14:14

    frontier will also require visual

  429. 14:16

    presence. So, we were excited to launch

  430. 14:18

    our first kind of pilot demo with City

  431. 14:20

    um in Cloud Next, which supports

  432. 14:22

    customized real-time avatars.

  433. 14:24

    Um and you can kind of like

  434. 14:26

    personalize anything from a

  435. 14:28

    hyperrealistic human to a cartoon

  436. 14:30

    appearance and everything in the middle.

  437. 14:32

    And powered by the same speech-to-speech

  438. 14:33

    model we've been showing, it allows to

  439. 14:35

    have, you know, low-latency,

  440. 14:37

    multilingual lip-syncing, and a really

  441. 14:38

    kind of like uh ongoing fluid

  442. 14:40

    conversation that has visual presence.

  443. 14:42

    So, let me show you a demo that also

  444. 14:44

    brings our Venn diagram together and

  445. 14:46

    what we're excited about.

  446. 14:48

    >> All this is reminding me of my

  447. 14:49

    daughter's college fund. How are we

  448. 14:51

    tracking on that?

  449. 14:55

    >> You're tracking well, Jackson. I've also

  450. 14:57

    identified a new opportunity that may

  451. 14:59

    get you there even sooner.

  452. 15:01

    Oh, and I can see Lisa just joined you.

  453. 15:03

    Hi, Lisa. She must be so excited about

  454. 15:06

    her college acceptance. Congratulations.

  455. 15:08

    It's wonderful to see your savings goals

  456. 15:10

    coming to life.

  457. 15:11

    >> Excelente. Las cosas han estado tan

  458. 15:14

    inestables últimamente.

  459. 15:18

    >> Sí. Los servicios de tecnología y

  460. 15:20

    comunicaciones han mostrado un desempeño

  461. 15:23

    sólido en lo que va del año.

  462. 15:25

    >> That's good news.

  463. 15:26

    >> I always joke that the user's audio in

  464. 15:28

    Spanish is worse than the audio model

  465. 15:30

    speaking back, but um this is kind of

  466. 15:32

    just to show how like our Venn diagram

  467. 15:34

    of combining, you know, multimodality in

  468. 15:36

    and out, you know, tool calling to like

  469. 15:38

    pull up the relevant examples from the

  470. 15:40

    user, and also conversational fluidity

  471. 15:42

    with ITNN are starting slowly to come

  472. 15:44

    together in these types of demos that

  473. 15:46

    we're excited to keep pushing the

  474. 15:47

    frontier of.

  475. 15:49

    Um so with this parting thought, I guess

  476. 15:51

    last

  477. 15:52

    kind of thought that we have for you is

  478. 15:53

    that we believe that AGI will not be

  479. 15:55

    typed, that it will be spoken. Um and

  480. 15:58

    for it to be spoken, there's a lot of

  481. 15:59

    things that need to work together in a

  482. 16:01

    single promptable, versatile model that

  483. 16:04

    allows a user to switch between all the

  484. 16:06

    sorts of conversation modes that we're

  485. 16:07

    looking at, right? From translation to

  486. 16:09

    taking action to brainstorming to

  487. 16:11

    rambling, and we truly believe in the

  488. 16:13

    power of these speech-to-speech models

  489. 16:14

    to achieve that like seamless switching.

  490. 16:17

    Um so we're excited to push the frontier

  491. 16:19

    on that. So if you're excited or want to

  492. 16:21

    learn more, please come talk to us, and

  493. 16:22

    thank you so much for coming.

  494. 16:38

    >> [music]