Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI

Bohan Li· EliseAI13:13

Read the talk

Realtime Voice Agents with Frontier Intelligence

Bohan Li explains how the presented voice-agent harness overlaps transcription, language-model generation, background tool work, and speech synthesis so a capable but slower model can participate in a natural phone conversation.

From a talk by Bohan Li

At a glance

Ideas worth remembering

  • A cascaded voice architecture creates separate opportunities to reduce perceived latency in transcription, response planning, and speech synthesis.

  • Speculation only works with revision: newer audio cancels stale transcription work, and useful background tool results can cancel and restart an early response.

  • Preparing a response and speaking it are different decisions. The harness generates while the caller talks but waits for turn completion before emitting audio.

  • A prefix cache can play a reusable opening while fresh synthesis handles personalized content. Sending the full sentence to the provider preserves prosody, after which duplicate opening audio is suppressed.

  • The clinic exchange is a prerecorded demonstration showing a pause, changed scheduling preference, alternatives, and booking confirmation; it does not establish live backend persistence or numerical performance.

A cascaded voice stack creates three places to hide latency

The presented harness’s starting constraint is awkward but common: the voice agent needs the judgment of a capable language model, yet a phone caller notices every pause. Bohan Li approaches the problem through an analogy to his previous self-driving work. Instead of treating voice as one opaque model call, the system separates perception, planning, and control so each stage can be accelerated on its own terms.

The mapping is direct:

  • Perception — transcription: Audio signals become text the reasoning system can process, much as cameras and lidar turn the road into usable observations.
  • Planning — the language model: The model consumes that transcription and decides what the agent should say or do.
  • Control — speech synthesis: Generated text becomes audio, analogous to turning a planned trajectory into physical controls.

This decomposition exposes three different waits: recognizing the caller, preparing a useful response, and producing audible speech. The harness can overlap work across those waits instead of demanding that one faster model solve the entire latency problem.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Fast transcription can be corrected without blocking every turn

The perception layer uses what Li calls a streaming speculative transcriber. A fast streaming engine emits text immediately, while a slower batch engine processes more audio and conversational context to produce a more accurate interpretation. Downstream work can begin from the fast result instead of waiting for the slower pass every time.

What happens when an accurate result arrives after the conversation has already moved on? The correction path is freshness-aware. If both engines produce the same text, there is nothing to revise. If new streaming audio arrives while an older corrective pass is running, that correction is canceled because it no longer represents all available audio. A carefully computed stale answer loses to a newer observation.

The concrete example begins after the agent asks for a name and date of birth. Early fragments such as “Sure” arrive quickly but do not contain the requested fields. Once the slower engine has enough audio and the context of the question, it can distinguish the name-like portion from the date-of-birth portion. Punctuation-only changes are ignored, and the resulting text is released to the agent.

The diagram answers a timing question: how can the slower path improve accuracy without making every turn wait for it? Both paths begin from the same audio, but only a still-relevant correction replaces the streaming hypothesis. Newer audio can invalidate work already in flight.

How it fits togetherSpeculative transcription with cancelable correction

Audio arrives incrementally during the caller’s utterance.

Fast text starts downstream work; the slower result is applied only if it remains relevant when it returns.

1:562:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

1:56 · section reference included

Generate early, but do not speak early

The planning layer attacks sequential model round trips. Tool calls are expensive here because a slow but intelligent language model may need to generate a request, wait for execution, ingest the result, and generate again. The presented harness moves suitable tool work into background agents, then inserts the result into the main model’s context as though that model had made the call itself. The talk describes the context insertion but does not specify its message format.

Every new transcription detection can also start an eager generation. Crucially, generation and emission are separate decisions: the model may prepare a candidate response while the caller is still speaking, but the harness does not play it until the end of the utterance is confirmed. This spends some computation speculatively in exchange for a chance to have the right response ready sooner.

The name-and-date example shows the revision process. “Sure” triggers an early generation, but the background extractor has no requested fields yet. More partial text arrives, still without a recognizable name. From that incomplete evidence, the main model starts preparing to ask the caller to spell the name because it suspects a transcription problem. Nothing has been spoken, so this wrong turn remains cheap to discard.

The observable change occurs when the corrected transcription arrives. The background agent can now identify the name and date of birth, including phonetic matching for a possible name mistranscription. The harness cancels the draft created without those results and starts a new generation with the extracted fields in context. Only after turn completion does the selected response leave the system.

Speculation does not remove inference; it deliberately creates work that may be thrown away. Its value depends on how often useful generation overlaps the caller’s remaining speech and how expensive canceled generations are. The talk provides no latency, cost, or cancellation-rate measurements.

3:263:56
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

3:26 · section reference included

A cached prefix lets the agent start speaking sooner

The control layer receives the response as a token stream. Waiting for the complete sentence before requesting speech would expose the remaining generation time to the caller, so the goal is to begin playback while the language model is still writing the continuation.

Captures the key cache miss at the personalized name, which releases the reusable opening while the remainder is still being generated.
Captures the key cache miss at the personalized name, which releases the reusable opening while the remainder is still being generated.

A prefix cache watches the stream for word sequences whose audio has already been generated, either during an earlier response or elsewhere in the same call. It does not immediately accept a one-word match; in the walkthrough it waits until three words before registering the first hit. At the same time, the growing text is sent to Cartesia over a WebSocket for fresh synthesis.

The developing sentence begins with the reusable phrase “You said your name is.” Such scripted openings recur, so their audio is a plausible cache target. The next token introduces the caller’s particular name and causes a miss. That miss becomes the release point: the harness immediately plays the cached generic opening while the personalized continuation is still being generated.

From the caller’s perspective, speech has started. Internally, the language model and speech provider are still finishing the sentence. The talk demonstrates why common openings are reusable, but it does not report a cache-hit rate, so the frequency of this latency win remains unspecified.

5:566:26
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:56 · section reference included

Preserve full-sentence prosody, then suppress the duplicate

Playing cached words creates a second problem: text-to-speech sounds more natural when the provider sees the whole sentence, but the opening has already been heard. The presented harness still sends Cartesia the complete text, including that opening. The provider therefore synthesizes the sentence with its normal context and prosody, unaware that a cache exists.

Illustrates the duplicate-suppression step: the provider’s already-played prefix is removed and only the fresh continuation is emitted.
Illustrates the duplicate-suppression step: the provider’s already-played prefix is removed and only the fresh continuation is emitted.

When the fresh audio returns, the harness suppresses the portion corresponding to the prefix and plays only the remaining audio after the cached clip. This preserves sentence-level context for synthesis while avoiding duplicate playback. Li allows that there may be a small hiccup; the talk does not explain how the two audio versions are aligned at the cut or measure whether listeners can detect the join.

The diagram answers the key control-layer question: how can cached speech start immediately without forcing the provider to synthesize an isolated, context-poor suffix? The same full sentence travels through both branches. One branch supplies early playback; the other supplies a natural continuation, with its duplicate prefix removed before output.

How it fits togetherCached-prefix playback joined to full-sentence synthesis

The response arrives word by word.

The cache starts speech early while the provider retains full sentence context; duplicate provider audio is suppressed before the continuation plays.

8:329:01
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

8:32 · section reference included

A prerecorded clinic-booking demonstration makes the hidden work disappear

The final demonstration is a prerecorded clinic-booking call. The supplied recording does not establish the caller’s identity or whether the patient details are real or synthetic. The caller asks to schedule an ultrasound. Elise requests a name and date of birth, confirms that the caller is a new patient, and asks permission to text an insurance-upload link. The captioned digits accompanying the birth-date request are not interpreted or expanded here.

The agent responds to the changed constraint by offering two appointment times for the following week.
The agent responds to the changed constraint by offering two appointment times for the following week.

The more revealing moment is a change of plan. The agent begins offering the earliest appointment, but the caller interrupts to check a calendar. Elise stops and says, “Sure. Take your time.” When he asks for the following week, it offers two times on Tuesday, July 7. He selects 2:00 p.m., and the agent confirms the booking.

That sequence matters because it is not merely a fixed questionnaire. The conversation pauses, abandons an in-progress offer, accepts a new scheduling constraint, presents alternatives, and reaches a selection. The caller hears short, conventional turns while transcription revisions, tentative generations, field extraction, tool work, and streaming speech can occur behind them.

The demonstration illustrates the intended conversational experience, not a full systems evaluation. It does not expose the underlying tool traces, establish that the demonstration used a live clinical scheduling system, or provide comparative latency and recognition measurements. What it does show is the product goal: substantial background coordination should collapse into an uneventful phone call.

9:319:55
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:31 · section reference included

The harness serves work in housing and healthcare

Li closes by placing the engineering work inside EliseAI’s focus on housing and healthcare—areas he describes as critical in people’s lives. The company is headquartered in New York and was working to expand its Bay Area presence. He ends with an invitation to follow the company and join the team.

Closing context for EliseAI’s focus on healthcare and housing, where natural realtime conversation supports consequential workflows.
Closing context for EliseAI’s focus on healthcare and housing, where natural realtime conversation supports consequential workflows.

That framing explains why the harness matters. Natural conversation is not merely a cosmetic layer when a caller is trying to schedule care or handle an important housing workflow. A frontier model alone does not create a realtime voice agent: the surrounding system decides when provisional work may begin, when stale work must be canceled, when tool results should revise a draft, and when cached and fresh audio can safely reach the caller.

11:5912:30
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:59 · section reference included

Resources

From the talk

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> My name is Bo. I'm going to be here

  3. 0:14

    presenting real-time voice agents with

  4. 0:17

    Frontier Intelligence. Effectively,

  5. 0:19

    going to be talking a little bit about

  6. 0:21

    how we at Xnor.ai architected our voice

  7. 0:24

    agent harness to get real-time voice

  8. 0:27

    with the Frontier level of intelligence

  9. 0:30

    that we need.

  10. 0:32

    Okay. So, before I start,

  11. 0:35

    I think I wanted to kind of draw some

  12. 0:37

    parallels about

  13. 0:39

    why we decided to go with cascaded voice

  14. 0:42

    agents and especially kind of

  15. 0:45

    comparing that to self-driving cars

  16. 0:47

    which I was working in before. So, to

  17. 0:50

    me, cascaded voice agents makes a lot of

  18. 0:52

    sense when you view it in lens of kind

  19. 0:54

    of breaking it down into perception

  20. 0:56

    which is

  21. 0:58

    for self-driving cars, it's you know,

  22. 1:00

    the bounding boxes, the camera, the

  23. 1:02

    lidar.

  24. 1:03

    For voice, it's going to be the

  25. 1:05

    transcription. Basically, effectively

  26. 1:07

    turning these like signals from the real

  27. 1:09

    world into

  28. 1:11

    elements of data that the language model

  29. 1:14

    or whatever brain you're working on

  30. 1:16

    can process.

  31. 1:19

    Second one is the planning step which is

  32. 1:21

    pretty straightforward.

  33. 1:23

    This is where the language model

  34. 1:25

    will take in the outputs from the

  35. 1:27

    perception stage and produce the outputs

  36. 1:30

    that you want to produce out back and

  37. 1:32

    out into the real world.

  38. 1:34

    And finally, there's the controls layer

  39. 1:36

    where

  40. 1:38

    on self-driving, you'd be taking the

  41. 1:39

    trajectory that the planner would output

  42. 1:42

    and kind of turn it into the real

  43. 1:45

    controls to kind of build like drive the

  44. 1:47

    car. Here, we're turning the text into

  45. 1:51

    audio that we use to express our voice

  46. 1:53

    agent's thoughts.

  47. 1:56

    And um

  48. 1:58

    yeah, so here I'm going to be like going

  49. 1:59

    to diving into each one of these

  50. 2:01

    elements and we've made a few kind of

  51. 2:04

    interesting tricks on each of these

  52. 2:06

    areas to

  53. 2:08

    improve the speed of our voice agents

  54. 2:10

    without sacrificing the intelligence.

  55. 2:13

    So, the first one is going to be uh the

  56. 2:16

    transcriber layer. So, we came up with

  57. 2:18

    this concept called like the streaming

  58. 2:19

    speculative transcriber where

  59. 2:21

    effectively we are layering a fast

  60. 2:24

    streaming transcriber like Flux on top

  61. 2:27

    of or kind of below a

  62. 2:29

    uh scribe V2 or a accurate batch

  63. 2:31

    transcription which kind of takes in

  64. 2:34

    more context. It's a little bit slower,

  65. 2:35

    but it will give you more accurate

  66. 2:37

    detections.

  67. 2:39

    So, we're going to walk through a

  68. 2:40

    scenario. So, in this in this case the

  69. 2:42

    agent just asked, you know, providing

  70. 2:44

    can you provide your name and date of

  71. 2:46

    birth and the user is going to say this

  72. 2:48

    and we'll see how that plays out um

  73. 2:51

    timing-wise. So, first we're going to

  74. 2:53

    get, you know, the short detection. Um

  75. 2:56

    we'll get it from we'll get it from the

  76. 2:57

    streaming layer. The accurate

  77. 3:00

    layer uh the corrective layer is not

  78. 3:02

    going to fire because it's the same

  79. 3:03

    text.

  80. 3:04

    Um we're going to get some more

  81. 3:07

    streaming text detections and in this

  82. 3:09

    case the corrective layer is actually

  83. 3:11

    canceled because we got new um new text.

  84. 3:14

    So, you know, more context, more audio

  85. 3:18

    is going to beat the old accurate one.

  86. 3:21

    And here's where kind of the first

  87. 3:22

    correction comes in. So, because the

  88. 3:25

    scribe V2 layer understands, you know,

  89. 3:28

    the the context of the question, it's

  90. 3:30

    able to understand that this is talking

  91. 3:31

    about name and this is a date of birth.

  92. 3:34

    Then a couple more detections, these are

  93. 3:35

    just punctuation, we don't care.

  94. 3:37

    And so, in the end we kind of release

  95. 3:39

    this text over to the agent.

  96. 3:44

    And moving on um to the language model

  97. 3:46

    layer.

  98. 3:47

    So, here since we're kind of using these

  99. 3:50

    slow but intelligent LLMs, we really

  100. 3:53

    want to reduce the number of round trips

  101. 3:55

    and the thing that causes us to do a lot

  102. 3:58

    of inferences is tool calling. So, one

  103. 4:00

    way to get rid of that is by having

  104. 4:03

    background agents do the tool calling

  105. 4:06

    for you and kind of

  106. 4:08

    um push the tools back into the context

  107. 4:11

    of the main agent so that it thinks it

  108. 4:13

    made the tool call, but

  109. 4:15

    um

  110. 4:16

    but it it it really didn't.

  111. 4:18

    So,

  112. 4:19

    uh so, we remember from like detections

  113. 4:21

    from before.

  114. 4:22

    So, well, what happened is each one of

  115. 4:24

    these detections is going to trigger a

  116. 4:27

    um an early

  117. 4:30

    kind of generation of the agent and we

  118. 4:33

    but we won't actually

  119. 4:35

    emit this out until we're confirming

  120. 4:38

    that the user has finished speaking. So,

  121. 4:40

    in this case, the user says, "Sure." The

  122. 4:42

    agent kind of knows that the user is

  123. 4:43

    about to say something else. Our

  124. 4:45

    background tool calling here, which is

  125. 4:47

    going to be helping us find figure out

  126. 4:49

    the name and the date of birth from the

  127. 4:50

    user detection, is not firing. So,

  128. 4:53

    nothing much there.

  129. 4:55

    Um the next instant detection comes in.

  130. 4:58

    It says that,

  131. 4:59

    you know, still not really a name. Um

  132. 5:02

    our agent kind of plays along and

  133. 5:03

    continues there.

  134. 5:06

    Now, kind of a more more context come

  135. 5:08

    comes back. The agent kind of feels like

  136. 5:10

    there should be a name. It's going to

  137. 5:12

    ask to spell it out because it's

  138. 5:14

    probably thinking there's some

  139. 5:15

    transcription error here. Still no name

  140. 5:17

    or date of birth.

  141. 5:19

    And then finally, this you remember this

  142. 5:20

    is kind of our corrected um final

  143. 5:22

    instant detection from the transcriber

  144. 5:25

    from the Scribe V2.

  145. 5:27

    Um here, our eager kind of agent

  146. 5:30

    generation that was made without any

  147. 5:32

    tool calls is going to get canceled

  148. 5:34

    because the background agent finally is

  149. 5:35

    able to find the name and date of birth

  150. 5:37

    it's looking for. So, it's going to

  151. 5:40

    retrigger and now the the agent actually

  152. 5:42

    has the context it needs.

  153. 5:44

    Um and you see here, it's kind of we're

  154. 5:46

    doing it the tool call here is a little

  155. 5:48

    bit um um

  156. 5:49

    some intelligence there. We're going to

  157. 5:51

    be like, you know, correcting

  158. 5:52

    mis-transcriptions of name, and doing

  159. 5:54

    some like phonetic matching here.

  160. 5:57

    Um and yeah, and then we'll kind of

  161. 5:59

    once we've understood that this is the

  162. 6:02

    end of the user utterance, we'll kind of

  163. 6:03

    emit it out. So, pretty standard.

  164. 6:06

    Okay, and then the next layer here is

  165. 6:08

    going to be text-to-speech. So, with

  166. 6:10

    text-to-speech

  167. 6:11

    the goal is to kind of take what the

  168. 6:13

    agent said, and the agent's going to be

  169. 6:15

    emitting this in a streaming fashion.

  170. 6:17

    So, we're going to need to

  171. 6:19

    um produce audio as quickly as possible.

  172. 6:21

    And ideally, what you can do is before

  173. 6:24

    the agent has even finished generating

  174. 6:26

    the full text you can have the audio

  175. 6:30

    play, so it's kind of hiding the latency

  176. 6:32

    of finishing the generation.

  177. 6:34

    So,

  178. 6:36

    um I'm going to kind of play the

  179. 6:37

    streaming um

  180. 6:38

    the stream the streaming uh agent output

  181. 6:40

    now. So, starts with you.

  182. 6:43

    And yeah, actually before I uh further,

  183. 6:46

    there's this new concept that we're

  184. 6:47

    introducing here called the prefix

  185. 6:48

    cache. So, the prefix cache is going to

  186. 6:51

    be looking at the

  187. 6:54

    um agent stream, and seeing if we

  188. 6:57

    already have generated audio for that

  189. 6:59

    sequence of words um from like a prior

  190. 7:03

    generation, or maybe like the same

  191. 7:04

    generation

  192. 7:05

    um in this

  193. 7:07

    uh in this call as as well.

  194. 7:10

    So, um it sees the word you. Uh we for

  195. 7:13

    this prefix cache, we're going to be,

  196. 7:15

    you know, we don't want to like

  197. 7:17

    immediately hit on every single word.

  198. 7:19

    We're going to be waiting for a little

  199. 7:20

    bit more words.

  200. 7:22

    Um so, after three words, the prefix

  201. 7:25

    cache gets our first hit.

  202. 7:27

    And um over here on the right, this is

  203. 7:30

    kind of our text-to-speech standard

  204. 7:32

    provider, you know, Cartesia is a

  205. 7:34

    text-to-speech engine with web socket

  206. 7:35

    support. So, we're we're piping the

  207. 7:38

    agent through the the cache, and also

  208. 7:41

    piping it through web socket.

  209. 7:44

    Um more tokens come in, more cache, more

  210. 7:48

    sending through web socket. Not much to

  211. 7:49

    say here.

  212. 7:51

    And okay, so now we get our first uh

  213. 7:55

    kind of first unique thing, which is we

  214. 7:58

    found the token that actually causes a

  215. 8:00

    cache miss. And it makes sense. If we're

  216. 8:02

    kind of caching previous generations,

  217. 8:05

    um you said your name is is a pretty

  218. 8:06

    common thing, but once you we add in the

  219. 8:08

    name, suddenly we're that's that's going

  220. 8:10

    to result in the cache miss.

  221. 8:12

    At this point, we're actually going to

  222. 8:13

    yield out our cached audio. So, you said

  223. 8:16

    your name is is going to be

  224. 8:18

    um emitted as the rest of the streaming

  225. 8:21

    text is coming back. So, at this point,

  226. 8:23

    the user hears the agent and user

  227. 8:25

    doesn't really know what's going on.

  228. 8:27

    They just looks like really fast

  229. 8:28

    response times to them.

  230. 8:31

    Um and now the kind of remaining text

  231. 8:33

    flows through.

  232. 8:35

    And at this point, we've already emitted

  233. 8:37

    from the cache. The cache has done its

  234. 8:39

    job. Um the rest we can kind of throw

  235. 8:42

    into Cartesia.

  236. 8:43

    And here's kind of the trick where

  237. 8:47

    Cartesia has seen the entire transcript

  238. 8:51

    up to this point. It

  239. 8:54

    to to Cartesia, like it doesn't know

  240. 8:56

    about the existence of this prefix

  241. 8:57

    cache. It's just going to generate this

  242. 8:59

    full sentence with, you know, standard

  243. 9:01

    natural prosody.

  244. 9:03

    But, what we do is when the generation

  245. 9:06

    comes back, since we've already played

  246. 9:08

    the audio here, we can actually suppress

  247. 9:09

    the audio from Cartesia here and just

  248. 9:12

    play out the remaining stuff. So,

  249. 9:15

    the user, there's might be a tiny bit of

  250. 9:18

    a hiccup. You know, I'll play some audio

  251. 9:20

    later and you'll know that you probably

  252. 9:22

    won't be able to notice.

  253. 9:24

    And effectively, we're kind of taking

  254. 9:26

    this and

  255. 9:28

    or taking this part and just emitting it

  256. 9:31

    directly after this audio has finished

  257. 9:33

    playing. So, it looks seamless to the

  258. 9:35

    user.

  259. 9:37

    And yeah, so I guess at this point kind

  260. 9:40

    of this is a lot of technical details.

  261. 9:42

    I'll I mean the best way is just kind of

  262. 9:44

    play a call and see

  263. 9:47

    see this in action. So here I recorded a

  264. 9:50

    call and let's just hear what happens.

  265. 9:55

    >> This is Bo OBGYN Elise speaking.

  266. 10:00

    >> Hi, I think I might be pregnant and I

  267. 10:02

    want to schedule an ultrasound to

  268. 10:03

    confirm.

  269. 10:06

    >> That's exciting news. Can you provide

  270. 10:08

    your name and date of birth?

  271. 10:12

    >> Sure, it's Elise Trial.

  272. 10:14

    Date of birth

  273. 10:15

    2303.

  274. 10:21

    >> Thank you. Are you a new patient with

  275. 10:23

    us?

  276. 10:26

    >> Yes.

  277. 10:29

    >> Welcome to Bo OBGYN. Can I text you a

  278. 10:32

    link to upload your insurance

  279. 10:33

    information?

  280. 10:36

    >> Sure.

  281. 10:39

    >> Okay, sent. Let me know when you've

  282. 10:41

    received it.

  283. 10:48

    I've received your insurance

  284. 10:49

    information. The earliest appointment I

  285. 10:52

    have is Thursday, July 2nd at 10:00 a.m.

  286. 10:55

    >> Actually, give me a second.

  287. 10:57

    Let me check my calendar.

  288. 11:01

    >> Sure, take your time.

  289. 11:04

    >> Um do you have something for next week?

  290. 11:09

    >> For next week, I can offer an ultrasound

  291. 11:11

    with Dr. Avery Stone at North Clinic on

  292. 11:13

    Tuesday, July 7th at 2:00 p.m. or 3:00

  293. 11:17

    p.m. Do either of those work for you?

  294. 11:20

    >> Yeah, 2:00 p.m. works.

  295. 11:23

    >> Great. Your appointment has been booked.

  296. 11:26

    We look forward to seeing you then.

  297. 11:29

    >> Thanks. Bye-bye.

  298. 11:32

    >> All right. Yeah, that's pretty much it.

  299. 11:34

    Um

  300. 11:35

    Yeah, you can kind of see our

  301. 11:38

    all this like uh

  302. 11:40

    streaming

  303. 11:42

    and you know, a lot of things are

  304. 11:43

    happening in the background and and you

  305. 11:44

    know, this is what really makes like

  306. 11:46

    voice agents interesting. And um there's

  307. 11:49

    a lot of effort that can be done in the

  308. 11:51

    harness to really kind of get a um

  309. 11:55

    natural conversation, which is what

  310. 11:57

    we're after.

  311. 11:59

    Uh okay. Yeah, so I guess briefly, you

  312. 12:01

    know, in the last part, I want to just

  313. 12:04

    talk a little bit about Elise. So I

  314. 12:06

    think Elise, you know, our headquarters

  315. 12:08

    are in New York and kind of we're trying

  316. 12:10

    to expand our presence here in the Bay

  317. 12:11

    Area. Um we I think it's maybe like a

  318. 12:15

    different style of company that I think

  319. 12:17

    people are

  320. 12:19

    like uh think of when they think about

  321. 12:21

    AI startups in San Francisco. Where

  322. 12:23

    we're actually very focused on um

  323. 12:26

    just like helping people and helping

  324. 12:30

    people where they need it, like kind of

  325. 12:32

    the life's most critical areas. We work

  326. 12:33

    on housing, health care and uh we're

  327. 12:37

    doing really well and you know, here's

  328. 12:40

    there's a link here to

  329. 12:42

    um you kind of join our team and there's

  330. 12:44

    going to we're going to be uh posting a

  331. 12:46

    lot on Twitter, so you can follow us at

  332. 12:48

    EliseAI as well.

  333. 12:50

    Um yeah, that's that's it.

  334. 12:53

    >> [applause]