Realtime Voice Agents with Frontier Intelligence — Bohan Li, EliseAI
Read the talk
Realtime Voice Agents with Frontier Intelligence
Bohan Li explains how the presented voice-agent harness overlaps transcription, language-model generation, background tool work, and speech synthesis so a capable but slower model can participate in a natural phone conversation.
From a talk by Bohan Li
At a glance
Ideas worth remembering
A cascaded voice architecture creates separate opportunities to reduce perceived latency in transcription, response planning, and speech synthesis.
Speculation only works with revision: newer audio cancels stale transcription work, and useful background tool results can cancel and restart an early response.
Preparing a response and speaking it are different decisions. The harness generates while the caller talks but waits for turn completion before emitting audio.
A prefix cache can play a reusable opening while fresh synthesis handles personalized content. Sending the full sentence to the provider preserves prosody, after which duplicate opening audio is suppressed.
The clinic exchange is a prerecorded demonstration showing a pause, changed scheduling preference, alternatives, and booking confirmation; it does not establish live backend persistence or numerical performance.
A cascaded voice stack creates three places to hide latency
The presented harness’s starting constraint is awkward but common: the voice agent needs the judgment of a capable language model, yet a phone caller notices every pause. Bohan Li approaches the problem through an analogy to his previous self-driving work. Instead of treating voice as one opaque model call, the system separates perception, planning, and control so each stage can be accelerated on its own terms.
The mapping is direct:
- Perception — transcription: Audio signals become text the reasoning system can process, much as cameras and lidar turn the road into usable observations.
- Planning — the language model: The model consumes that transcription and decides what the agent should say or do.
- Control — speech synthesis: Generated text becomes audio, analogous to turning a planned trajectory into physical controls.
This decomposition exposes three different waits: recognizing the caller, preparing a useful response, and producing audible speech. The harness can overlap work across those waits instead of demanding that one faster model solve the entire latency problem.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fast transcription can be corrected without blocking every turn
The perception layer uses what Li calls a streaming speculative transcriber. A fast streaming engine emits text immediately, while a slower batch engine processes more audio and conversational context to produce a more accurate interpretation. Downstream work can begin from the fast result instead of waiting for the slower pass every time.
What happens when an accurate result arrives after the conversation has already moved on? The correction path is freshness-aware. If both engines produce the same text, there is nothing to revise. If new streaming audio arrives while an older corrective pass is running, that correction is canceled because it no longer represents all available audio. A carefully computed stale answer loses to a newer observation.
The concrete example begins after the agent asks for a name and date of birth. Early fragments such as “Sure” arrive quickly but do not contain the requested fields. Once the slower engine has enough audio and the context of the question, it can distinguish the name-like portion from the date-of-birth portion. Punctuation-only changes are ignored, and the resulting text is released to the agent.
The diagram answers a timing question: how can the slower path improve accuracy without making every turn wait for it? Both paths begin from the same audio, but only a still-relevant correction replaces the streaming hypothesis. Newer audio can invalidate work already in flight.
Audio arrives incrementally during the caller’s utterance.
Fast text starts downstream work; the slower result is applied only if it remains relevant when it returns.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Generate early, but do not speak early
The planning layer attacks sequential model round trips. Tool calls are expensive here because a slow but intelligent language model may need to generate a request, wait for execution, ingest the result, and generate again. The presented harness moves suitable tool work into background agents, then inserts the result into the main model’s context as though that model had made the call itself. The talk describes the context insertion but does not specify its message format.
Every new transcription detection can also start an eager generation. Crucially, generation and emission are separate decisions: the model may prepare a candidate response while the caller is still speaking, but the harness does not play it until the end of the utterance is confirmed. This spends some computation speculatively in exchange for a chance to have the right response ready sooner.
The name-and-date example shows the revision process. “Sure” triggers an early generation, but the background extractor has no requested fields yet. More partial text arrives, still without a recognizable name. From that incomplete evidence, the main model starts preparing to ask the caller to spell the name because it suspects a transcription problem. Nothing has been spoken, so this wrong turn remains cheap to discard.
The observable change occurs when the corrected transcription arrives. The background agent can now identify the name and date of birth, including phonetic matching for a possible name mistranscription. The harness cancels the draft created without those results and starts a new generation with the extracted fields in context. Only after turn completion does the selected response leave the system.
Speculation does not remove inference; it deliberately creates work that may be thrown away. Its value depends on how often useful generation overlaps the caller’s remaining speech and how expensive canceled generations are. The talk provides no latency, cost, or cancellation-rate measurements.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A cached prefix lets the agent start speaking sooner
The control layer receives the response as a token stream. Waiting for the complete sentence before requesting speech would expose the remaining generation time to the caller, so the goal is to begin playback while the language model is still writing the continuation.
A prefix cache watches the stream for word sequences whose audio has already been generated, either during an earlier response or elsewhere in the same call. It does not immediately accept a one-word match; in the walkthrough it waits until three words before registering the first hit. At the same time, the growing text is sent to Cartesia over a WebSocket for fresh synthesis.
The developing sentence begins with the reusable phrase “You said your name is.” Such scripted openings recur, so their audio is a plausible cache target. The next token introduces the caller’s particular name and causes a miss. That miss becomes the release point: the harness immediately plays the cached generic opening while the personalized continuation is still being generated.
From the caller’s perspective, speech has started. Internally, the language model and speech provider are still finishing the sentence. The talk demonstrates why common openings are reusable, but it does not report a cache-hit rate, so the frequency of this latency win remains unspecified.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Preserve full-sentence prosody, then suppress the duplicate
Playing cached words creates a second problem: text-to-speech sounds more natural when the provider sees the whole sentence, but the opening has already been heard. The presented harness still sends Cartesia the complete text, including that opening. The provider therefore synthesizes the sentence with its normal context and prosody, unaware that a cache exists.
When the fresh audio returns, the harness suppresses the portion corresponding to the prefix and plays only the remaining audio after the cached clip. This preserves sentence-level context for synthesis while avoiding duplicate playback. Li allows that there may be a small hiccup; the talk does not explain how the two audio versions are aligned at the cut or measure whether listeners can detect the join.
The diagram answers the key control-layer question: how can cached speech start immediately without forcing the provider to synthesize an isolated, context-poor suffix? The same full sentence travels through both branches. One branch supplies early playback; the other supplies a natural continuation, with its duplicate prefix removed before output.
The response arrives word by word.
The cache starts speech early while the provider retains full sentence context; duplicate provider audio is suppressed before the continuation plays.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A prerecorded clinic-booking demonstration makes the hidden work disappear
The final demonstration is a prerecorded clinic-booking call. The supplied recording does not establish the caller’s identity or whether the patient details are real or synthetic. The caller asks to schedule an ultrasound. Elise requests a name and date of birth, confirms that the caller is a new patient, and asks permission to text an insurance-upload link. The captioned digits accompanying the birth-date request are not interpreted or expanded here.
The more revealing moment is a change of plan. The agent begins offering the earliest appointment, but the caller interrupts to check a calendar. Elise stops and says, “Sure. Take your time.” When he asks for the following week, it offers two times on Tuesday, July 7. He selects 2:00 p.m., and the agent confirms the booking.
That sequence matters because it is not merely a fixed questionnaire. The conversation pauses, abandons an in-progress offer, accepts a new scheduling constraint, presents alternatives, and reaches a selection. The caller hears short, conventional turns while transcription revisions, tentative generations, field extraction, tool work, and streaming speech can occur behind them.
The demonstration illustrates the intended conversational experience, not a full systems evaluation. It does not expose the underlying tool traces, establish that the demonstration used a live clinical scheduling system, or provide comparative latency and recognition measurements. What it does show is the product goal: substantial background coordination should collapse into an uneventful phone call.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The harness serves work in housing and healthcare
Li closes by placing the engineering work inside EliseAI’s focus on housing and healthcare—areas he describes as critical in people’s lives. The company is headquartered in New York and was working to expand its Bay Area presence. He ends with an invitation to follow the company and join the team.
That framing explains why the harness matters. Natural conversation is not merely a cosmetic layer when a caller is trying to schedule care or handle an important housing workflow. A frontier model alone does not create a realtime voice agent: the surrounding system decides when provisional work may begin, when stale work must be canceled, when tool results should revise a draft, and when cached and fresh audio can safely reach the caller.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
Watch the recording alongside the supplied timestamped captions; chapters are not supplied.
Product context for the housing and healthcare communication workflows discussed at the end of the talk, including VoiceAI and career information.
Related talks
- Building Effective Voice Agents
Compares cascaded transcription–LLM–speech pipelines with speech-to-speech systems and examines the accuracy, determinism, latency, and tool-use tradeoffs behind production voice agents.
- Designing Voice Agents for Real Conversations
Extends the latency discussion into turn detection, voice activity detection, interruption cancellation, and the risk of mistaking silence for a completed turn.
- 200 Million Patient Interactions Later: What the Generic Voice Stack Misses
Provides a healthcare-specific companion focused on specialist models, clinical evaluation, tool verification, safety, empathy, and escalation.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> My name is Bo. I'm going to be here
- 0:14
presenting real-time voice agents with
- 0:17
Frontier Intelligence. Effectively,
- 0:19
going to be talking a little bit about
- 0:21
how we at Xnor.ai architected our voice
- 0:24
agent harness to get real-time voice
- 0:27
with the Frontier level of intelligence
- 0:30
that we need.
- 0:32
Okay. So, before I start,
- 0:35
I think I wanted to kind of draw some
- 0:37
parallels about
- 0:39
why we decided to go with cascaded voice
- 0:42
agents and especially kind of
- 0:45
comparing that to self-driving cars
- 0:47
which I was working in before. So, to
- 0:50
me, cascaded voice agents makes a lot of
- 0:52
sense when you view it in lens of kind
- 0:54
of breaking it down into perception
- 0:56
which is
- 0:58
for self-driving cars, it's you know,
- 1:00
the bounding boxes, the camera, the
- 1:02
lidar.
- 1:03
For voice, it's going to be the
- 1:05
transcription. Basically, effectively
- 1:07
turning these like signals from the real
- 1:09
world into
- 1:11
elements of data that the language model
- 1:14
or whatever brain you're working on
- 1:16
can process.
- 1:19
Second one is the planning step which is
- 1:21
pretty straightforward.
- 1:23
This is where the language model
- 1:25
will take in the outputs from the
- 1:27
perception stage and produce the outputs
- 1:30
that you want to produce out back and
- 1:32
out into the real world.
- 1:34
And finally, there's the controls layer
- 1:36
where
- 1:38
on self-driving, you'd be taking the
- 1:39
trajectory that the planner would output
- 1:42
and kind of turn it into the real
- 1:45
controls to kind of build like drive the
- 1:47
car. Here, we're turning the text into
- 1:51
audio that we use to express our voice
- 1:53
agent's thoughts.
- 1:56
And um
- 1:58
yeah, so here I'm going to be like going
- 1:59
to diving into each one of these
- 2:01
elements and we've made a few kind of
- 2:04
interesting tricks on each of these
- 2:06
areas to
- 2:08
improve the speed of our voice agents
- 2:10
without sacrificing the intelligence.
- 2:13
So, the first one is going to be uh the
- 2:16
transcriber layer. So, we came up with
- 2:18
this concept called like the streaming
- 2:19
speculative transcriber where
- 2:21
effectively we are layering a fast
- 2:24
streaming transcriber like Flux on top
- 2:27
of or kind of below a
- 2:29
uh scribe V2 or a accurate batch
- 2:31
transcription which kind of takes in
- 2:34
more context. It's a little bit slower,
- 2:35
but it will give you more accurate
- 2:37
detections.
- 2:39
So, we're going to walk through a
- 2:40
scenario. So, in this in this case the
- 2:42
agent just asked, you know, providing
- 2:44
can you provide your name and date of
- 2:46
birth and the user is going to say this
- 2:48
and we'll see how that plays out um
- 2:51
timing-wise. So, first we're going to
- 2:53
get, you know, the short detection. Um
- 2:56
we'll get it from we'll get it from the
- 2:57
streaming layer. The accurate
- 3:00
layer uh the corrective layer is not
- 3:02
going to fire because it's the same
- 3:03
text.
- 3:04
Um we're going to get some more
- 3:07
streaming text detections and in this
- 3:09
case the corrective layer is actually
- 3:11
canceled because we got new um new text.
- 3:14
So, you know, more context, more audio
- 3:18
is going to beat the old accurate one.
- 3:21
And here's where kind of the first
- 3:22
correction comes in. So, because the
- 3:25
scribe V2 layer understands, you know,
- 3:28
the the context of the question, it's
- 3:30
able to understand that this is talking
- 3:31
about name and this is a date of birth.
- 3:34
Then a couple more detections, these are
- 3:35
just punctuation, we don't care.
- 3:37
And so, in the end we kind of release
- 3:39
this text over to the agent.
- 3:44
And moving on um to the language model
- 3:46
layer.
- 3:47
So, here since we're kind of using these
- 3:50
slow but intelligent LLMs, we really
- 3:53
want to reduce the number of round trips
- 3:55
and the thing that causes us to do a lot
- 3:58
of inferences is tool calling. So, one
- 4:00
way to get rid of that is by having
- 4:03
background agents do the tool calling
- 4:06
for you and kind of
- 4:08
um push the tools back into the context
- 4:11
of the main agent so that it thinks it
- 4:13
made the tool call, but
- 4:15
um
- 4:16
but it it it really didn't.
- 4:18
So,
- 4:19
uh so, we remember from like detections
- 4:21
from before.
- 4:22
So, well, what happened is each one of
- 4:24
these detections is going to trigger a
- 4:27
um an early
- 4:30
kind of generation of the agent and we
- 4:33
but we won't actually
- 4:35
emit this out until we're confirming
- 4:38
that the user has finished speaking. So,
- 4:40
in this case, the user says, "Sure." The
- 4:42
agent kind of knows that the user is
- 4:43
about to say something else. Our
- 4:45
background tool calling here, which is
- 4:47
going to be helping us find figure out
- 4:49
the name and the date of birth from the
- 4:50
user detection, is not firing. So,
- 4:53
nothing much there.
- 4:55
Um the next instant detection comes in.
- 4:58
It says that,
- 4:59
you know, still not really a name. Um
- 5:02
our agent kind of plays along and
- 5:03
continues there.
- 5:06
Now, kind of a more more context come
- 5:08
comes back. The agent kind of feels like
- 5:10
there should be a name. It's going to
- 5:12
ask to spell it out because it's
- 5:14
probably thinking there's some
- 5:15
transcription error here. Still no name
- 5:17
or date of birth.
- 5:19
And then finally, this you remember this
- 5:20
is kind of our corrected um final
- 5:22
instant detection from the transcriber
- 5:25
from the Scribe V2.
- 5:27
Um here, our eager kind of agent
- 5:30
generation that was made without any
- 5:32
tool calls is going to get canceled
- 5:34
because the background agent finally is
- 5:35
able to find the name and date of birth
- 5:37
it's looking for. So, it's going to
- 5:40
retrigger and now the the agent actually
- 5:42
has the context it needs.
- 5:44
Um and you see here, it's kind of we're
- 5:46
doing it the tool call here is a little
- 5:48
bit um um
- 5:49
some intelligence there. We're going to
- 5:51
be like, you know, correcting
- 5:52
mis-transcriptions of name, and doing
- 5:54
some like phonetic matching here.
- 5:57
Um and yeah, and then we'll kind of
- 5:59
once we've understood that this is the
- 6:02
end of the user utterance, we'll kind of
- 6:03
emit it out. So, pretty standard.
- 6:06
Okay, and then the next layer here is
- 6:08
going to be text-to-speech. So, with
- 6:10
text-to-speech
- 6:11
the goal is to kind of take what the
- 6:13
agent said, and the agent's going to be
- 6:15
emitting this in a streaming fashion.
- 6:17
So, we're going to need to
- 6:19
um produce audio as quickly as possible.
- 6:21
And ideally, what you can do is before
- 6:24
the agent has even finished generating
- 6:26
the full text you can have the audio
- 6:30
play, so it's kind of hiding the latency
- 6:32
of finishing the generation.
- 6:34
So,
- 6:36
um I'm going to kind of play the
- 6:37
streaming um
- 6:38
the stream the streaming uh agent output
- 6:40
now. So, starts with you.
- 6:43
And yeah, actually before I uh further,
- 6:46
there's this new concept that we're
- 6:47
introducing here called the prefix
- 6:48
cache. So, the prefix cache is going to
- 6:51
be looking at the
- 6:54
um agent stream, and seeing if we
- 6:57
already have generated audio for that
- 6:59
sequence of words um from like a prior
- 7:03
generation, or maybe like the same
- 7:04
generation
- 7:05
um in this
- 7:07
uh in this call as as well.
- 7:10
So, um it sees the word you. Uh we for
- 7:13
this prefix cache, we're going to be,
- 7:15
you know, we don't want to like
- 7:17
immediately hit on every single word.
- 7:19
We're going to be waiting for a little
- 7:20
bit more words.
- 7:22
Um so, after three words, the prefix
- 7:25
cache gets our first hit.
- 7:27
And um over here on the right, this is
- 7:30
kind of our text-to-speech standard
- 7:32
provider, you know, Cartesia is a
- 7:34
text-to-speech engine with web socket
- 7:35
support. So, we're we're piping the
- 7:38
agent through the the cache, and also
- 7:41
piping it through web socket.
- 7:44
Um more tokens come in, more cache, more
- 7:48
sending through web socket. Not much to
- 7:49
say here.
- 7:51
And okay, so now we get our first uh
- 7:55
kind of first unique thing, which is we
- 7:58
found the token that actually causes a
- 8:00
cache miss. And it makes sense. If we're
- 8:02
kind of caching previous generations,
- 8:05
um you said your name is is a pretty
- 8:06
common thing, but once you we add in the
- 8:08
name, suddenly we're that's that's going
- 8:10
to result in the cache miss.
- 8:12
At this point, we're actually going to
- 8:13
yield out our cached audio. So, you said
- 8:16
your name is is going to be
- 8:18
um emitted as the rest of the streaming
- 8:21
text is coming back. So, at this point,
- 8:23
the user hears the agent and user
- 8:25
doesn't really know what's going on.
- 8:27
They just looks like really fast
- 8:28
response times to them.
- 8:31
Um and now the kind of remaining text
- 8:33
flows through.
- 8:35
And at this point, we've already emitted
- 8:37
from the cache. The cache has done its
- 8:39
job. Um the rest we can kind of throw
- 8:42
into Cartesia.
- 8:43
And here's kind of the trick where
- 8:47
Cartesia has seen the entire transcript
- 8:51
up to this point. It
- 8:54
to to Cartesia, like it doesn't know
- 8:56
about the existence of this prefix
- 8:57
cache. It's just going to generate this
- 8:59
full sentence with, you know, standard
- 9:01
natural prosody.
- 9:03
But, what we do is when the generation
- 9:06
comes back, since we've already played
- 9:08
the audio here, we can actually suppress
- 9:09
the audio from Cartesia here and just
- 9:12
play out the remaining stuff. So,
- 9:15
the user, there's might be a tiny bit of
- 9:18
a hiccup. You know, I'll play some audio
- 9:20
later and you'll know that you probably
- 9:22
won't be able to notice.
- 9:24
And effectively, we're kind of taking
- 9:26
this and
- 9:28
or taking this part and just emitting it
- 9:31
directly after this audio has finished
- 9:33
playing. So, it looks seamless to the
- 9:35
user.
- 9:37
And yeah, so I guess at this point kind
- 9:40
of this is a lot of technical details.
- 9:42
I'll I mean the best way is just kind of
- 9:44
play a call and see
- 9:47
see this in action. So here I recorded a
- 9:50
call and let's just hear what happens.
- 9:55
>> This is Bo OBGYN Elise speaking.
- 10:00
>> Hi, I think I might be pregnant and I
- 10:02
want to schedule an ultrasound to
- 10:03
confirm.
- 10:06
>> That's exciting news. Can you provide
- 10:08
your name and date of birth?
- 10:12
>> Sure, it's Elise Trial.
- 10:14
Date of birth
- 10:15
2303.
- 10:21
>> Thank you. Are you a new patient with
- 10:23
us?
- 10:26
>> Yes.
- 10:29
>> Welcome to Bo OBGYN. Can I text you a
- 10:32
link to upload your insurance
- 10:33
information?
- 10:36
>> Sure.
- 10:39
>> Okay, sent. Let me know when you've
- 10:41
received it.
- 10:48
I've received your insurance
- 10:49
information. The earliest appointment I
- 10:52
have is Thursday, July 2nd at 10:00 a.m.
- 10:55
>> Actually, give me a second.
- 10:57
Let me check my calendar.
- 11:01
>> Sure, take your time.
- 11:04
>> Um do you have something for next week?
- 11:09
>> For next week, I can offer an ultrasound
- 11:11
with Dr. Avery Stone at North Clinic on
- 11:13
Tuesday, July 7th at 2:00 p.m. or 3:00
- 11:17
p.m. Do either of those work for you?
- 11:20
>> Yeah, 2:00 p.m. works.
- 11:23
>> Great. Your appointment has been booked.
- 11:26
We look forward to seeing you then.
- 11:29
>> Thanks. Bye-bye.
- 11:32
>> All right. Yeah, that's pretty much it.
- 11:34
Um
- 11:35
Yeah, you can kind of see our
- 11:38
all this like uh
- 11:40
streaming
- 11:42
and you know, a lot of things are
- 11:43
happening in the background and and you
- 11:44
know, this is what really makes like
- 11:46
voice agents interesting. And um there's
- 11:49
a lot of effort that can be done in the
- 11:51
harness to really kind of get a um
- 11:55
natural conversation, which is what
- 11:57
we're after.
- 11:59
Uh okay. Yeah, so I guess briefly, you
- 12:01
know, in the last part, I want to just
- 12:04
talk a little bit about Elise. So I
- 12:06
think Elise, you know, our headquarters
- 12:08
are in New York and kind of we're trying
- 12:10
to expand our presence here in the Bay
- 12:11
Area. Um we I think it's maybe like a
- 12:15
different style of company that I think
- 12:17
people are
- 12:19
like uh think of when they think about
- 12:21
AI startups in San Francisco. Where
- 12:23
we're actually very focused on um
- 12:26
just like helping people and helping
- 12:30
people where they need it, like kind of
- 12:32
the life's most critical areas. We work
- 12:33
on housing, health care and uh we're
- 12:37
doing really well and you know, here's
- 12:40
there's a link here to
- 12:42
um you kind of join our team and there's
- 12:44
going to we're going to be uh posting a
- 12:46
lot on Twitter, so you can follow us at
- 12:48
EliseAI as well.
- 12:50
Um yeah, that's that's it.
- 12:53
>> [applause]