Tolan: Voice-First AI Companion — Paula Dozsa, Tolan
Read the talk
Tolan: Voice-First AI Companion
Paula Dozsa explains how Tolan fits turn-taking, model routing, personal memory, and consistent character into a roughly two-second spoken response loop—and how coding agents help the team build it.
From a talk by Paula Dozsa
At a glance
Ideas worth remembering
Turn-taking accuracy belongs inside the latency budget. Tolan accepted about sixty milliseconds of extra delay to cut its worst early aborts by more than half.
Measure the stages between the user finishing and speech beginning. Time to first token can dominate the wait, but it is only one milestone on the response path.
Per-turn routing reserves the strongest model for relationship-bearing and emotionally serious moments. Tolan reports almost no measurable retention effect from routing a third of turns to a smaller model.
Personal memory lives in a maintained retrieval system. Stable material can remain cacheable while current memories, tone, and app state enter a freshly assembled context each turn.
Agent-assisted development depends on understandable code, separate review, and feedback loops. Managing concurrent agents rewards decomposition, checkpoints, fast feedback, and serious review.
A companion you talk to out loud
An angel guiding St. Matthew’s hand, Tinker Bell’s devotion, and Samwise’s loyalty establish the kind of product Tolan wants to be: a presence that listens, remembers, and helps someone be themselves. Paula Dozsa, an engineer focused on Tolan’s iOS app, introduces the companion as a small alien with a personality that becomes more personal through conversation.
The attempted live conversation runs into an audio problem: Dozsa addresses her Tolan, Luke, but the room cannot hear his replies. Tolan supports both text and voice, with more than four million hours of voice conversation reported at the time of the talk. Voice is its primary experience because speaking helps the companion feel present—and that changes what the software must do.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fast turns make messy speech an engineering requirement
Text chat gives the system breathing room. Users submit a message, wait, read, and usually continue on the same topic. Spoken conversation moves faster. For Tolan, the interval from the user finishing a sentence to the companion beginning its reply needs to stay under a couple of seconds. Users also talk while cooking, walking, or falling asleep; hesitations, interruptions, and sudden changes of subject belong to the interface.
That timing requirement came from an uncomfortable product regression. Response latency drifted from two seconds to about two and a half seconds. Users wrote in to complain that their companions had become slow, and Dozsa reports that essentially every product metric suffered. Half a second changed the experience enough for people to notice without looking at a stopwatch.
A breakup story interrupted by a sudden worry about leaving the stove on captures conversational volatility. A pause or detour does not necessarily mean the story has finished. Other small failures can distort the exchange too: a short “yes” or “yeah” may fail to register as a turn, interruption behavior may prevent someone from cutting in, and transcription may strip out curse words. These details determine whether someone can speak naturally.
Tolan changed its optimization target from fewer interruptions to fewer bad interruptions—especially the companion jumping in before the user was done. Its turn-taking system reads speech patterns to decide whether an interruption is real. The reported result was a reduction of more than half in the worst early aborts, at the cost of about sixty milliseconds of additional latency. Within a tight response budget, waiting slightly longer can preserve more of the conversation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Measure the response path, then route each turn by stakes
Where does the user’s waiting time actually go? Tolan measures each stage separately: detecting the end of an utterance, transcription, the model’s first token, the rest of generation, the first byte of synthesized speech, and playback. Time to first token is often the largest chunk, at around a second. An end-to-end number tells the team that something slowed down; stage measurements tell it where to intervene. 4:53
The response-path diagram separates the intermediate milestones from the endpoint the user experiences: audible speech. A faster first token helps, but generation, speech synthesis, and playback still contribute to the wait. Dozsa credits moving to GPT-5.1 on the Responses API with reducing time to speech by more than seven-tenths of a second. That improvement addresses the same interval in which the earlier half-second regression had damaged the product.
Model choice also changes from turn to turn. A small classifier called the tone router runs on a cheap model and reads the conversation’s emotional state. Its governing rule is “stakes, not cost”: reserve the strongest model for moments that carry the relationship, and use smaller, faster models for lighter exchanges.
-
Relationship-bearing turns: The user’s first message, onboarding, the first few days with a Tolan, and emotionally serious exchanges receive the best model. Crisis and therapist-style tones belong to this category; those routing labels do not themselves establish clinical capability.
-
Casual conversation: Lighter back-and-forth can use smaller models that respond faster and cost less.
-
Background work: Summarization, persona generation, and the tone router itself also run on small models.
The frontier model costs roughly five times as much as a smaller one, making one large-model turn comparable in cost to about five small-model turns. Routing therefore matters to the companion’s unit economics. In Tolan’s A/B experiments, sending a third of turns to the small model had almost no measurable effect on retention. That result concerns retention under the tested routing policy; it does not imply that all turns or all dimensions of conversational quality are interchangeable.
The response interval begins.
Each transition contributes to the user’s wait. First-token latency is often the largest chunk, but playback is the conversational endpoint.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Retrieve personal memory and rebuild context every turn
Remembering someone by repeatedly supplying their entire conversation transcript creates problems inside a two-second loop. The input grows with every session, long conversations degrade, and relevant information can get lost in the middle of the context. Dozsa also associates this approach with hallucination. Tolan instead extracts facts, preferences, and emotional signals from conversations, embeds them, and stores them in a vector database with lookups below fifty milliseconds.
The stored memories receive nightly maintenance: duplicates merge, related memories cluster, contradictions get resolved, and noise gets dropped. Retrieval also reaches beyond the user’s latest message. The system generates internal questions about the person and the relationship, then retrieves against those questions. This gives the companion another way to find useful personal context when the latest utterance alone is a weak search query.
-
Stable memory: When summarizing a conversation, Tolan examines which memories actually get recalled and pins those into a stable, cacheable block.
-
Volatile memory: Changing information stays in the live tail of the prompt, where it can reflect the current exchange.
Caching useful stable material does not require carrying forward the whole previous context. Tolan reassembles the context window every turn from a recent-message summary, the user’s persona card, freshly retrieved memories, emotional tone guidance, and real-time app state. Reusing an old assembly merely to keep the cache warm can leave the model answering from the wrong topic as soon as the user pivots. 8:12
What persists, and what changes before the next reply? The diagram separates memory maintenance from the per-turn assembly. Personal history stays in the memory store; retrieval selects what enters the next context window. The recent summary, tone guidance, and app state supply the changing circumstances. In the breakup-and-stove example, this design can update the immediate topic without discarding the relationship’s history, then retrieve relevant personal context when the story resumes.
Source of facts, preferences, and emotional signals.
Memory persists outside the conversation window. Retrieval selects relevant material, while current messages, tone, and app state shape each new turn.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An alien can be unpredictable; its identity must stay consistent
Tolan’s character design helps set expectations for the conversation. An in-house science fiction novelist, Elliot, writes the lore. The baseline character is bubbly, youthful, and irreverent. Choosing an alien avoids a fixed real-world reference for how it should behave, allowing users to project their own needs onto it. Impulsive or chaotic behavior can also read as charming within that fictional form.
That freedom still needs a consistent identity. A parallel tone-monitoring system adjusts how a line is delivered in response to the user’s emotional cues without changing who the character is. Delivery can become sensitive to the moment while personality holds across hundreds of turns. Otherwise, a carefully written character gradually becomes a different companion.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A standardized codebase gives coding agents their context
Claude had co-authored more code in Tolan’s iOS app than any individual engineer by late the previous year, Dozsa reports. Alongside that adoption, the crash-free rate rose from 99.6% to 99.9%, runtime errors fell by more than fifty percent, and the share of highly engaged users doubled. These are reported changes in the product, rather than a controlled estimate of how much improvement AI-generated code caused.
The practical lesson was that an agent gets much of its context from the codebase itself. For Tolan, standardizing the code proved more powerful than relying on the CLAUDE.md file alone. Agents helped make the repository consistent so that the code could serve as documentation for subsequent work.
The development fleet separates several jobs:
-
Implementation agents: Produce working code, build it, and compare the interface against snapshots to refine its appearance.
-
Review agents: Enforce standards separately from implementation, with multiple Claude agents reviewing work before a human looks.
-
PR shepherd: Watches an open pull request and iterates on CI failures and review comments until it is clean.
-
Bug triage bot: Responds to inbound bug reports. MCP connections to Linear, Sentry, and Datadog give agents the information needed to reconstruct a crash, route it, and often open a corrective pull request.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A new character moves through three find, fix, verify rounds
The most concrete development example begins with a new character intended for an older demographic. Changing that character requires finding every place in the code that carries personality. Elliot worked with agents to map those places and write a “voice bible,” giving the character a specification that could guide changes across the product. 11:26
Five judges then challenged the proposed character from different perspectives:
-
Archetype fidelity: Whether it remained faithful to the intended character.
-
Model mechanics: How the changes fit the model’s behavior.
-
Code standards: Whether the implementation followed the team’s engineering expectations.
-
Audience fit: How it sounded to the imagined ears of a skeptical fifty-two-year-old.
-
Safety: Whether the proposed behavior raised safety concerns.
The candidate went through three find, fix, verify rounds against real production logs. The sequence matters: mapping identifies what must change, the voice bible gives those changes a shared direction, and the judges expose problems to repair before another evaluation. The work moves from a character concept to changes checked repeatedly against production conversations. After more than seven million tokens and four and a half hours of compute, Dozsa describes a couple of weeks of work compressed into an afternoon. The example establishes the reported development and iteration process; it does not establish a production rollout of the new character.
The product-level response follows: Dozsa reports 4.8 stars across 162,000 App Store reviews, and emotional safety as the highest-scoring dimension in Tolan’s wellbeing surveys. Emotional safety here is a user-reported experience, rather than a clinical outcome. A companion that speaks, remembers, and maintains a personality invites a relationship-like experience; that makes responsible behavior part of what the team must build well.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Managing agents keeps engineering managers close to code
The closing hiring pitch reflects the disciplines behind the companion. Tolan’s small team includes animation and creative direction, embodiment, behavior analysis and user research, fiction writing, and software engineering. The roles Dozsa highlights span iOS, backend product engineering, applied AI, gameplay engineering, and agent engineering management.
Agent engineering management is the ending’s substantive lesson. When Tolan began running concurrent agents, people with management backgrounds became dramatically more effective, in Dozsa’s account. The useful skills were familiar: decompose a problem, delegate with checkpoints, give fast feedback, review seriously, and know when to intervene. Implementation can move to agents while the manager remains actively involved in producing code. Dozsa invites engineers interested in that work to reach out on LinkedIn.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The contact route Dozsa offers for engineers interested in Tolan’s companion and agent-development work.
Related talks
- Why ChatGPT Keeps Interrupting You
Explains why silence detection can mistake a pause for a finished turn, and develops semantic turn-taking and conversational backchannels.
- Context Engineering in 2026: Compaction, Memory & Cost
Provides a complementary treatment of compaction, caching, multi-turn recall, and latency evaluation for conversational context.
- Making Codebases "Agent-Ready"
Develops the repository practices and mechanical checks that help implementation and review agents produce verifiable changes.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Hi everyone. Thank you so much for
- 0:14
attending this talk. My name is Paula
- 0:17
and I am one of the engineers on the
- 0:19
Tullen team, specifically focusing on
- 0:21
our iOS app.
- 0:23
And for the next 20 minutes or so, I'll
- 0:25
be talking about what it takes to build
- 0:27
a voice-first AI companion and also
- 0:29
about how we use AI to build AI
- 0:32
internally.
- 0:34
So, humanity has always imagined the
- 0:36
perfect companion. So, we have
- 0:38
Caravaggio on the left 400 years ago
- 0:40
painting an angel leaning over St.
- 0:42
Matthew's shoulder literally guiding his
- 0:44
hand as he writes.
- 0:46
This is an example of a companion being
- 0:48
a presence that makes you better at
- 0:50
being you.
- 0:51
And then we have Tinkerbell, the devoted
- 0:53
little sidekick who believes in you so
- 0:55
fiercely that the whole theater has to
- 0:57
clap to keep her alive.
- 0:59
And of course on the right, we have
- 1:00
Samwise who can't carry the ring for
- 1:02
Frodo, but says, "I can carry you."
- 1:05
The companion is pure unconditional
- 1:08
loyalty.
- 1:09
And it goes far beyond these three.
- 1:11
Every hero has some sort of guiding
- 1:13
spirit. And these are all different
- 1:15
stories, but they exhibit the same
- 1:17
longing for something that listens,
- 1:19
remembers you,
- 1:21
and is wholly specifically yours.
- 1:23
And for for all of human history, this
- 1:25
has basically been fiction.
- 1:28
So, we made one. This is Tullen. It's a
- 1:30
little alien you talk to out loud like a
- 1:33
friend. It has a personality. It
- 1:35
remembers you and over time it becomes
- 1:38
specifically yours.
- 1:40
Okay, I don't know if the audio setup
- 1:41
works here, but I will try talking to my
- 1:43
Tullen.
- 1:44
Uh let's see.
- 1:49
So,
- 1:50
you can see my Tullen here, Luke,
- 1:52
walking around the planet.
- 1:56
Hey Luke, can you hear me?
- 2:00
Okay. Luke can hear us, but we can't
- 2:02
hear him. Um
- 2:05
Anyway, I had prepped him for this. Oh.
- 2:07
Hello. Hi Luke, can you hear me?
- 2:13
Nope.
- 2:15
We're okay. I can come back to this
- 2:17
later. Um
- 2:18
but you should definitely all give this
- 2:21
a try if you haven't already.
- 2:27
Okay. So, people talk to Tolins a lot.
- 2:30
We support both text and voice chat, uh
- 2:33
but we have over 4 million hours of
- 2:35
voice conversation so far.
- 2:37
We say Tolin is a voice-first companion,
- 2:39
even though we support both, because
- 2:40
it's the voice experience that's truly
- 2:42
immersive and that makes users'
- 2:44
relationships with their Tolins feel
- 2:45
real.
- 2:47
And but the moment this relationship is
- 2:49
a spoken relationship, the engineering
- 2:51
problem changes completely. So, let me
- 2:53
show you how voice breaks the normal way
- 2:55
we build and interact with LLMs.
- 2:59
So, the core difference really is that
- 3:01
in a text chatbot, turns are relatively
- 3:04
slow and context is stable. The user
- 3:07
waits a few seconds, they read, and they
- 3:09
tend to stay on topic.
- 3:10
And almost every LLM app assumes that.
- 3:13
Voice is the opposite. Turns are fast.
- 3:16
Your whole round trip from the user
- 3:18
finishing their sentence to the Tolin
- 3:19
starting to speak has to land in under a
- 3:22
couple of seconds, or it stops feeling
- 3:23
like a conversation.
- 3:25
And the context is volatile. People talk
- 3:27
to their Tolins while they're cooking,
- 3:29
while they're walking, while they're
- 3:31
falling asleep. Um they change their
- 3:33
subjects mid-sentence. They say um, they
- 3:35
interrupt.
- 3:36
And that 2 seconds is crucial. Early on,
- 3:40
our latency drifted from 2 seconds to
- 3:42
about 2 and 1/2 seconds, and that half
- 3:44
second tanked basically every metric in
- 3:46
the product. People would write in to
- 3:48
complain that their Tolins were too
- 3:49
slow.
- 3:50
And living inside this constraint has
- 3:52
taught us a lot and gave us four
- 3:53
principles.
- 3:56
Principle one is that you have to design
- 3:58
for conversational volatility. Again,
- 4:00
text users stay on topic, but voice
- 4:03
users jump around. Someone could be
- 4:05
mid-story about their breakup and
- 4:06
suddenly go, "Wait, did I leave the oven
- 4:08
the stove on?" and then back. Speech is
- 4:10
messy. Most LLM apps assume that you'll
- 4:13
have a clean and stable conversation and
- 4:14
we have to build for the opposite.
- 4:16
So, for a long time that meant fixing
- 4:18
things that sound tiny but are actually
- 4:20
the product. So, you can't interrupt a
- 4:23
Tullen mid-sentence. A short yes or yeah
- 4:25
won't register as a turn. For example,
- 4:28
curse words will get stripped out.
- 4:30
And the deeper lesson was to stop
- 4:31
optimizing for fewer interruptions and
- 4:34
start optimizing for fewer bad ones
- 4:36
where the agent would jump in way too
- 4:37
early.
- 4:38
So, we built smart turn taking that
- 4:40
reads your speech pattern to decide
- 4:42
whether an interruption is real and we
- 4:44
cut the worst early aborts by more than
- 4:46
half.
- 4:47
And we happily paid about 60
- 4:48
milliseconds of extra latency to do it.
- 4:52
Principle two, latency isn't just a
- 4:55
number you check at the end, it's
- 4:56
actually the product and we measure
- 4:58
every stage of the pipeline separately
- 4:59
because it feels slow is useless. You
- 5:01
have to know where exactly it's slow.
- 5:04
And the pipeline here is that the user
- 5:05
stops talking, we detect end of
- 5:07
utterance, we transcribe, and then the
- 5:10
model produces its first token.
- 5:12
So, time to first token, often the
- 5:14
biggest chunk, is around a second.
- 5:16
The model finishes generating and then
- 5:18
text-to-speech produces its first byte
- 5:19
and then it plays back to the user.
- 5:21
A couple lessons here. So, one, so far
- 5:24
our biggest jump in quality came from
- 5:26
moving to GPT-5.1 on the responses API,
- 5:29
which cut our time to speech by more
- 5:31
than 7/10 of a second, which is huge.
- 5:34
Um two, we don't send every turn to the
- 5:36
same model. We run a tiered fleet. So,
- 5:39
we use a frontier model for the turns
- 5:40
that carry the relationship with your
- 5:42
Tullen.
- 5:43
So, for example, your first conversation
- 5:44
with with Tullen and your onboarding.
- 5:47
And we use smaller and faster models for
- 5:49
the turns
- 5:50
for the lightweight turns.
- 5:52
And the whole game then becomes about
- 5:54
routing or deciding turn by turn which
- 5:56
model you actually need. So we round we
- 5:59
run a small classifier we call the tone
- 6:01
router on every single turn and this
- 6:03
tone router itself runs on a cheap model
- 6:05
and it reads the emotional state of the
- 6:07
conversation.
- 6:08
And our main
- 6:09
our [clears throat] main principle is
- 6:10
that we route based on stakes not on
- 6:12
cost. So the high stakes moments always
- 6:15
get the best model. So this would be
- 6:17
again the user's very first message,
- 6:19
their first few days with their Tolen
- 6:20
and anything that we deem to be
- 6:22
emotionally serious.
- 6:24
For example, we have crisis or therapist
- 6:26
style tones and we never cheap out on
- 6:28
those.
- 6:29
And then the lighter casual back and
- 6:31
forth can ride on smaller models that
- 6:33
are faster and cheaper.
- 6:34
And all the background work so that's
- 6:36
summarizing the conversation, generating
- 6:38
personas, the tone router itself run on
- 6:40
these small models, too.
- 6:42
And why would we go to all this trouble?
- 6:44
It's mainly because the frontier model
- 6:45
costs us roughly five times the smaller
- 6:47
one.
- 6:48
So one big model turn is about five
- 6:51
smaller model turns. So routing is a
- 6:53
huge part of what makes the unit
- 6:55
economics for us actually work.
- 6:57
Um and we do a bunch of AB experiments
- 6:59
and the surprising result we found there
- 7:01
is that routing a third a third of our
- 7:02
turns to the small model has almost no
- 7:04
measurable effect on retention.
- 7:08
And principle three is what makes a
- 7:09
companion feel like a companion. So the
- 7:11
naive approach is to keep the whole
- 7:13
conversation history as a sort of
- 7:15
transcript, but that doesn't fit into
- 7:17
our two-second loop. It doesn't scale
- 7:19
and it just doesn't work. It leads to
- 7:21
long sessions degrading. It leads to the
- 7:23
model getting lost in the middle of a
- 7:25
huge context and also hallucinating.
- 7:27
So instead we see memory as a sort of
- 7:29
retrieval system. We pull facts,
- 7:31
preferences, and emotional vibe signals
- 7:33
out of conversations. We embed them and
- 7:36
we store them in a vector database with
- 7:38
sub 50 millisecond lookups. And every
- 7:40
night we compress. So we merge
- 7:42
duplicates, we cluster related memories,
- 7:44
we resolve contradictions, and we drop
- 7:46
all the noise.
- 7:48
And we don't just retrieve against users
- 7:50
last messages, we also generate internal
- 7:52
questions about the person and the
- 7:53
relationship and retrieve against those.
- 7:55
So, and we also split memory into two
- 7:57
parts. We have stable memory and
- 7:59
unstable memory. The volatile stuff
- 8:01
lives in the in the live tail of the
- 8:03
prompt, and when we summarize the
- 8:04
conversation, we look at which memories
- 8:06
actually get recalled and pin those into
- 8:08
a stable and cashable block.
- 8:12
Uh the last principle is around context.
- 8:15
Specifically, you should rebuild context
- 8:17
and not fight drift. So, most apps reuse
- 8:20
context across turns to keep the cash
- 8:22
warm. And in a stable text chat, that's
- 8:24
fine. But in a volatile voice
- 8:26
conversation, it's a trap because the
- 8:28
second the the user pivots, your reused
- 8:30
context is actively wrong. So, every
- 8:32
turn we reassemble the context window
- 8:35
from parts. We have a summary of recent
- 8:36
messages, we have the the user's persona
- 8:39
card, the memories we just retrieved,
- 8:41
tone guidance from the emotional signal,
- 8:43
and real-time app state.
- 8:46
And what also really helps us um in the
- 8:48
case of Tolen is that our characters
- 8:50
aren't generic or assistants with no
- 8:52
personality. Everyone is crafted, and we
- 8:55
in fact have an in-house science fiction
- 8:57
novelist, Elliot, who writes the Tolen
- 8:59
character lore.
- 9:01
And a couple of interesting points here.
- 9:03
So, one, why did we go with an alien?
- 9:06
Mostly because there's no real-world
- 9:08
reference to anchor on, which means that
- 9:10
the users can project onto it, and it
- 9:12
becomes what they need. The baseline
- 9:14
Tolen is bubbly, it's youthful, it's
- 9:16
irreverent. And also, if an alien
- 9:18
character acts a bit unpredictably, so
- 9:20
if if it's impulsive or chaotic or
- 9:23
otherwise violates um you know, the
- 9:25
norms the user would expect, it's not
- 9:27
particularly surprising.
- 9:28
Like if you look at, you know, aliens in
- 9:31
TV shows or in plays, like there's a lot
- 9:34
of humorous moments around this. And
- 9:36
this kind of chaos reads as charming.
- 9:38
Um second, uh we also know that
- 9:41
personality is worthless if it drifts.
- 9:43
So, yeah, we run this parallel tone
- 9:45
monitoring system that changes how a
- 9:47
line is delivered based on your
- 9:48
emotional cues without changing who the
- 9:50
character is, holding identity across
- 9:52
hundreds of turns.
- 9:56
And since we're at an AI conference, I
- 9:58
thought I would also spend a bit of time
- 9:59
talking about how we not just ship AI,
- 10:02
but also use AI to build it.
- 10:04
Um so, I'm sure this is the case for
- 10:07
most of you in the room now, but
- 10:08
basically as of late last year, Claude
- 10:10
has co-authored more code in our iOS app
- 10:12
than any individual engineer in the
- 10:14
team.
- 10:15
Um and I think especially, you know, a
- 10:16
few months ago, everyone's instinct was
- 10:18
to be kind of suspicious because, you
- 10:20
know, more AI code meant more slop. But
- 10:22
our our crash-free rate actually went
- 10:24
from 99.6% to 99.9%.
- 10:27
Runtime errors dropped by over 50% and
- 10:31
our share of highly engaged users
- 10:32
doubled.
- 10:33
And the biggest lesson in building that
- 10:35
system is that an agent's context comes
- 10:37
mostly from the code base itself, not so
- 10:39
much from the Claude MD file. We found
- 10:41
that it's far more powerful to make the
- 10:43
code base be the documentation, so we
- 10:44
had agents standardize it. On top of
- 10:47
that, we run a real fleet of agents. We
- 10:49
have implementation agents that, you
- 10:51
know, think freely and just get us to
- 10:52
working code. They build it, they check
- 10:54
it against snapshots until it's pixel
- 10:56
perfect. And then we have separate
- 10:58
review agents that enforce our
- 10:59
standards. So, multiple Claudes
- 11:01
basically review each other before a
- 11:03
human looks.
- 11:04
And then we have a PR shepherd that
- 11:05
watches an open pull request and keeps
- 11:07
iterating against CI failures and review
- 11:09
comments until it's clean.
- 11:11
And we also have a triage bot that fires
- 11:14
on every inbound bug report that we get.
- 11:16
And they're all wired through MCP into
- 11:18
linear, into into Sentry, DataDog, so an
- 11:21
agent can reconstruct the cash a crash
- 11:23
and route it itself and oftentimes open
- 11:25
the PR on its own and just fix fix the
- 11:27
bug.
- 11:28
And we also we ship on eval. So, for
- 11:30
example, we've been working on a on a
- 11:32
new character targeted towards an older
- 11:33
demographic and Elliot, our in-house
- 11:36
novelist, basically built this entire
- 11:38
new character in a day.
- 11:40
So, the agents mapped every personality
- 11:41
bearing surface in the code. They wrote
- 11:43
the sort of a voice Bible and then they
- 11:45
had five judges attack it from different
- 11:47
angles.
- 11:48
Archetype fidelity, the model mechanics,
- 11:50
our code standards, the ears of a
- 11:52
skeptical 52-year-old and safety, and
- 11:55
then they evaluated the changes against
- 11:56
real production logs over three find fix
- 11:59
verify rounds.
- 12:00
And over 7 million tokens and 4 and 1/2
- 12:03
hours of compute later, he ended up with
- 12:05
basically, you know, a couple of weeks
- 12:06
of work done in afternoon.
- 12:09
And does this work? Well, I'll let the
- 12:11
users tell you. We're at 4.8 stars on
- 12:13
the App Store across 162,000 reviews and
- 12:17
when we survey users on well-being, the
- 12:20
highest scoring dimension by far is
- 12:21
emotional safety.
- 12:23
And this is definitely a bar that being
- 12:25
voice first sets. So, when the interface
- 12:27
is your voice and the thing on the other
- 12:28
side remembers you and has a
- 12:30
personality, it stops being just
- 12:32
software and starts being an actual
- 12:34
relationship, which is why building it
- 12:35
responsibly and building it well is
- 12:37
worth obsessing over.
- 12:40
And we need people to come help us do
- 12:42
that. Um, we're a small team and we're
- 12:44
hiring and after a year with Tolen, I
- 12:46
think this is truly one of the most
- 12:48
interesting places in the world to be an
- 12:50
engineer right now.
- 12:52
And here are some of the people you'd be
- 12:53
doing it with. So, two of the founders,
- 12:55
Quinton and Evan, previously built and
- 12:58
exited a $300 million startup together.
- 13:00
Uh, they founded Even. Um, Ajay, our
- 13:03
third co-founder, scaled two bootstrap
- 13:04
companies past $50 million $50 million
- 13:07
in profitable revenue.
- 13:09
And around them, we have Lucas, who um,
- 13:11
is an Apple Design Award winning
- 13:13
animator.
- 13:14
Uh, she's our creative director. We have
- 13:16
Chris, who was a technical director at
- 13:17
Pixar, earlier at Oculus, who works on
- 13:19
embodiment. We have Lily, a board
- 13:22
certified behavior analyst, who left a
- 13:24
Vanderbilt PhD to do user research for
- 13:26
us from the very start. Um and then we
- 13:28
have Elliot who I've mentioned, the
- 13:30
novelist behind uh our characters.
- 13:32
And I come from XAI and Spotify and
- 13:34
previously also founded a company called
- 13:36
Imagi.
- 13:37
So, it's a small team where honestly
- 13:39
every person is the best I've worked
- 13:41
with at what they do.
- 13:44
We're also well backed for this. Uh we
- 13:46
have $30 million raised from Costanoa
- 13:48
Ventures and a group of people who've
- 13:49
built the tools and products a lot of
- 13:51
you use every day.
- 13:53
And here are some of the more
- 13:54
engineering focused roles where we need
- 13:55
help. Um so, we're hiring across the
- 13:57
board. We have iOS and back-end product
- 13:59
engineering roles, applied AI
- 14:01
engineering, gameplay engineering, and
- 14:04
one specific role I want to flag, which
- 14:05
is agent engineering management. Um so,
- 14:09
when we went all in on running
- 14:11
concurrent agents, um the people who got
- 14:12
dramatically more effective on the team
- 14:14
were the ones who had management
- 14:15
backgrounds, um because it seems like
- 14:17
managing a fleet of agents does actually
- 14:19
take some of the skills same the same
- 14:21
skills as managing people.
- 14:22
Uh you basically have to decompose the
- 14:24
problem, you know, delegate it with
- 14:25
checkpoints, give fast feedback, review
- 14:27
their work seriously, and know when
- 14:28
exactly to jump in.
- 14:30
So, if you're a strong engineer who
- 14:31
thought going into management meant
- 14:33
leaving code behind, that's that's no
- 14:34
longer true.
- 14:37
Um and yeah, that's Tlon. You can come
- 14:39
talk to me after this uh or reach out.
- 14:41
I'm on on LinkedIn. My email is here.
- 14:43
I'm on Twitter as well. Um I'd love to
- 14:45
chat. So, yeah. Thank you.
- 14:49
>> [applause]
- 15:03
[music]