Building ambitious software — Jonathan Kelley, Dioxus Labs & Cognition
Read the talk
Building Ambitious Software in the Age of AI
Jonathan Kelley’s Dioxus team learned that coding agents can absorb Rust’s mechanical difficulty and accelerate specialized integrations, debugging, and release work. The harder problems—choosing architecture, designing meaningful tests, communicating intent, and deciding what deserves to ship—remain stubbornly human.
From a talk by Jonathan Kelley
At a glance
Ideas worth remembering
A large volume of plausible code is not progress if maintainers cannot justify merging it; unchecked agent output can become a “slop cannon.”
Rust’s difficulty can become an advantage when agents absorb borrow-checker work and edge cases while developers retain its structural constraints.
Agents pay off as patient specialists on knowledge-heavy integrations and specification-driven debugging, especially when architecture and desired behavior are already clear.
Release checks, backports, documentation audits, and fuzzing harnesses are strong automation targets; selecting meaningful end-to-end tests still requires human judgment.
Code is cheap, but quality is not. Architecture, explicit intent, verification, and careful review become more important as implementation accelerates.
Dioxus started as a bet on one cross-platform stack
In 2021, Jonathan Kelley spent his final undergraduate summer starting Dioxus, a cross-platform application framework written in Rust. The proposition was appealingly direct: build web, desktop, and mobile interfaces with Rust, HTML, and CSS instead of moving among platform-specific languages, IDEs, and build systems. Rust would produce native applications without a virtual machine, JavaScript runtime, or inter-process bridge, while familiar web markup and React-inspired reactivity would preserve access to web concepts and tooling. 0:22
The easy-sounding abstraction concealed an enormous implementation burden. In the ecosystem available at the time, Dioxus could not simply assemble ready-made pieces for reactivity, font rendering, hot reloading, and bundling. The team had to build much of the application stack itself. “Building a web browser” became one of the necessary intermediate steps rather than an adjacent research project. 1:53
Five years later, Dioxus supported the major capabilities in that original plan: shared full-stack and native code, cross-platform rendering, hot reload, and bundle splitting. Kelley reported nearly 37,000 GitHub stars, millions of downloads, and an estimated cumulative audience above 200 million across applications including AI assistants, voting software, data-science tools, and satellite collision-avoidance systems. The end-user figure is an estimate presented in the recording, not a measured deployment census. 2:29
The developer experience was deliberately compressed around a conventional Rust project: fewer files, unified tooling, optimized assets, little platform-specific code, and a main.rs entry point. That simplicity mattered because the product was not only its visible output. Developers had to understand the project structure, compile it across platforms, and build businesses against its APIs.
Two projects show how far the team pushed the stack. Blitz is a lightweight HTML and CSS renderer built from a browser-grade CSS engine extracted from Firefox, a custom HTML/DOM implementation, and a hybrid GPU pipeline. Kelley reported Blitz applications below five megabytes and below 50 megabytes of runtime memory. Subsecond watches Rust, C, or C++ source files, recompiles only the changed portions, and patches the running native application in place in roughly 100 milliseconds; the same approach extends to Rust compiled to WebAssembly. 3:59
Those systems established the quality culture that shaped the later agent experiment. A small team had spent years reading every contribution, maintaining an ambitious release cadence, and writing essentially all of Dioxus by hand. Generated code would be judged against that existing standard, not merely against whether it compiled or demonstrated a feature.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fast Rust generation produced a backlog, not a release
The team had been skeptical that agentic coding and high-quality systems work belonged together. That changed when coding agents became markedly better at Rust. The team responded by exhausting its coding-agent subscriptions and generating tens of thousands of lines for long-desired features, integrations, and bug fixes. Almost none of that output cleared the question that mattered: should this be merged? 5:37
The code remained in drafts because implementation volume had outrun reviewable quality. Kelley calls this failure mode becoming a “slop cannon”: an agent can rapidly produce plausible changes, but every unmerged change still consumes attention, creates uncertainty, and competes with real maintenance work. The observable result was not a faster product. It was a much larger queue of code that the maintainers did not trust.
The postmortem produced a surprising reversal. Dioxus had spent years making Rust easier for people through readable code, good tools, and helpful errors. Agents were less sensitive to that ergonomic work. Instead, they handled Rust’s mechanical burden directly: following edge cases, satisfying the type system, and repeatedly negotiating with the borrow checker. The learning curve the team had tried to flatten for humans became useful because an agent could absorb it while preserving Rust’s compile-time constraints. 7:00
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ambitious software needs a different quality bar
The lessons apply most strongly to foundational software, where users build on the code itself. Research prototypes can prioritize discovery, and an application may hide its internals behind a polished interface. Dioxus has to keep working, remain fixable when it fails, preserve APIs across patch releases, and maintain accurate examples, documentation, tests, and benchmarks. A shortcut in the framework becomes somebody else’s broken build or stalled business. 8:00
Velocity therefore compounds through the existing architecture. New work lands on a “substrate” made of module boundaries, APIs, tests, and maintenance conventions. A sound substrate makes later features easier to add and repair. A poor one turns every additional feature into another dependency on an unstable base. Dioxus still has dozens of desired features, but shipping them faster cannot justify breaking APIs relied on by millions of users.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Agents work well as patient specialists
The first strong use case was not indiscriminate feature generation. It was knowledge-heavy engineering. A three-person core team cannot retain every detail of every operating system, build system, runtime, language, API, binary format, and implementation quirk. Coding agents can patiently search documentation, inspect bespoke interfaces, examine binaries, and compare unfamiliar implementations without becoming tired of the investigation. 9:56
A concrete example was adding deeply integrated Kotlin and Swift plugins to Dioxus’s build system. Comparable native-module infrastructure had taken years of handwritten development elsewhere. With agents, Dioxus shipped its version in roughly two to three weeks. Kelley recalls that the initial implementation may have arrived in about a day; the remaining time—roughly two weeks—went into designing test cases and exercising the integration on real devices. Generation accelerated the first draft, but verification still dominated the path to a shippable feature. 10:56
Blitz supplied another good fit. CSS painting and layout bugs often require reading specifications and checking how established browser engines handle obscure cases. An agent can retrieve the relevant rule and compare Chrome, Safari, or WebKit behavior while the engineer stays focused on the local rendering problem. That reduces the pressure to insert a quick workaround simply because investigating the correct behavior would otherwise take too long.
These examples share a useful boundary: the system already has an architectural destination, while the agent supplies patient research and implementation labor. The agent makes doing the careful version cheaper. It does not decide whether the integration belongs in the product or how the surrounding system should evolve.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Release chores scale better than test judgment
The less glamorous use case may be more valuable: agents can repeatedly perform work that a small team cannot afford to do by hand. Checking that a release tarball extracts into the correct directory or confirming that examples and documentation still match the code consumes time that could otherwise go into architectural work. Editor extensions illustrate the harder edge of this problem: a team may struggle to retest an extension after every patch, but verifying that it installs in Zed and works during actual use is difficult for agents too. Automating release chores does not by itself solve that end-to-end check. 12:15
Three recurring mechanisms paid off:
- Release verification: Agents can walk release checklists and inspect packaging details consistently.
- Maintenance across branches: They can backport bug fixes onto stable releases, reducing the manual cost of supporting shipped versions.
- Documentation audits: They can find missing examples and stale comments, including comments that still describe behavior changed by a human code edit.
Kelley links this automation to a concrete operational change: the latest Dioxus version had received more patch releases than any previous one, with releases arriving weekly or multiple times per week. The talk does not isolate which automation contributed how much, but the team became willing to release at a cadence it previously considered too risky.
Testing exposed the boundary of this success. An agent can write a test for nearly any API, but that does not mean it has chosen a consequential behavior. Given a constructor, it may simply test that constructor. Foundational end-to-end behavior requires product knowledge, an appropriate test interface, and a runner capable of reproducing the environment. Humans still enumerate the important conditions and build much of that testing structure. 14:21
Agents remain useful as a sounding board for candidate conditions and edge cases, and they excel at constructing fuzzing harnesses. A fuzzing harness repeatedly drives software with millions of ordinary, malformed, or adversarial inputs to expose crashes and assumptions that hand-written examples miss. Here the target behavior is already clear—feed varied inputs into a defined interface and report failures—so the agent can handle the repetitive harness construction without being asked to decide what product behavior matters most.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Architecture and review remain the bottleneck
Once implementation becomes cheap, architecture consumes more of the schedule. Agent contributions inherit the quality of the system they enter: a poor foundation still produces poor contributions, even when code arrives at exceptional speed. When a feature does not fit, an agent may readily undertake a huge refactor or redesign the architecture rather than hesitate over the scope of the change. Humans and agents can both write spaghetti code; agents can now write it faster. Dioxus consequently spends much of its development time deciding which future features the architecture must accommodate and how today’s interfaces should leave room for them. Kelley expects architecture to take an even larger share of engineering time as implementation improves, provided developers communicate their intent clearly. 16:03
What separates generated code from shipped software? The diagram synthesizes the practices described here: architectural intent shapes implementation, tests and real-device checks examine behavior, and maintainers read the code they ship. Its arrows show relationships rather than a fixed release order. In the Kotlin and Swift integration, an initial implementation arrived in about a day, while test cases and real-device testing took roughly two weeks. That contrast makes the remaining work visible: producing code quickly still leaves substantial work to decide whether it belongs in the product.
Define how the feature fits today and how the system may evolve.
An explanatory synthesis of the talk’s practices, not a prescribed sequence. Architecture shapes implementation; testing examines behavior; maintainers review code and decide what clears the quality bar.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Reading code is still the job
Dioxus still reviews every pull request line by line, using AI review as another bug-finding input rather than a replacement for maintainers reading what they ship. Open-source contributions make this especially important: contributors often want one bug fixed or one feature accepted, while maintainers have to consider the codebase’s long-term shape. A locally successful patch can still be glued into the wrong abstraction.
Intent must also travel through text. Agents cannot infer the maintainers’ unwritten model of how Dioxus should evolve, and contributors may not communicate that model clearly. Prompt quality therefore affects implementation quality in a practical sense: a better description carries more architectural decisions, constraints, and acceptance criteria into the agent’s work. It is not mind reading, and it does not remove the need to inspect the result.
The technical conclusion is deliberately old-fashioned: reading code eventually matters more than writing it. Coding agents make lines of code cheap, but they do not make quality cheap. Engineering value remains in choosing an elegant shape for a complex system, anticipating how it will change, preserving flexibility, and taking responsibility for what ships. Faster implementation raises that bar because weak decisions can now propagate much faster. 18:03
Kelley closes by connecting that work to Cognition: Cognition acquired Dioxus, and the Dioxus team joined the company. He also notes that Cognition is hiring people to work on the next generation of software-development tools. 18:41
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
From the talk
The Rust full-stack application framework discussed throughout the talk. Its README shows the cross-platform project structure, hot-reloading workflow, native integrations, examples, and contribution paths.
Related talks
- Full Walkthrough: Workflow for AI Coding — Matt Pocock
Develops the planning, vertical-slice, testing, review, and architectural practices needed to turn coding-agent output into maintainable changes.
- Why Rust is the Ideal Language for Vibe-Coding
Offers another perspective on Rust’s suitability for AI-assisted coding, alongside Dioxus’s experience of agents handling Rust’s learning curve.
- Harness Engineering is not Enough: Why Software Factories Fail
Pairs Kelley’s experience with a discussion of why harness engineering alone may be insufficient for software factories.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
Hello, my name is Jonathan Kelly and
- 0:15
today we're going to talk about what it
- 0:17
means to build ambitious software in the
- 0:19
age of AI.
- 0:22
5 years ago, I made the first commit
- 0:24
ever to a project called Diosis. I used
- 0:28
the last summer I had as an
- 0:29
undergraduate and instead of getting an
- 0:31
internship at Google or doing research
- 0:33
in AI like many of my friends at the
- 0:35
time, I spent it exploring an idea I had
- 0:38
for a crossplatform app framework
- 0:40
written in the Rust programming
- 0:42
language.
- 0:44
In 2021, Rust was still pretty niche,
- 0:47
but the ecosystem was growing, the
- 0:49
tooling was improving, and the pitch of
- 0:51
native performance, a solid type system,
- 0:54
and simple cross compilation really sold
- 0:56
me. It's extremely nerdy.
- 1:00
The idea for Diosis was straightforward.
- 1:03
What if we had an crossplatform app
- 1:05
framework? Instead of waiting through
- 1:08
dozens of tool chains, programming
- 1:09
languages, and idees, what if we simply
- 1:12
wrote our all of our apps in Rust using
- 1:16
HTML and CSS as the markup language?
- 1:19
This was back in the day 2021 where
- 1:21
React Native was janky, Flutter was too
- 1:24
slow, and neither performed well with
- 1:26
native APIs. On the flip, with Rust, we
- 1:30
could build native abstractly with no
- 1:32
VM, no IPC, no JavaScript. And if we
- 1:36
used a little bit of HTML and CSS for
- 1:39
the UI and take some inspiration from
- 1:41
React for the reactivity, we could reuse
- 1:44
vast amounts of web components and web
- 1:46
tooling. The goal was an extremely
- 1:48
powerful app framework that was still
- 1:50
quite familiar to the average developer.
- 1:53
Sounds easy, right? Well, as they say,
- 1:57
we choose to build an app framework from
- 1:58
scratch, not because it's easy, but
- 2:00
because we thought it would be easy.
- 2:03
In reality, trying to challenge React
- 2:04
Native and Flutter is extremely
- 2:06
ambitious. In 2021, there were very few
- 2:09
off-the-shelf components you could use
- 2:11
to build Dioxis. Everything from
- 2:13
reactivity to font rendering to hot
- 2:16
reloading and application bundling had
- 2:18
to be built from scratch. There's
- 2:20
nothing we could use. For us, tasks like
- 2:23
building a web browser were just
- 2:25
necessary steps along the way.
- 2:29
Now today in 2026, Diosis has achieved
- 2:33
and far surpassed its original mission.
- 2:36
We support all the features we
- 2:37
originally set out to build from
- 2:39
crossplatform support to native
- 2:41
rendering to Rust hot reload and bundle
- 2:44
splitting. We've basically reinvented
- 2:47
and improved the entire app development
- 2:49
stack.
- 2:50
Users can ship a powerful full stack web
- 2:53
application in the exact same codebase
- 2:56
sharing components as their iOS and
- 2:58
Android apps.
- 3:00
The Daxis project now has nearly 37,000
- 3:04
stars on GitHub with millions of
- 3:05
downloads. Apps built in Dioxis are
- 3:08
rolled across the globe with a
- 3:11
cumulative estimate of over 200 million
- 3:12
end users. Users have built things like
- 3:15
AI assistance, software for voting, data
- 3:19
science tools, and even collision
- 3:21
avoidance system for satellites in
- 3:23
space.
- 3:25
We've put a ton of effort into making
- 3:27
Dioxis as userfriendly as possible.
- 3:30
Fewer files, build tooling, hot
- 3:32
reloading, asset optimization,
- 3:35
everything you need to easily ship
- 3:36
across all platforms.
- 3:38
Because Dioxis apps are written in Rust,
- 3:40
they are structurally very simple. You
- 3:43
rarely need to drop into platform
- 3:45
specific code because all rest projects
- 3:48
are alike. It's very easy for developers
- 3:50
to dive into a new project. You can
- 3:52
completely skip annoying build system
- 3:53
setup. All you need is a main.rs to get
- 3:57
started.
- 3:59
One of the most ambitious goals we had
- 4:01
for ship lightweight but fully featured
- 4:05
HTML and CSS rendering engine called
- 4:07
Blitz. We extracted the browser grade
- 4:09
CSS engine out of Firefox, built our own
- 4:12
HTML DOM, and developed a hybrid GPU
- 4:15
rendering pipeline.
- 4:17
Compared to Electron apps, which are RAM
- 4:19
and storage hogs, Blitz apps are
- 4:21
lightweight, coming in at less than 5
- 4:23
megabytes bundle sizes and consume less
- 4:25
than 50 megabytes of RAM at runtime. And
- 4:28
they're pretty cool. You can write your
- 4:29
own custom components, spinning cubes,
- 4:31
you can customize the browser however
- 4:32
you want. It's a very cool project.
- 4:35
We also worked on a tool called
- 4:36
Subsecond uh which is our generic hot
- 4:39
reload engine for Rust, C and C++.
- 4:42
Subsecond watches your code for edits,
- 4:44
recompiles parts of the code that
- 4:46
changed and patches the running app in
- 4:48
place all in 100 100 milliseconds. This
- 4:51
was an incredibly difficult technical
- 4:53
challenge and is the only hot reload
- 4:55
engine for native compiled code to have
- 4:58
such wide language and runtime support.
- 5:00
It works on every major system and even
- 5:02
the web is compiled into web assembly.
- 5:05
No one has done this before
- 5:09
because these projects we've worked on
- 5:11
along the way over the past 5 years are
- 5:13
incredibly ambitious and are the result
- 5:15
of a tiny but
- 5:20
over every line of code by with our own
- 5:23
two eyes and maintain a frequent but
- 5:25
ambitious release cadence.
- 5:28
The most amazing thing, every line of
- 5:30
code in Dioxis until very recently has
- 5:33
been painstakingly written by hand.
- 5:37
Why do I say recently? Well, if you
- 5:40
aren't aware, software engineering and
- 5:42
development has taken a massive turn in
- 5:44
the past 6 months. AI coding agents got
- 5:48
really, really good. And specifically,
- 5:50
they got really good at Rust. Our team,
- 5:53
a bunch of cracked rest engineers, has
- 5:56
been quite skeptical of AI for a long
- 5:58
time. We had not felt the AGI, so to
- 6:01
speak. And we definitely weren't using
- 6:03
AI in our day-to-day work. We thought
- 6:05
the two things were incompatible,
- 6:07
shipping high quality code
- 6:10
tools. Seeing them get really good at
- 6:12
Rust was a huge surprise to us. So, we
- 6:15
were finally excited.
- 6:18
With this newfound excitement, we
- 6:20
started building our team. maxed out our
- 6:22
cloud code subscriptions, turned out
- 6:24
tens of thousands of lines of Rust, and
- 6:26
built all sorts of features we had long
- 6:27
wished to have. Unfortunately, very
- 6:30
little of the code cleared our quality
- 6:32
bar of should we merge this in.
- 6:35
Thousands of lines of new features, bug
- 6:37
fixes, and integrations we had wanted
- 6:40
for years sat there and draft and
- 6:42
continue to sit there and draft. We
- 6:45
definitely did not know how to properly
- 6:47
wield these tools and it was way too
- 6:49
easy to become what we call a slop
- 6:52
cannon.
- 6:54
Um, so we reflected a bit and studied
- 6:57
what worked and what didn't.
- 7:07
to read, easy to write, good tools, good
- 7:10
error messages. The coding agents
- 7:12
generally don't care about this. We
- 7:14
tried to make Rust easel
- 7:20
with Dioxis still
- 7:22
because Rust is harder to write. The
- 7:24
coding agents deal with the development
- 7:26
burden for you. They handle the edge
- 7:28
cases and they fight the borrow checker,
- 7:30
saving you from the cognitive burden of
- 7:32
writing Rust apps. The learning curve
- 7:35
which we fought to reduce is now a
- 7:38
feature.
- 7:40
So throughout the process of adopting
- 7:43
the coding tools to work on Dioxis, we
- 7:45
learned a wide array of lessons.
- 7:48
Many of the thing the coding agents do
- 7:50
really well today and many things they
- 7:52
just aren't there yet. So the next
- 7:55
couple slides I want to talk about some
- 7:56
of the things we learned and what it
- 7:58
means to to build ambitious software
- 8:00
projects in the age of agent coding.
- 8:05
Um, it's important to talk about first
- 8:07
what it means to build an ambitious
- 8:09
software project. Um, there's many
- 8:12
different types of software out there.
- 8:14
Depends on what you ship every day. Um,
- 8:17
you might be doing research and the
- 8:20
quality of your code isn't the most
- 8:21
important thing. You might be doing
- 8:22
prototyping code and and iterating fast,
- 8:25
moving quickly is important. You might
- 8:27
be building applications which people
- 8:29
don't see the code internally. They just
- 8:31
see what it looks like on the outside.
- 8:34
But for us and for Diosis,
- 8:37
we care about a few different things.
- 8:40
Primarily, of course, we care that our
- 8:42
code works all the time and that if it
- 8:44
breaks, we can easily fix it. I think
- 8:47
this is something people don't think
- 8:48
about enough these days that you need to
- 8:50
continue to build easily maintainable
- 8:52
code and like the velocity that you ship
- 8:56
lays down on this substrate that you've
- 8:58
built and if the substrate isn't good,
- 9:00
nothing you build on top is going to be
- 9:02
good.
- 9:03
Secondarily, we care about shipping new
- 9:06
features.
- 9:07
Our road map is really long. It extends
- 9:10
into the far future. There's dozens of
- 9:13
features we still have yet to build for
- 9:15
Dioxis. And we want to ship these
- 9:17
quickly to keep up with the times, but
- 9:20
we also want to maintain quality. When
- 9:23
building a large ambitious project like
- 9:25
Deiosis, there's a constant tension of
- 9:27
shipping fast, adding new features, and
- 9:30
then also making sure you don't break
- 9:31
things and that in a patch release,
- 9:33
you're not like breaking APIs that
- 9:36
millions of people rely on. For a
- 9:38
project that people build their
- 9:40
businesses on, there's also a high bar
- 9:42
for releases. We need to maintain uh
- 9:44
high quality of our documentation, of
- 9:47
our examples, of our tests, of our
- 9:49
benchmarks. If it if anything is out of
- 9:51
place, people figure it out pretty
- 9:53
quickly.
- 9:56
So, we really do like coding agents as a
- 10:01
a form of an excellent assistant for
- 10:03
very hard technical problems. Coding
- 10:05
agents bring a level of patience and
- 10:07
massive knowledge that is very hard to
- 10:10
muster as an individual working on a
- 10:12
very large software project.
- 10:14
Many problems in Dioxis are knowledge
- 10:16
problems. Our team can't feasibly know
- 10:19
every detail about every build system,
- 10:21
every runtime, every operating system,
- 10:24
every programming language, every API,
- 10:26
every quirk. Fortunately, this is
- 10:29
exactly where the coding agents excel.
- 10:32
They can quickly quickly sift through
- 10:34
thousands of pages of documentation,
- 10:36
read all the bespoke APIs, dig into
- 10:39
binaries, reverse engineer APIs. They
- 10:42
have so much more patience than an
- 10:44
individual developer does.
- 10:46
We were able to implement things like
- 10:48
cotlin and swift plugins for dioxysis
- 10:51
deeply integrated into our build system
- 10:54
which is a really hard feature. If you
- 10:56
know react native turbo modules these
- 10:58
things took many years of development to
- 11:00
get right by people writing them by
- 11:02
hand. We were able to ship this in like
- 11:05
two to three weeks with coding agents
- 11:07
and we probably could have gone faster.
- 11:08
I think implementation was done in like
- 11:10
the first day and we spent two weeks
- 11:12
building test cases and testing on real
- 11:13
devices.
- 11:15
And in Blitz, the thing on the right,
- 11:17
our custom web engine, web agents have
- 11:20
accelerated debugging hard CSS, styling,
- 11:23
and layout issues for us. The agents
- 11:25
know the CSS spec exceptionally well.
- 11:28
You might be writing a line of code
- 11:29
that's trying to resolve some sort of
- 11:31
painting or layout issue. And the agents
- 11:33
can instantly recall exactly how Google
- 11:35
Chrome and Safari do it. Can tell you
- 11:37
the right way of handling it for your
- 11:39
problem. and you don't have to go open
- 11:41
the the WebKit source code that's nested
- 11:44
deep somewhere in Apple's Git
- 11:47
repositories.
- 11:49
We're able to invest time in doing
- 11:51
things the right way, not the hacky way,
- 11:54
which interestingly is a turn and
- 11:57
compared to how we used to do it. We
- 11:59
would always gauge a project based on
- 12:00
its complexity and tend to take
- 12:02
shortcuts as humans to ship things
- 12:05
faster but not at a high quality bar. So
- 12:08
coding agents give us the ability to
- 12:10
maintain quality and do things the right
- 12:12
way which is very interesting.
- 12:15
Um a less sexy application of coding
- 12:18
agents for ambitious projects is
- 12:20
actually doing the extremely mundane
- 12:22
tasks. Uh our team is very small. We
- 12:24
have three core engineers working on
- 12:26
Diosis. uh any time that we spend like
- 12:29
verifying the tarball extracts into the
- 12:32
right directory structure is like time
- 12:34
wasted from us thinking about the
- 12:35
architecture and the the hard problems
- 12:37
of our software. Um Dioxis is a large
- 12:40
project and it's been a challenge to
- 12:42
maintain a high quality bar across the
- 12:44
entire codebase across every release. In
- 12:46
one release we might add an extension
- 12:48
for a new editor like zed. We might not
- 12:51
be able to test that editor every time
- 12:53
we do a patch release. And it might be
- 12:54
easy to break that. Applying agents to
- 12:57
the problem actually lets us automate
- 12:58
many of these like hard tedious tasks
- 13:00
that would have taken like countless
- 13:02
hours before.
- 13:04
And then for us like the code is the
- 13:06
product. People download the code, they
- 13:08
build on the code, users interact with
- 13:11
our APIs, they read our docs, and they
- 13:12
build on our architecture. So any
- 13:14
laziness in the the quality of the code,
- 13:17
the SDKs that we ship to users
- 13:20
translates directly into a worse
- 13:22
developer experience and people either
- 13:24
getting upset, their businesses being
- 13:26
stalled, or them turning off the
- 13:27
product. So coding agents have been
- 13:30
excellent at maintaining uh tasks like
- 13:33
verifying release checklists,
- 13:35
backporting bug fixes onto stable
- 13:36
releases, and ensuring our docs and
- 13:38
documents are of extremely high quality.
- 13:41
We still do write a lot of docu comments
- 13:44
ourselves, but it's very easy to give
- 13:46
the agent a task of making sure
- 13:47
everything is documented properly.
- 13:49
Everything has an example and everything
- 13:51
actually is documenting the thing that
- 13:53
it says in the way that it says. Um, as
- 13:56
humans, you know, you'll go edit the
- 13:58
code, but you won't edit the comment.
- 13:59
So, a lot of your comments will actually
- 14:01
be out of date over time and things get
- 14:02
very confusing. Um, and if you just look
- 14:05
at the numbers, we've shipped more patch
- 14:07
releases in our most recent DAXis
- 14:09
version than we ever had before. So,
- 14:10
we've been able to maintain weekly or
- 14:12
multiple times a week release cadence
- 14:14
for a large ambitious piece of software
- 14:17
in a way that we would be scared to do a
- 14:18
release earlier.
- 14:21
Um, one thing I'm not 100% convinced
- 14:24
yet, uh, we have found varying levels of
- 14:27
success is using AI to write tests or at
- 14:30
least blindly writing tests. Um, one
- 14:34
place we've struggled with Diosis is
- 14:36
testing. It can be very hard to test
- 14:38
foundational software, especially like
- 14:40
end to end for complex systems. It's
- 14:43
hard to test that your extension
- 14:44
installs into zed and works the way you
- 14:46
want it to do with literally opening zed
- 14:48
and like using the extension. Um, the
- 14:52
coding agents struggle here too to an
- 14:54
extent. Uh, they also are, you know,
- 14:58
have a tendency to write kind of sloppy
- 14:59
tests. you'll give it a constructor and
- 15:01
then it will go test the constructor and
- 15:02
that's not a very interesting test. Um
- 15:05
they can easily write tests for any
- 15:07
given API but much like humans they fail
- 15:09
to write the right tests. So we still
- 15:12
find ourselves enumerating test
- 15:14
conditions manually um crafting test
- 15:16
APIs our ourselves and handling test
- 15:19
runners um but it is sometimes a great
- 15:22
sounding board to come up with the test
- 15:24
ideas for a particular thing you're
- 15:26
trying to to make sure has coverage and
- 15:28
then enumerating the the edge conditions
- 15:31
um but one place that we have actually
- 15:32
really enjoyed using coding agents to do
- 15:34
testing is building test harnesses. So
- 15:37
fuzzing is a critical part of building
- 15:40
like production grade software which
- 15:43
means taking your application and
- 15:44
putting it under millions of different
- 15:46
inputs and quite often adversarial
- 15:49
inputs basically like malformed inputs
- 15:52
or uh ways of using the software that
- 15:54
users should not be using the software
- 15:56
but they they can use the software and
- 15:58
coding agents are excellent at building
- 16:00
these harnesses.
- 16:03
Um, one thing we've found that code
- 16:06
extra code architecture is still an art.
- 16:08
Um,
- 16:11
coding agents enable you to ship at an
- 16:13
exceptionally high velocity. I mentioned
- 16:14
this earlier. If the substrate on which
- 16:16
your coding agents code lands is bad,
- 16:19
their contributions will be bad as well.
- 16:22
Unlike a human engineer, coding agents
- 16:23
aren't typically afraid to voluntarily
- 16:26
go on a huge refactor of a system or
- 16:28
redesign the architecture when a feature
- 16:29
doesn't quite fit. They'll typically
- 16:31
just ship. Most of our development time
- 16:33
is actually now spent thinking about
- 16:35
software architecture about what
- 16:37
features we'll want in the future and
- 16:38
how the system will evolve. Just like
- 16:40
human engineers can write spaghetti
- 16:42
code, so can the agents, but now just
- 16:44
faster. However, I will say with fable
- 16:48
level tools, the actual code quality
- 16:50
itself is so high, provided you properly
- 16:52
communicate your intent, that proper
- 16:55
software architecture will probably take
- 16:57
the vast majority of time in the future.
- 16:59
Actual code writing, not so much.
- 17:02
Um, one thing we do for Daxis, which
- 17:05
maybe you guys still do, maybe you
- 17:07
don't, uh, is we review every PR line by
- 17:10
line. Um, we definitely use AI review to
- 17:13
spot bugs ahead of time, but we still do
- 17:16
like to read the code that we ship. We
- 17:19
receive lots and lots of PRs from
- 17:21
strangers. Actually, uh, Daxis is big
- 17:23
open source project. Uh, and not every
- 17:25
PR is made the same. Uh we find that
- 17:28
users can be quite bad at communicating
- 17:29
their intent to the models. Contributors
- 17:32
don't usually think deeply about how the
- 17:33
codebase should evolve over time. They
- 17:35
just want their bug fix or their feature
- 17:36
in. Uh and many solutions are glued in
- 17:39
place. So we're we're not quite at the
- 17:41
point where the coding agents can read
- 17:43
our minds. Uh and thus we're still
- 17:45
limited by the medium of text. And as
- 17:48
ridiculous as it sounds, prompt
- 17:49
engineering is quite real. The quality
- 17:51
of an implementation can be very much
- 17:52
dependent on the the prompt that you
- 17:54
give the model.
- 17:56
But in a sense, nothing really has
- 17:58
changed. Reading code has always been
- 18:00
more important than writing code. Um
- 18:02
maybe not in the beginning, but
- 18:04
eventually as the project evolves, uh it
- 18:06
does.
- 18:07
So my closing thoughts on on using
- 18:10
coding agents to build ambitious
- 18:12
software is that code is now cheap, but
- 18:14
quality is not. Um the the job of a
- 18:18
software engineer has never really been
- 18:20
about putting lines of code on the
- 18:21
screen. It's it's been about
- 18:24
architecting elegant solutions to
- 18:26
complex problems to thinking 10 steps
- 18:28
ahead about how a system will evolve
- 18:30
about retaining flexibility in the face
- 18:32
of changing requirements. These facts
- 18:35
have not changed and the bar for
- 18:37
software engineering is higher than
- 18:39
ever.
- 18:41
Uh if you would like to work on the
- 18:43
tools of the next generation of
- 18:44
software, Cognition, the people who have
- 18:46
acquired Dioxis are hiring. Uh the Daxis
- 18:49
team joined Cognition to be part of the
- 18:50
future and hopefully you will too. Thank
- 18:53
you.
- 18:55
[applause]
- 19:10
>> [music]