AI Engineer Code 2025
No Vibes Allowed: Solving Hard Problems in Complex Codebases
Read the talk
Context Engineering for Complex Codebases: Research, Plan, Implement, and Keep Humans Thinking
Dex Horthy explains how intentional context compaction, targeted research, explicit implementation plans, and human review can help coding agents work more effectively in existing codebases without sacrificing team understanding.
From a talk by Dex Horthy
At a glance
Ideas worth remembering
Optimize an agent’s context for correctness, completeness, manageable size, and constructive trajectory; incorrect information is more damaging than missing information or excess noise. 4:52
Use intentional compaction to preserve relevant findings, files, and implementation details while starting fresh context windows without repeating exploratory work. 2:52
Treat subagents as isolated research contexts that return concise findings, not as fictional frontend, backend, or QA teammates. 6:34
Ground research in the current code, then create explicit plans with concrete changes and tests; review those plans to maintain mental alignment and catch mistakes before they multiply. 13:57
Scale the workflow to the task: simple changes may need no formal research, while complex or cross-repository work can benefit from deeper investigation and repeated compaction. 17:40
Keep humans responsible for architectural reasoning and organizational change: coding agents can amplify sound judgment, but they cannot substitute for it. 10:00
Why existing codebases expose the limits of casual AI coding
Dex Horthy frames the central problem as a mismatch between apparent output and durable engineering progress. He describes AI-assisted software development that produces more shipped code while also creating rework, code churn, and technical debt. Straightforward greenfield projects may suit coding agents, but large, established systems present a different challenge: understanding existing architecture, navigating accumulated constraints, and making changes that do not simply create another cleanup project. 0:24
His proposed response is context engineering: improving what current models can accomplish by deliberately managing the information available in their context windows. Horthy says his three-person team spent eight difficult weeks changing how it collaborated and built software, ultimately reporting approximately two to three times more throughput. He presents that experience as a team-specific result and motivation for the workflow, not as a guaranteed outcome for every organization. 1:09
The objective extends beyond generating correct-looking code. Horthy identifies several requirements at once: agents must handle brownfield codebases and complex problems, avoid low-quality output, preserve mental alignment across the team, and meaningfully offload work to AI. Those goals can conflict when increased code generation outpaces human understanding, making collaboration and review part of the technical problem rather than an afterthought. 2:02
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Treat the context window as a limited engineering resource
A common failure pattern begins when someone asks an agent to complete a task, repeatedly corrects its mistakes, and continues until the conversation becomes unwieldy. Restarting with a fresh context and clearer steering can help, but Horthy recommends a more deliberate technique: intentional compaction. The agent compresses useful information from an existing conversation into a reviewable Markdown document, allowing a new agent or context window to begin from the important findings without repeating the entire investigation. 2:52
The value of compaction depends on selecting information that materially affects the next decision. Searching for files, understanding control flow, editing code, and collecting test or build output all consume context, as do verbose tool responses. A useful compacted artifact should preserve the task, relevant files, and precise locations while removing unnecessary noise. Horthy characterizes coding models as stateless across the information not present in the current conversation, so the quality of their next action depends on the quality of the tokens available to them. 3:59
He evaluates context along four dimensions: correctness, completeness, size, and conversational trajectory. Incorrect information is the most damaging, followed by missing information and excessive noise; repeated failed attempts can also establish an unhelpful interaction pattern. His informal dumb zone describes degraded results as a context window fills: using Claude Code as an example, he suggests that diminishing returns may begin around 40 percent utilization, while emphasizing that the threshold varies with the model and task complexity. 4:52
Investigation and accumulated conversation
A working conversation becomes a concise artifact that starts a new agent with relevant findings.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Use subagents and research-plan-implement to compress work
Horthy argues that subagents are most useful as a mechanism for controlling context, not as simulated organizational roles. Instead of assigning agents anthropomorphic titles, a parent agent can delegate a bounded investigation to a separate context window. That worker performs the expensive searching and reading, then returns a concise finding—such as the relevant file—so the parent can proceed without absorbing every exploratory step. 6:34
He builds on that pattern with frequent intentional compaction, organized into three practical phases: research, plan, and implement. Research establishes how the existing system actually works and identifies relevant files. Planning converts that understanding and the requested change into explicit steps, including filenames, code snippets, and a testing approach after each change. Implementation then executes the reviewed plan while keeping the active context focused. 7:28
The workflow is not an argument that these three labels are universally necessary or permanent. Horthy explicitly says that research-plan-implement may not remain the defining sequence; the enduring principles are compaction, context engineering, and avoiding overloaded context windows. He also rejects the idea that better results come from accumulating Markdown artifacts for their own sake: documents matter only when they accurately compress relevant truth or intent and improve execution. 17:40
Identify relevant files and system behavior
Context stays focused as investigated codebase truth becomes an actionable plan and then implementation.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Ground research in live code and review plans before they become code
Persistent repository onboarding documents can give agents useful background, but they introduce a scaling problem. As a codebase grows, comprehensive instructions either become too long or omit important details, potentially consuming much of the effective context budget before the agent begins useful work. Progressive disclosure can improve this arrangement by distributing guidance across repository levels so an agent loads root-level information and only the additional context relevant to its current area. 11:54
Even carefully organized documentation can drift away from the implementation. Horthy therefore favors on-demand compressed context: steer the agent toward the relevant area, use targeted investigations to examine vertical slices of the system, and assemble a research document based on the current code itself. He describes this as compressing truth rather than relying on documentation that teams may fail to keep synchronized with shipped changes. 13:26
Planning then becomes compression of intent. A useful plan combines the research with a product requirement, bug report, or other requested change and specifies what will happen concretely enough for a human to assess the approach. Horthy says his team increasingly includes actual proposed code snippets because vague plans do not provide sufficient confidence about the resulting changes. The tradeoff is that longer plans can improve execution reliability while becoming harder to read, so each team must find a workable balance. 13:57
This review stage protects mental alignment: a shared understanding of how the system is changing and why. Horthy says code still gets reviewed, but reviewing plans can help technical leaders track evolving architecture and catch problems earlier without relying exclusively on large diffs. He also highlights the value of showing reviewers the implementation steps, prompts, and successful build evidence so they can understand the path behind a change rather than seeing only its final code. 15:03
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Recognize the limits and scale the process to the problem
Horthy illustrates the approach with work on a 300,000-line Rust codebase for a programming language associated with Boundary ML. He describes comparing plans produced with and without research, discarding weak research outputs, and receiving a positive response to a proposed fix. In a separate extended session on BAML, he reports shipping 35,000 lines of code over seven hours, while explicitly noting that some of that volume came from generated code and updated golden files; one pull request was merged approximately a week later. These examples are presented as specific experiences, not controlled measurements or evidence that raw line counts equal productivity. 8:15
The limitations are equally important. An attempt to remove Hadoop dependencies from Parquet Java did not initially succeed, and Horthy says the collaborators ultimately returned to a whiteboard after identifying the system’s pitfalls. His conclusion is that AI cannot replace the underlying engineering judgment: it amplifies the thinking already invested in the task, and a misunderstanding in research can misdirect an entire implementation while a flawed plan can produce many flawed lines of code. 9:02
The appropriate amount of process depends on the change. Altering a button color may need only a direct instruction; a small feature may require a simple plan; work spanning multiple repositories may justify dedicated research followed by planning; especially difficult problems can demand still more context engineering. Horthy argues that learning this calibration takes repeated practice, recommends becoming proficient with one tool instead of continually optimizing across several, and identifies organizational adaptation as the longer-term challenge: teams need leadership, review practices, and shared workflows that prevent faster generation from becoming someone else’s cleanup burden. 17:40
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Read the complete timestamped transcript
- 0:12
[upbeat electronic music] Hi, everybody. [audience applauding] How y'all doing?
- 0:23
All right.
- 0:24
It's exciting. I'm Dex. Uh, as they did in the great intro, I've been hacking on agents for a while. Um, our talk, 12-Factor Agents at AI Engineer in June was one of the top talks of all time.
- 0:34
Uh, I think top eight or something, one of the best ones from, from AI Engineer in June. May or may not have said something about context engineering. Um, why am I here today?
- 0:42
What am I here to talk about? Um, I want to talk about one of my favorite talks from AI Engineer in June, and I know we all got the update from Igor yesterday, but they wouldn't let me change my slides, so this is gonna be about what Igor talked about in June.
- 0:53
Uh, basically that they surveyed 100,000 developers across all company sizes, and they found that most of the time you use AI for software engineering, you're doing a lot of rework, a lot of codebase churn.
- 1:04
Uh, and it doesn't really work well for complex tasks, brownfield codebases. Um, and you can see in the chart basically you are shipping a lot more, but a lot of it is just reworking the slop that you shipped last week.
- 1:15
So, uh, and then the other side, right, was that, uh, if you're doing greenfield little Vercel dashboard, something like this, then it's gonna work great. Uh, if you're going to go in a ten-year-old Java codebase, maybe not so much.
- 1:29
And this matched my experience. Personally, and talking to a lot of founders and great engineers, too much slop, uh, tech debt factory. It's just, it's not gonna work for our codebase.
- 1:37
Like, maybe someday when the models get better. But that's what context engineering is all about. How can we get the most out of today's models? How do we manage our context window?
- 1:46
So we talked about this in August. Um, I have to confess something. The first time I used Claude Code, I was not impressed. It was like, okay, this is a little bit better.
- 1:54
I get it. I like the UX. Um, but since then, we as a team figured something out, um, that we were actually able to get, you know, two to three X more throughput.
- 2:02
And we were shipping so much that we had no choice but to change the way we collaborated. We rewired everything about how we build software. Uh, it was a team of three.
- 2:12
It took eight weeks. It was really frickin' hard. Uh, but now that we've solved it, we're, we're never going back. This is the whole no slop thing. I think, I think we got somewhere with this.
- 2:20
Went super viral on Hacker News in September. Uh, we have thousands of folks who have gone onto GitHub and grabbed our, you know, Research, Plan, Implement prompt system. Um, so the goals here, which we kind of backed our way into, we need AI that can work well in brownfield codebases, that can solve complex problems.
- 2:38
No slop, right? No more slop. Uh, and we had to maintain mental alignment. I'll talk a little bit more about what that means in a minute. And of course, we want to spend-- with everything, we want to spend as many tokens as possible.
- 2:48
What we can offload meaningfully to the AI is really, really important. Um, super high leverage. So this is advanced context engineering for coding agents. Um, I'll start with kind of, like, framing this.
- 2:59
The most naive way to use a coding agent is to ask it for something and then tell it why it's wrong and resteer it and ask and ask and ask until you run out of context or you give up or you cry.
- 3:09
Um, we can be a little bit smarter about this. Most people discover this pretty early on in their AI, like, exploration, uh, is that it might be better if you start a conversation and you're off track that, uh, you just start a new context window.
- 3:24
You say, "Okay, we went down that path. Let's start again. Same prompt, same task, but this time we're gonna go down this path. And, like, don't go over there 'cause that doesn't work."
- 3:31
So, uh, how do you know when it's time to start over? If you see this- [laughing]
- 3:39
... it's probably time to start over, right? This is what Claude says when you tell it it's screwing up. Um, so we can be even smarter about this. We can do what I call intentional compaction.
- 3:50
Um, and this is basically whether you're on track or not, you can take, uh, your existing context window and ask the agent to compress it down into a markdown file.
- 3:59
You can review this. You can tag it. And then when the new agent starts, it gets straight to work instead of having to do all that searching and codebase understanding and getting caught up.
- 4:07
Um, what goes into compaction? Well, the question is, like, what takes up space in your context window? So, um, it's looking for files. It's understanding code flow. It's editing files.
- 4:18
It's test and build output. And if you have one of those MCPs that's dumping JSON and a bunch of UUIDs in your context window, you know, God help you.
- 4:27
Uh, so what should we compact? I'll give more on the specifics here, but this is a really good compaction. This is exactly what we're working on, the exact files and line numbers that matter to the problem that we're solving.
- 4:37
Um, why are we so obsessed with context? Because LLMs are pu-- I actually got roasted on YouTube for this one. They're not pure functions 'cause they're non-deterministic, but they are stateless.
- 4:46
And the only way to get better, better performance out of an LLM is to put better tokens in, and then you get better tokens out. And so every turn of the loop when Claude is picking the next tool or any coding agent is picking the next tool, and there could be hundreds of right next steps and hundreds
- 4:59
of wrong next steps. But the only thing that influences what comes out next is what is in the conversation so far. So we're gonna optimize this context window for correctness, completeness, size, and a little bit of trajectory.
- 5:11
And the trajectory one is interesting because a lot of people say, "Well, the-- I, I told the agent to do something, and it did something wrong, so I corrected it and I yelled at it, and then it did something wrong again, and then I yelled at it."
- 5:20
And then the LLM is looking at this conversation, says, "Okay, cool, I did something wrong and the human yelled at me, and I did something wrong and the human yelled at me."
- 5:26
So the next most likely convers-- token in this conversation is, "I better do something wrong so the human can yell at me again." So what-- mind, be mindful of your trajectory.
- 5:35
If you were going to invert this, the worst thing you can have is incorrect information, then missing information, and then just too much noise. Um, if you like equations, there's a dumb equation if you want to think about it this way.
- 5:46
Um, Geoff Huntley, uh, did a lot of research on coding agents. Uh, he put it really well. Just the more you use the context window, the worse outcomes you'll get.
- 5:55
This leads to a concept, um, and a very, very, uh, academic concept called the dumb zone. So you have your context window. You have 168,000 tokens, roughly. Some are re-reserved for output and compaction.
- 6:05
This varies by model, um, but we'll use Claude Code as an example here. Around the forty percent line is where you're going to start to see some diminishing returns depending on your task.
- 6:14
Um, if you have too many MCPs in your coding agents, you are doing all your work in the dumb zone and you're never gonna get good results. People talked about this, I'm not gonna talk about that one.
- 6:22
Your mileage may vary. Forty percent is like it depends on how complex the task is, but this is kind of a good guideline. Um, so back to compaction, or as I will call it from now on, cleverly avoiding the dumb zone.
- 6:35
Um, we can do sub-agents. Um, if you have a front-end sub-agent and a back-end sub-agent and a QA sub-agent and a data, data scientist sub-agent, please stop. Sub-agents are not for anthropomorphizing roles, they are for controlling context.
- 6:49
And so what you can do is if you wanna go find how something works in a large codebase, um, you can steer the coding agent to do this if it supports sub-agents, or you can build your own sub-agent system.
- 6:58
But basically you say, "Hey, go find how this works." And it can fork out a new context window that is gonna go do all that reading and searching and finding and reading entire files and understanding the codebase, and then just return a really, really succinct message back up to the parent agent of just like, "Hey, the file
- 7:15
you want is here." Parent agent can read that one file and get straight to work. And so this is really powerful. If you wield these correctly, you can get good responses like this, and then you can manage your context really, really well.
- 7:28
Um, what works even better than sub-agents or like a layer on top of sub-agents is a workflow I call frequent intentional compaction. We're gonna talk about research, plan, implement in a minute, but like the point is you're constantly stay-- keeping your context window small.
- 7:41
You're building your entire workflow around context management. So comes in three phases, research, plan, implement, um, and we're gonna try to stay in the smart zone the whole time.
- 7:51
So the research is all about understanding how the system works, finding the right files, staying objective. Here's a prompt you can use to do research. Here's the output of, um, a research prompt.
- 8:00
These are all open source. You can go grab them and play with them yourself. Um, y- planning, you're gonna outline the exact steps. You're gonna include file names and line snippets.
- 8:08
You're gonna be very explicit about how we're gonna test things after every change. Here's a good planning prompt. Here's one of our plans. It's got actual code snippets in it.
- 8:15
Um, and then we're gonna implement. And if you've read one of these plans, you can see very easily how the dumbest model in the world is probably not gonna screw this up.
- 8:22
Um, so we just go through and we run the plan, and we keep the context low. As a planning prompt, like I said, it's the least exciting part of the process.
- 8:29
Um, I wanted to put this into practice. So working for us, uh, I do a podcast with my buddy, uh, Vaibhav, who's the CEO of a company called Boundary ML.
- 8:36
Uh, and I said, "Hey, I'm gonna try to one-shot a fix to your three hundred thousand line Rust codebase for a programming language." [laughs]
- 8:44
Um, and the whole episode goes in, it's like an hour and a half. Uh, I'm not gonna talk through it right now. But we built a bunch of research, then we threw them out 'cause they were bad, and then we made a plan, and we made a plan without research and with research and compared all the results.
- 8:53
It's a fun time. Uh, by-- That was Monday night. By Tuesday morning, we were on the show, and the CTO had like seen the PR and like didn't realize I was doing it as a bit for a podcast, and basically was like, "Yeah, this looks good.
- 9:04
We'll get it in the next release." He, I think he was a little confused. Um, here's the, the plan. But anyways, uh, yeah, confirmed. Works in brownfield codebases and no slop.
- 9:15
But I wanted to see if we could solve complex problems. [laughs] So Vaibhav was still a little skeptical. I sat down, we sat down for like seven hours on a Saturday and we shipped thirty-five thousand lines of code to BAML.
- 9:24
One of the PRs got merged like a week later. I will say some of this is code gen. You know, you update your behavior, all the golden files update and stuff.
- 9:31
But we shipped a lot of code that day. Um, he estimates it was about one to two weeks in seven hours, and, uh, so cool. We can solve complex problems.
- 9:40
There are limits to this. I sat down with my buddy Blake. We tried to remove Hadoop dependencies from Parquet-Java. If you know what Parquet-Java is, I'm sorry, [audience laughing] uh, for whatever happened to you to get you to this point in your career.
- 9:53
Uh, it did not go well. Uh, here's the plans. Here's the research. Uh, at a certain point, we threw everything out and we actually went back to the whiteboard.
- 10:00
We had to actually-- Once we had learned where were the, where all the foot guns were, we r- we went back to, okay, how is this actually gonna fit together?
- 10:07
Um, and this brings me to a really interesting point that, uh, Jake's gonna talk about later. Uh, do not outsource the thinking. AI cannot replace thinking, it can only amplify the thinking you have done or the lack of thinking you have done.
- 10:19
So people ask, "So, so Dex, this is spec-driven development, right?" No. Spec-driven development is broken. Not the idea, but the phrase. Um,
- 10:32
it's not well defined. This is Birgitta from ThoughtWorks. Um, and a lot of people just say spec, and they mean a more detailed prompt. Does anyone remember this picture?
- 10:41
Does anyone know what this is from? All right, that's a deep cut. Uh, there will never be a year of agents because of semantic diffusion. Martin Fowler said this in two thousand and six.
- 10:49
We come up with a good term with a good definition, and then everybody gets excited and everybody starts meaning it to mean a hundred things to a hundred different people, and it becomes useless.
- 10:59
We had an agent is a person, an agent is a microservice, an agent is a chatbot, an agent is a workflow. And thank you, Simon. We're back to the beginning.
- 11:07
An agent is just tools in a loop. Um, this is happening to spec-driven dev. I used to have Sean's, uh, slide in the beginning of this talk, but it caused a bunch of people to focus on the wrong things.
- 11:17
His thing of like, forget the code, it's like assembly now, and you just focus on the markdown. Very cool idea, but people say spec-driven dev is writing a better prompt, a product requirements document.
- 11:26
Sometimes it's using like verifiable feedback loops and backpressure. Maybe it is treating the code like assembly like Sean taught us. Um, but a lot of people is just using a bunch of markdown files while you're coding.
- 11:37
Or my favorite, I just stumbled upon this last week, uh, a spec is, uh, documentation for an open source library. So it's gone. It's-- As spec-driven dev is overhyped, it's useless now.
- 11:48
It's semantically diffused. Um, so I wanna t-talk about like four things that actually work today, the tactical and practical steps that we found working internally and with a bunch of users.
- 11:58
Um, we do the research, we figure out how the system works. Um, you remember Memento? This is the best, the best movie on context engineering, as Peter says it. [audience laughing]
- 12:07
This guy wakes up, he d-- has no memory. He has to like read his own [REDACTED:physical_attribute] to figure out who he is and what he's up to. [audience laughing] If you don't onboard your agents, they will make stuff up.
- 12:17
And so if this is your team, this is very simplified for most of you. Most of you have much bigger orgs than this. But let's say you wanna do some work over here.
- 12:23
Um, one thing you could do is you could put onboarding into every repo. You put a bunch of context, here's the repo, here's how it works. This is a compression of all the context in the codebase that the agent can see ahead of time before actually getting to work.
- 12:36
This is challenging because sometimes it gets too long. As your codebase gets really big, you either have to make this longer or you have to leave information out. And so as you, uh, are reading through this, you're gonna read the context of this big five million line monorepo, and you're gonna use all the smart zone just to
- 12:53
learn how it works, and you're not gonna be able to do any good tool calling in the dumb zone. So that's, uh, you can-- [laughs]
- 13:01
you can shard this down the stack. You can do the-- Just talking about progressive disclosure, you could split this up, right? You could just put a file in the root of every repo, and then, like, at every level, you have, like, additional context based on if you're working here, this is what you need to know.
- 13:15
Uh, we don't document the files themselves 'cause they're the source of truth. But then as your agent is working, you know, you pull in the root context, and then you pull in the sub-context.
- 13:22
And we won't talk about any specific... Like, you could use CLAUDE.md for this, you can use hooks for this, whatever it is. Um, but then you still have plenty of room in the smart zone 'cause you're only pulling in what you need to know.
- 13:31
Um, the problem with this is that it gets out of date. And so every time you ship a new feature, you need to kinda, like, cache, invalidate, and rebuild large parts of this internal documentation.
- 13:42
And you could use a lot of AI and make it part of your process to update this. Um, but I wanna ask a question. Between the actual code, the function names, the comments, and the documentation, does anyone wanna guess what is on the Y-axis of this chart?
- 13:57
Slop.
- 13:57
Slop. It's actually the amount of lies you can find [audience laughing] in any one part of your codebase. Um, so you could make it part of your process to update this, but you probably shouldn't 'cause you probably won't.
- 14:08
What we prefer is on-demand compressed context. So if I'm building a feature that me- relates to SCM providers in Jira and Linear, um, I would just give it a little bit of steering.
- 14:17
I would say, "Hey, we're going over in, like, this, like, part of the codebase over here." Um, and a good research, uh, prompt or, or slash command might take you-- or skill even, uh, launch a bunch of sub-agents to take these vertical slices through the codebase, and then build up a research document that is just a snapshot
- 14:34
of the actually true, based on the code itself, parts of the codebase that matter. We are compressing truth. Um, planning is leverage. Planning is about compression of intent. Um, and in plan, we're gonna outline the exact steps.
- 14:48
We'll take our research and our PRD or our bug ticket or our whatever it is, and we create a plan, and we create a plan file. So we're compacting again.
- 14:55
And I wanna pause to talk about mental alignment. Um, does anyone know what code review is for? [chattering] [laughs]
- 15:04
Mental alignment. Mental alignment. It's, it is about finding, making sure things are correct and stuff, but the most important thing is how do we keep everybody on the team on the same page about how the codebase is changing and why.
- 15:14
And I can read a thousand lines of Golang every week. Uh, sorry, I can't read a thousand. It's hard. I can do it. I don't want to. Um, and as our team grows, I-- all the code gets reviewed.
- 15:23
We don't not read the code. But I, as, you know, a technical leader in the, in the, on the team, I can read the plans, and I can keep up to date, and I can...
- 15:30
That's enough. I can catch some problems early, and I maintain understanding of how the system is evolving. Um, Mitchell had this really good post about how he's been putting his Amp threads on his pull requests so that you can see not just, "Hey, here's a wall of green text in GitHub," but, "Here's the exact steps.
- 15:44
Here's the prompts. And hey, I ran the build at the end and it passed." This takes the reviewer on a journey in a way that a GitHub PR just can't.
- 15:51
And as you're shipping more and more in two to three times as much code, it's really on you to find ways to keep your team on the same page and show them, "Here's the steps I did, and here's how we tested it manually."
- 16:02
Um, your goal is leverage, so you want high confidence that the model will actually do the right thing. I can't read this plan and know what actually is gonna happen and what code changes are gonna happen.
- 16:11
So we've, over time, iterated towards our plans include actual code snippets of what's gonna change. So your goal is leverage. You want compression of intent, and you want reliable execution.
- 16:21
Um, and so I don't know. I have a physics background. We like to draw lines through the center of peaks and curves. Uh, [laughs] as your plans get longer, reliability goes up, readability goes down.
- 16:31
There's a sweet spot for you and your team and your codebase. You should try to find it. Because when we review the research and the plans, if they're good, then we can get mental alignment.
- 16:40
Um, don't outsource the thinking. I've said this before. This is not magic. There is no perfect prompt. You still-- It will not work if you do not read the plan.
- 16:50
So we built our entire process around you, the builder, are in back and forth with the agent, reading the plans as they're created. And then if you need peer review, you can send it to someone and say, "Hey, does this plan look right?
- 17:00
Is this the right approach? Is this the right order to look at these things?" Um, Jake, again, wrote a really good blog post about, like, the thing that makes research plan implement valuable is you, the human, in the loop, making sure it's correct.
- 17:11
So if you take one thing away from this talk, it should be that a bad line of code is a bad line of code. And a bad part of a plan is-- could be a hundred bad lines of code.
- 17:22
And a bad line of research, like a misunderstanding of how the system works and where things are, your whole thing's gonna be hosed. You're gonna be telling-- sending the model off in the wrong direction.
- 17:31
And so when we're working internally and with users, we're constantly trying to move human effort and focus to the highest leverage parts of this pipeline. Um, don't outsource the thinking.
- 17:41
Watch out for tools that just spew out a bunch of Markdown files just to make you feel good. I'm not gonna name names here. Uh, sometimes this is overkill, and the way I like to think about this is like, yeah, you don't always need a full research plan implement.
- 17:54
Sometimes you need more, sometimes you need less. If you're changing the color of a button, just talk to the agent and tell it what to do. Um, if you're doing, like, a simple plan and it's a small feature, if you're doing medium features across multiple repos, then do one research, then build a plan.
- 18:09
Basically, the hardest problem you can solve, the ceiling goes up, the more of this context engineering compaction you're willing to do. Um, and so if you're in the top right corner, you're probably gonna have to do more.
- 18:19
A lot of people ask me, "How do I know how much context engineering to use?" It takes reps. You will get it wrong. You have to get it wrong over and over and over again.
- 18:27
Sometimes you'll go too big. Sometimes you'll go too small. Pick one tool and get some reps. I recommend against min-maxing across Claude and Codex and all these different tools.
- 18:36
Um- So I'm not a big acronym guy. Uh, we said spec-driven dev was broken. Uh, research, plan, and implement I don't think will be the steps. The important part is compaction and context engineering and staying in the smart zone.
- 18:48
But people are calling this RPI, and there's nothing I can do about it. So, uh, just be wary. There is no perfect prompt. There is no silver bullet. Um, if you really want a hypey word, you can call this harnage, harness engineering, which is part of context engineering, and it's how you integrate with the integration points on
- 19:04
Codex, Claude, Cursor, whatever, how you customize your codebase. Um, so what's next? I think the coding agent stuff is actually gonna be commoditized. People are gonna learn how to do this and get better at it, and the hard part is gonna be how do you adapt your team and your workflow in the SDLC to work in a
- 19:21
world where 99% of your code is shipped by AI? Uh, and if you can't figure this out, you're hosed because there's kind of a rift growing where, like, staff engineers don't adopt AI because it doesn't make them that much faster, and then junior mid-level engineers use a lot 'cause it fills in skill gaps, and then it also
- 19:35
produces some slop, and then the senior engineers hate it more and more every week because they're cleaning up slop that was shipped by Cursor the week before. Uh, this is not AI's fault.
- 19:44
This is not the mid-level engineer's fault. Like, if ... Cultural change is really hard, and it needs to come from the top if it's gonna work. So if you're a technical leader at your company, pick one tool and get some reps.
- 19:54
If you wanna help, we are hiring. We're building an agentic IDE to help teams of all sizes speed run the journey to 99%, uh, AI-generated code. Uh, if you ...
- 20:04
We'd love to, we'd love to talk if you wanna work with us. Uh, go, go hit our website. Send us an email. Come find me in the hallway. Uh, thank you all so much for your energy. [audience applauds] [upbeat music]