AI Engineer World's Fair 2026
From AI-Assisted to AI-Native: Building a Frontier Development Team — Clare Liguori, AWS
Read the talk
From AI-Assisted to AI-Native: Building a Frontier Development Team
Clare Liguori explains how Amazon teams changed their daily work to support hours of independent agent execution—and why faster implementation shifts the bottleneck toward review, organizational learning and decisions.
From a talk by Clare Liguori
At a glance
Ideas worth remembering
The 50-team pilot associated stronger deployment gains with deliberate workflow changes. Commit counts, revised delivery estimates and deployment velocity describe different outcomes.
Agent context needs maintenance in both directions: preserve missing team knowledge and remove obsolete model workarounds.
Independent execution depends on clear intent and actionable checks. Local deterministic mocks shorten the correction loop, allowing agents to repair more mistakes before returning.
Allow time for codebase investment and bounded organizational learning. Parallel agents also increase coordination and review demands, particularly for less experienced reviewers.
When implementation shrinks, product decisions and launch approvals can dominate delivery time. Decisions that are easy to reverse are a specific opportunity to move faster.
What changes when engineers leave the continuous loop
Inline completion helps write the next line or function. Chat answers questions about a codebase. Vibe coding turns implementation into a running conversation. Yet Clare Liguori, Senior Principal Engineer at AWS working mostly on Kiro, had personally felt only about a 10–20% productivity improvement through those earlier phases. That is her anecdotal starting point. Amazon’s internal pilots were beginning to report much larger gains: a median of 4.5×, sometimes exceeding 10×. The question is what changed in the work itself.
Amazon calls the emerging practice frontier development and defines it through observable behavior:
- Hands-off coding: Engineers write perhaps 1–2% of the code they produce; agents write the rest.
- Infrequent intervention: Engineers aim to give agents enough direction to run for hours without interruption.
- Concurrent work: Multiple agents work through a backlog in parallel, reducing idle time.
These behaviors change where human attention goes. An engineer must prepare work that can continue without a new instruction after every generated change.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An exceptional team, then a carefully prepared sprint
The first striking example was Bedrock Mantle. Bedrock’s model-hosting service needed a new inference data plane, with an original estimate of 30 people over 18 months. The anticipated work included building the service and migrating customers and models. Instead, six engineers built the new data plane in 76 days with Kiro. Liguori describes this as the pathfinder result, reporting up to 20× improvement from an assessment that looked at commits.
The six included two distinguished engineers and some of Amazon’s strongest experts in distributed systems and LLM architecture. Their result established a possibility, but its staffing made reproduction an immediate concern. Other teams could adopt the tool; they could not simply acquire the same concentration of expertise.
Prime Video tested a different group of six engineers in a 10-day sprint. Their progress changed the project delivery estimate from 90 weeks to 24 weeks, and the team compared sprint commits with its earlier commit history. The 24-week figure was a revised forecast, rather than a completed delivery. The conditions were also unusually favorable: no on-call duties, limited meetings and few distractions.
The preparation explains part of what made that sprint executable. A senior engineer had spent the preceding three weeks creating small, well-scoped tasks with detailed requirements. Six engineers could then spend the protected sprint churning through already-defined work. The fast execution rested on earlier human effort to remove ambiguity. That left the next question: could ordinary teams sustain this pattern while maintaining existing systems and handling their normal obligations?
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Fifty ordinary teams expose the difference in working practices
Amazon Stores observed 50 teams for the better part of a year. These teams had normal mixes of early-career, mid-career and senior engineers. They worked in existing codebases, unlike Mantle’s greenfield build. The measure also moved closer to customer delivery: deployment velocity to production, or how quickly teams could get changes out to customers.
Half the teams achieved less than a 3× increase. The stronger-performing half saw a median of 4.5×, with some exceeding 10×. Ninety percent used Kiro alongside other internal tools. Liguori attributes the split to deliberate changes in working practices: the teams pulling ahead reorganized how they worked, while the others added assistants to their existing routines. This observational account does not isolate the causal effect of individual habits or establish gains every team should expect. Its deployment measure also differs from the earlier commit assessments and revised project estimate.
Interviews with the pilot teams, Bedrock Mantle and Prime Video produced five recurring habits. The emphasis on habits matters: a protected sprint can demonstrate potential, while everyday development requires a repeatable way to define, delegate and check work. Building that routine takes time.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Write down team knowledge, then make the codebase easier to work in
The first habit is investing in agent context. Much of a team’s operating knowledge lives in people’s heads and moves through Slack conversations, onboarding, mentoring, code reviews, stand-ups and sprint planning. Agents need access to that knowledge too. Whenever an agent makes a mistake or chooses an unwanted approach, the practical question is what it needed to know that was missing from the skills or steering files. Those files turn a correction into guidance available on future tasks.
Context maintenance includes subtraction. Earlier models required numerous prohibitions to work around recurring quirks; improved model behavior made some of those instructions unnecessary. Keeping every old workaround adds context without necessarily improving the next run. The habit therefore has two directions: add missing team knowledge, and prune instructions whose original purpose has disappeared.
The second habit is accepting an initial slowdown. Almost every interviewed team reported lower productivity while intentionally adopting its new workflow. Existing codebases needed engineering work before agents could succeed independently. That investment addressed several distinct obstacles:
- Opaque failures: Better tool error messages helped models understand what went wrong and attempt a correction.
- Missing operations: New tools and MCP servers gave agents ways to perform work they otherwise could not complete.
- Difficult navigation: Restructuring code helped agents find and understand the parts they needed to change.
These are concrete preparation costs, rather than benefits that arrive automatically with access to a coding assistant.
Some teams went as far as changing programming languages. Liguori had seen teams struggle with Python and JavaScript when their workflow lacked compiler feedback that would expose mistakes before the agent returned its work. Moves to TypeScript and use of Rust offered more useful compiler diagnostics. The mechanism is actionable feedback: the agent gets a signal it can use to repair its own output. A language migration is optional, and these examples do not establish that every codebase needs one.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Give agents a completion standard and resolve intent before code
The third habit addresses the engineer’s own waiting time. A conversation that repeatedly generates code, waits 30 seconds to a minute, and returns for human review keeps the engineer occupied throughout execution. Each response requires reading and redirection. That attention cannot also go to other work, and running several such conversations in parallel becomes difficult.
An independent assignment gives the agent both the work and the means to validate it. The agent should keep correcting its changes until they meet a stated quality bar: the code runs, compiles, passes tests, is testable and has high coverage. Putting these expectations in steering files makes them recurring instructions. The human no longer has to supply the same completion standard in every prompt.
The fourth habit is making intent explicit. A high-level prompt can produce extensive changes before anyone notices that the requirements or technical design were misunderstood. Correcting code scattered across a repository is then an expensive way to discover what the feature should do. For ambiguous, complex work, Amazon engineers refine a specification before implementation.
Kiro can generate the initial specification, so explicit intent does not require manually writing the whole document. The engineer and model can work through disagreements in one place before those disagreements become implementation changes. Conversation remains useful here: its purpose is to settle the task, giving later autonomous execution a clearer destination.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Move testing earlier so the agent can correct itself locally
The fifth habit, shifting testing left, supplies the feedback that makes a long independent run useful. Agents will make mistakes. Linters and unit, integration, performance and security tests give them signals they can act on before returning to an engineer. Familiar engineering hygiene gains another use: the same checks can support repeated autonomous correction throughout implementation.
A concrete change is the move from integration tests that depend on live services to mock services with deterministic responses running entirely locally. In the earlier arrangement, testing required starting other services and connecting to cloud systems. With local mocks, the agent can perform the test loop on a laptop. A failed check supplies feedback; the agent changes the code and checks again. Reducing the wait for each result permits more correction cycles in the same period. The benefit concerns speed and repeatability in this local loop; a mock does not establish every property of the live integration.
Where does the engineer stop being required at every turn? The diagram follows that local correction loop: the assignment and quality bar enter at the start, while failed checks return to implementation rather than immediately returning to the human.
The loop makes the relationship between testing and autonomy visible. Clear instructions alone cannot tell the agent whether its implementation works. Local checks supply that information quickly enough for it to continue repairing the change. Meeting the completion standard becomes the point at which it returns to the engineer.
Specify what to do and how to self-validate.
Deterministic local mocks reduce dependency setup and cloud connections in the test loop. Failed checks guide another correction; passing the stated quality bar ends the independent assignment.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Parallel agents still create human work
These habits do not produce effortless productivity. Engineers may stay up late trying to perfect a prompt so an agent can run overnight and leave a finished change in the morning. Multiple concurrent agents also mean switching between terminal tabs and tasks. Independence at the agent level can increase pressure and cognitive load for the person coordinating the work.
Review is a particular difficulty. Senior engineers have often spent years assessing other people’s code. Early-career engineers may not yet have developed that skill, and reviewing generated output can feel harder than writing the code themselves. Reducing manual implementation therefore changes the human task; it does not necessarily make that task easier.
Organizations must also allow the initial investment. Liguori admits sharing the leadership temptation to ask why a team with strong models and tools is not already shipping faster. She describes taking two months to invest in a codebase, discover team practices and establish new habits. Constant pressure to ship features every month can consume the time needed to make agents effective.
Expansion needs the same patience. Amazon learned through an exceptional pathfinder, a focused sprint and the 50-team pilot before attempting broader adoption. At the time of the talk, the 2026 challenge was extending the approach to the next 2,000 teams—a scaling objective, rather than an achieved result. An immediate organization-wide mandate would send teams into the new workflow before they had learned which practices and context their own environment required.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Faster implementation makes decision delays dominate
The ending moves beyond coding practices to the time a customer waits for a product. Liguori describes work that previously took 9–12 months to implement and can now take 1–2 months. If the organization still spends two months deciding whether to build it and another two months approving launch, those decisions become the longest part of delivery. Their duration did not have to increase for them to become the bottleneck; implementation became much shorter.
Frontier teams can consequently spend more time making decisions than writing code. The practical recommendation is to make decisions faster, especially when they are easy to reverse. Reversibility identifies choices that can be revisited, giving organizations a specific place to reduce deliberation as the implementation loop accelerates.
Frontier development requires a deliberate change in daily work across engineers and their organizations. The closing invitation is to examine each interaction with an AI tool: what knowledge, clearer intent or feedback would let the work continue without another intervention? The same question eventually reaches product decisions and launch processes. Freed attention still needs somewhere useful to go.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- Spec-Driven Development: Agentic Coding at FAANG Scale and Quality — Al Harris, Amazon Kiro
Continue with the specification approach that gives autonomous implementation a clearer destination.
- Making Codebases "Agent-Ready"
Explore the codebase preparation behind the initial slowdown and later independent execution.
- Let’s Talk About FOMAT – Fear of Missing Agent Time
Follow the human pressure created by trying to keep agents working continuously.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> My name is Claire La Gory and I'm a
- 0:15
senior principal engineer at AWS. I
- 0:18
mostly work on Kuro, our agent encoding
- 0:20
assistant, but today I want to talk
- 0:22
about some of the practices we've been
- 0:24
seeing inside of Amazon and Amazon teams
- 0:27
where we've been seeing really exciting
- 0:29
results of productivity increases that
- 0:33
are step function improvements since
- 0:34
what what we've been seeing with AI so
- 0:36
far.
- 0:38
So, I've been working on agentic AI for
- 0:42
over 3 years now and I've kind of seen
- 0:44
the evolution that's happened in our
- 0:46
industry when it comes to coding
- 0:48
assistance with AI. First, we had this
- 0:51
inline code completion helping us to
- 0:54
write the next line, maybe the next
- 0:56
function. We moved on to chat, asking
- 0:59
questions about our code. Everybody
- 1:01
started doing vibe coding sometime last
- 1:03
year, but now we're starting to see kind
- 1:06
of an early adopter phase of what we've
- 1:08
been calling frontier development.
- 1:11
And completely anecdotally, based on my
- 1:13
own experience, I've really only felt
- 1:16
maybe 10 to 20% more productive with all
- 1:19
of these phases that have come before.
- 1:22
But now inside of Amazon, we've been
- 1:24
running pilots with different teams
- 1:26
across the company and we've been seeing
- 1:29
a median of 4.5x productivity
- 1:31
improvement and sometimes more than 10x.
- 1:34
So, something has really changed here
- 1:36
now that we're seeing these step
- 1:38
function improvements in productivity.
- 1:40
And I like to
- 1:43
define what we've been calling frontier
- 1:45
developers inside of Amazon by three
- 1:48
behaviors that I've been seeing. One is
- 1:51
hands-off coding. Frontier developers
- 1:53
write maybe 1 to 2% of the code that
- 1:56
they produce. The rest is agents.
- 2:00
The second is that they interact with
- 2:01
their agents infrequently. They'll aim
- 2:04
to get their coding assistant to run for
- 2:06
up to hours at a time without their
- 2:09
intervention.
- 2:10
And third is that they minimize idle
- 2:12
time.
- 2:13
These frontier developers tend to run
- 2:15
multiple agents in parallel churning
- 2:18
through a backlog of tasks.
- 2:22
The first time that I saw a frontier
- 2:24
developer team was the Bedrock Mantle
- 2:27
team. Bedrock is our model hosting
- 2:30
service.
- 2:31
Hosts LLMs like Claude and GPT. And
- 2:36
sometime last year we knew or I say we
- 2:40
but the Bedrock team
- 2:42
knew that they were going to need to
- 2:43
build a new inference data plane. But
- 2:46
they had estimated it at 30 people over
- 2:50
18 months. This is a big big service and
- 2:53
it was going to take time to build the
- 2:55
new one, migrate customers over, migrate
- 2:58
models over. They decided to take a step
- 3:00
back. They took six people and they
- 3:03
built [snorts] it in 76 days with Kiro.
- 3:06
So this was a huge achievement. This was
- 3:08
the first time we've we'd seen anything
- 3:10
of the kind inside of Amazon. So this
- 3:13
was truly the pathfinder team that
- 3:15
proved that it was possible to get up to
- 3:18
20X improvement. Now they looked at
- 3:21
commits and I'll talk about a couple of
- 3:23
other ways that we are uh measuring
- 3:25
productivity improvements.
- 3:27
But there was one problem with this
- 3:29
story which was that yes, it was built
- 3:32
with six people. It was built with some
- 3:35
of the top engineers literally in the
- 3:37
company including two distinguished
- 3:39
engineers. So this was not just any team
- 3:42
of six people. These were experts in
- 3:45
distributed systems, experts at LLMs and
- 3:48
their architecture.
- 3:51
So this the story was amazing and it
- 3:53
kind of spread like wildfire across
- 3:55
Amazon, but it was also very
- 3:57
unachievable for a lot of teams. There
- 3:59
were a lot of questions about can this
- 4:02
actually be reproduced on another team?
- 4:05
So, another experiment that I want to
- 4:07
talk about is an experimental sprint
- 4:10
that was done in the Prime Video
- 4:11
organization.
- 4:13
They took a 10-day sprint and they did
- 4:16
an experiment where they put, again, six
- 4:18
engineers in a room and they let them go
- 4:21
wild with Kiro.
- 4:23
Uh they brought down the project
- 4:26
delivery time estimate from what was
- 4:29
going to be 90 weeks down to 24 based on
- 4:32
all of the progress they had made in
- 4:34
this 10-day sprint. And they they looked
- 4:37
at their commit history and they looked
- 4:40
at what did they used to do prior to
- 4:42
this 10-day sprint and how many commits
- 4:44
did they produce just in this 10 days.
- 4:48
And so, this sprint really proved that
- 4:50
we can achieve, again, at least
- 4:53
something close to what the Bedrock
- 4:54
Mantle team had uh had achieved with a
- 4:58
different set of engineers.
- 5:00
But again, there was a challenge with
- 5:02
this story, which was it was six
- 5:05
engineers in a room, but they had no
- 5:07
on-call duties, limited meetings, very
- 5:10
few distractions, which we all know are
- 5:13
regular in the lives of an engineer.
- 5:16
And the senior engineer on the team had
- 5:19
spent the previous 3 weeks creating very
- 5:22
detailed, small, well-scoped tasks with
- 5:25
detailed requirements for these
- 5:28
six [clears throat] engineers to just go
- 5:29
churn on for those 2 weeks.
- 5:32
So, this was again not necessarily real
- 5:34
life. This was a structured sprint, uh a
- 5:37
a point in time that they were able to
- 5:39
achieve this, but again, the question is
- 5:42
is this achievable on real teams on
- 5:45
day-to-day
- 5:47
work?
- 5:48
So, Amazon stores which encompasses
- 5:51
amazon.com, all of our retail websites,
- 5:54
as well as our physical stores,
- 5:56
did a more structured pilot. They
- 5:59
watched 50 teams that were totally
- 6:01
normal normal distribution of um early
- 6:06
career folks, mid-career, senior
- 6:08
engineers, and that worked on existing
- 6:11
systems. Nothing green field like the
- 6:13
mantle team got to build from the ground
- 6:15
up, but existing systems with existing
- 6:17
code bases.
- 6:19
And they they watched them for the
- 6:21
better part of last year, and they found
- 6:24
something super interesting.
- 6:26
They found that there was a big
- 6:28
difference in the productivity gains
- 6:30
that they saw between half of the teams
- 6:32
and the other half.
- 6:34
And in this case, they used a
- 6:36
productivity metric of deployment
- 6:38
velocity to production. So, not just
- 6:40
commits, how many commits are they
- 6:43
producing, but how quickly are we
- 6:45
getting changes out to customers? How
- 6:47
how quickly are we able to ship things?
- 6:50
And they saw that for half of the teams,
- 6:52
they achieved less than 3x increase.
- 6:55
And what they found that was the
- 6:57
difference between seeing less than 3x
- 6:59
productivity increase, these teams that
- 7:01
saw a median of 4.5x, and and in some
- 7:04
cases more than 10,
- 7:06
was how they used the tools. 90% of
- 7:09
these teams used Kiro, among other
- 7:11
internal tools that we have, and what
- 7:14
they found was it wasn't about the
- 7:16
tools, it was about the way that they
- 7:18
worked.
- 7:19
The teams that achieved step function
- 7:21
improvements
- 7:23
intentionally changed the way that they
- 7:25
worked, and the other simply kind of
- 7:27
sprinkled Kiro and some of the other
- 7:29
tools that we have on top of their
- 7:31
existing way of working. And for me at
- 7:34
least, this was the big aha moment. That
- 7:37
why I hadn't been feeling potentially
- 7:40
the massive gains that productive that
- 7:43
in in productivity that AI has promised.
- 7:46
It's about changing the way that we
- 7:47
work.
- 7:49
So, across this pilot, they went and
- 7:51
interviewed uh the teams that were
- 7:53
involved in the pilot as well as some of
- 7:55
these other teams on the Bedrock mantel
- 7:56
team, on uh Prime Video, and they found
- 8:00
five habits. And and I use the word
- 8:03
habits very specifically because again,
- 8:05
it's not about that one sprint. It's
- 8:08
about doing this day-to-day. And it And
- 8:10
what they found when they interviewed
- 8:12
with these teams was that it really was
- 8:14
habits that they had to build
- 8:16
day-to-day. When we change our way of
- 8:18
working, it's it's hard to build these
- 8:21
habits. It takes time to build these
- 8:22
habits.
- 8:24
So, let's go through each of these one
- 8:25
by one.
- 8:26
Habit number one is investing in agent
- 8:28
context. We have a lot of stuff in our
- 8:32
head. We tend to transfer all of that
- 8:34
stuff in our head to other people
- 8:35
through Slack conversations, through
- 8:38
onboarding, mentors, things like that,
- 8:40
through code reviews, through
- 8:43
stand-ups and sprint planning, and they
- 8:45
had to write all of that down. And the
- 8:48
habit that they built was every time the
- 8:51
agent makes a mistake or does something
- 8:53
not the way that you would have done it,
- 8:55
what am I missing in my skills files?
- 8:57
What am I missing in my steering files
- 9:00
that the agent needed?
- 9:02
But then, as we know, across last year,
- 9:04
we saw leaps and bounds in models'
- 9:07
abilities and their behaviors.
- 9:09
Uh the Sonnet 3.7 in the middle of last
- 9:12
year had a lot of quirks that we had to
- 9:15
put a lot of do nots in our uh in our
- 9:17
steering files, and now we don't have to
- 9:19
do that as much with Opus 4.5 as of last
- 9:22
November, and then we've had 6 months
- 9:25
more than 6 months of improvement since
- 9:27
then
- 9:28
uh with all of the new versions of
- 9:29
models that have come out since then.
- 9:31
And so, the question, the new habit,
- 9:33
again, is do I still need this in my
- 9:36
steering files or is this just bloating
- 9:37
context?
- 9:39
The second one is slowing down to speed
- 9:41
up. In almost every team that was
- 9:44
interviewed, they reported that their
- 9:46
productivity actually went down as they
- 9:49
intentionally adopted a new way of
- 9:51
working.
- 9:52
That's counterintuitive, right? You have
- 9:55
to do intentional engineering work
- 9:57
before you're going to see that hockey
- 9:59
stick curve in productivity improvement.
- 10:02
Because we have to do real work in our
- 10:04
code base first for agents to be
- 10:06
successful there, especially in
- 10:08
brownfield existing code bases. So they
- 10:10
had to build that agent context up. They
- 10:13
had to improve existing tools error
- 10:15
messages so that the model knew what was
- 10:17
going on when it failed. They built new
- 10:20
tools, new MCP servers for helping that
- 10:23
model to actually get done what it
- 10:25
needed to get done. A lot of teams ended
- 10:27
up restructuring their code base so that
- 10:29
agents could actually navigate it more
- 10:31
easily. And I've even seen drastic
- 10:34
changes like changing the programming
- 10:36
language of the code base.
- 10:38
Um often I've seen teams struggle with
- 10:40
Python, with JavaScript because they're
- 10:43
untyped languages. It's hard to test.
- 10:46
There's no compiler errors. So the model
- 10:48
kind of guesses and give it gives it
- 10:50
back to you. And so I've seen teams
- 10:53
moving to TypeScript. Um Rust has become
- 10:56
very popular inside of Amazon. The
- 10:57
compiler gives great error messages.
- 11:00
Um you don't have to do that, but I've
- 11:02
seen a lot of teams making those
- 11:04
intentional changes for the productivity
- 11:06
gains that they're able to see.
- 11:09
The third one is feeding agents, not
- 11:12
babysitting agents. And for me this was
- 11:14
one of those aha moments of why we're
- 11:17
seeing this step function improvement in
- 11:19
productivity.
- 11:21
If you are vibe coding, if you are
- 11:23
having a back-and-forth conversation
- 11:25
with your agent all day long, of course
- 11:28
you're not going to see four to five x
- 11:31
productivity improvements because you
- 11:33
are in the loop the entire time. You're
- 11:35
probably sitting there for 30 seconds to
- 11:37
a minute waiting for it to generate code
- 11:40
and come back to you with with the code
- 11:42
to review.
- 11:44
If you're sitting there waiting for it,
- 11:46
then you can't go off and do other
- 11:48
stuff. It's really difficult to run
- 11:50
agents in parallel. It's very difficult
- 11:53
to get to to clone yourself into
- 11:55
multiple agents. And so if your
- 11:58
conversations look a bit like this on
- 12:00
the left, then you're babysitting that
- 12:02
agent. As opposed to the right side
- 12:05
where you're feeding it what it needs to
- 12:07
do and how it can self-validate. And
- 12:09
that's really the key so that agents can
- 12:11
self-correct and only come back to you
- 12:14
when it meets a certain quality bar,
- 12:16
when it when it actually runs and
- 12:18
compiles and passes tests, when it's
- 12:20
testable, when it it actually has high
- 12:23
coverage. And of course the next level
- 12:25
is put all of this content into your
- 12:27
steering file so it does it every time
- 12:29
without you having to prompt it.
- 12:33
The fourth habit is to make intent
- 12:36
explicit. At Amazon we practice a lot of
- 12:39
behavior-driven development. We've built
- 12:41
that into the Q product and so it's very
- 12:44
natural for Amazon engineers to adopt it
- 12:46
in Q. Um what what I've typically seen
- 12:50
with live coding as opposed to frontier
- 12:52
engineering is giving a very high-level
- 12:56
prompt, letting the agent generate a ton
- 12:59
of code, and then having a
- 13:01
back-and-forth conversation saying, "Oh,
- 13:04
that's not really what I meant. That you
- 13:07
haven't you haven't exactly gotten the
- 13:09
the requirements right. No, I didn't
- 13:11
actually want to build it that way.
- 13:12
Here's a technical design." And it is
- 13:15
less I find less productive to iterate
- 13:18
with the agent on code when the intent
- 13:21
itself was incorrect. So often will have
- 13:25
will see Amazon engineers go through
- 13:28
this process for for ambiguous complex
- 13:31
features of writing the specification.
- 13:34
And in Kiro, of course, you don't have
- 13:35
to write this whole specification. You
- 13:37
can have the model generate it, but it's
- 13:40
a lot easier to to iterate with the
- 13:43
model in kind of a back and forth
- 13:44
conversation about a document than it is
- 13:48
about code that's code changes that are
- 13:50
spread across a code base.
- 13:53
The fifth one is shift testing left. One
- 13:58
of the keys here is to give the agent
- 14:00
that fast feedback loop.
- 14:02
Because that's what lets it go off for
- 14:04
hours at a time and self-correct. The
- 14:06
agent is going to make mistakes and
- 14:08
that's fine. But if you give it the
- 14:11
right signals, it can self-correct and
- 14:13
it can spend a while doing that.
- 14:16
So, I've seen teams adding linters,
- 14:19
adding unit tests, integration tests,
- 14:21
performance tests, security tests. These
- 14:23
are all things we all know we should
- 14:24
have been doing all along. This is good
- 14:27
engineering hygiene and practices. But
- 14:29
now the ROI is, I think, finally high
- 14:33
enough for actually us to actually
- 14:34
invest in it. Um one thing that I've
- 14:37
been seeing a lot of teams do is mock
- 14:39
out services. Often with integration
- 14:42
tests, we would test kind of end-to-end
- 14:44
an entire system including live
- 14:46
services. But we've been investing a lot
- 14:49
in in mock services that run entirely
- 14:51
locally with deterministic responses
- 14:54
because it lets the agent do everything
- 14:57
locally. Um doing everything on your
- 15:00
laptop without having to spin up a bunch
- 15:02
of other services and and connect to
- 15:04
cloud services makes everything a lot
- 15:07
faster because the the more that your
- 15:10
agent can get fast feedback means the
- 15:13
more loops that it can can do and the
- 15:15
more productive your own agent can be.
- 15:19
So, across all of these, these are some
- 15:21
of the habits we've seen, but of course
- 15:23
I would be remiss if I would tell you if
- 15:26
you adopt all of these habits, you will
- 15:30
achieve nirvana. You will be the most
- 15:31
productive engineering organization the
- 15:34
world has ever seen. Things are still
- 15:36
hard. We are still very much in an early
- 15:38
adopter phase and teams are still
- 15:42
figuring it out.
- 15:43
So, one thing that we've been seeing
- 15:45
across our teams just organizationally
- 15:48
is the risk of burnout. I did not coin
- 15:51
this term. I forget who did at what
- 15:53
conference, but flow mat is real. We've
- 15:56
been seeing engineers staying up late
- 15:58
late at night
- 16:00
trying to get that perfect prompt that's
- 16:02
going to make their agent run for hours
- 16:04
overnight so that they wake up in the
- 16:05
morning with a code change ready.
- 16:08
The cognitive load increases as you run
- 16:11
these multiple agents in parallel.
- 16:13
You're constantly shifting between
- 16:15
terminal tabs.
- 16:17
And then we do see that reviewing AI
- 16:20
output is often harder for some than
- 16:22
than actually writing it, especially
- 16:24
early in career.
- 16:26
Senior engineers have have already spent
- 16:28
a large portion of their career
- 16:30
reviewing others code.
- 16:32
But early career engineers don't have
- 16:35
that muscle yet and so reviewing it can
- 16:38
can feel like a lot more cognitive load
- 16:41
than they're used to and actually
- 16:42
writing it.
- 16:44
The other one is organizational change.
- 16:47
So, it's already hard to change the way
- 16:50
we work as engineers. The way that we
- 16:51
spend our entire day completely changes
- 16:55
when we're frontier engineers, but also
- 16:57
organizations have to change to enable
- 17:00
frontier engineering teams.
- 17:02
One that I've seen very commonly is
- 17:06
accepting slowing down to speed up.
- 17:09
And I've been guilty of this myself. My
- 17:11
my fellow leaders have been guilty of of
- 17:13
this of saying, "Well, you have the AI
- 17:15
tools now and the models are so amazing
- 17:18
now. Why are you not going faster?
- 17:22
Um and that's because you have to take
- 17:25
those two months to invest in your code
- 17:27
base, to figure out the best practices
- 17:29
for your team, to make hard habit
- 17:33
changes on your team.
- 17:35
Um and and if you're constantly
- 17:37
expecting
- 17:38
shipping features every month because
- 17:40
now we have these amazing models and
- 17:42
we're seeing um all of these these
- 17:45
companies on X saying how they're
- 17:47
shipping 20 PRs a day, um we have to
- 17:51
slow down to speed up.
- 17:54
The second one is actually going too
- 17:55
broad in the organization too fast. I
- 17:58
think that if we had um expected all
- 18:02
teams in massive organizations to be
- 18:04
frontier teams immediately, we would not
- 18:07
have had the learnings that we had from
- 18:10
the Pathfinder, from the from the sprint
- 18:13
experiment, from the pilot uh teams
- 18:16
within Amazon. And now the challenge for
- 18:19
us is how do we scale it out? And that's
- 18:21
what 2026 is about for Amazon is how do
- 18:23
we scale this out to more and more
- 18:25
teams, to the next uh 2,000 teams
- 18:28
instead of uh 50 teams.
- 18:31
Um and so I think that when you roll it
- 18:33
out too quickly, you have a lot of teams
- 18:36
who don't know what they're doing. You
- 18:38
haven't had time to find the best
- 18:40
practices for your own organizations,
- 18:42
the the context that your organization
- 18:44
needs.
- 18:45
And the last one is that you're going to
- 18:47
find new bottlenecks.
- 18:49
Previously, code writing code manually
- 18:52
was the bottleneck. Um I find that
- 18:55
within Amazon, we've found um the speed
- 18:58
of decision-making becomes a new
- 19:00
bottleneck. Um the more that you spend
- 19:03
reviewing the decision to actually build
- 19:06
a new product, the slower it is to build
- 19:09
the product now because the code only
- 19:11
takes 1 to two months to write.
- 19:13
>> [snorts]
- 19:13
>> Um all of the review processes
- 19:16
associated with the launch of a product
- 19:19
become the bottleneck. When it used to
- 19:21
take 9 to 12 months to build a new
- 19:24
product, it didn't matter so much in the
- 19:27
in the overall wash of things if it took
- 19:29
two months to make the decision to build
- 19:31
the product and then two months to
- 19:32
approve the launch. But now those are
- 19:36
the bottlenecks. Those are the long
- 19:37
pole. And so you find all of these all
- 19:41
of these things that slow you down.
- 19:44
Often I find that frontier engineering
- 19:46
teams spend more time making decisions
- 19:49
than they do writing code. And so the
- 19:51
more that you can make fast decisions,
- 19:53
especially ones that are easy to be
- 19:55
reversed, the better.
- 19:57
So my one big takeaway for for everyone
- 20:00
here is that
- 20:02
frontier engineering is about
- 20:04
intentionally changing the way that you
- 20:06
work. And that is difficult. That takes
- 20:09
time. It is forming new habits and a new
- 20:12
way of working.
- 20:14
And that goes across any engineering
- 20:17
team as well as your organization. Um so
- 20:20
I encourage you to think about
- 20:23
um how you're interacting with AI tools
- 20:25
and how that can change to free yourself
- 20:29
up from being in the loop.
- 20:31
Um thanks. I'm going to I'll hang out uh
- 20:33
a little bit if anyone has questions in
- 20:35
the back. Um but thanks for the time
- 20:37
today.
- 20:54
>> [music]