Tokens Should Have Jobs — Katelyn Lesse & Angela Jiang, Anthropic
Read the talk
Tokens Should Have Jobs
Katelyn Lesse and Angela Jiang show why an agent’s token budget is also an allocation problem: advice, evaluation, and reflection can outperform spending the same allowance entirely on execution.
From a talk by Katelyn Lesse and Angela Jiang
At a glance
Ideas worth remembering
Treat an agent budget as an allocation across jobs, not merely a quantity of execution tokens.
Advice, grading, and dreaming intervene at different times: during an attempt, after an attempt, and between runs through memory.
Control the token allowance when comparing strategies. On the reported benchmark, execution scored 76 and advising scored 89 under the same roughly 600,000-token maximum.
Measure the outcome the user can actually use. For the P&L example, anything short of a perfectly scored answer still requires correction or another run.
Choose a strategy for the objective: advising is favored here for token efficiency, while grading or dreaming is favored for single-run reliability.
Claude Managed Agents supplies the concrete individual-agent layer described in the talk; a meta-harness composes roles above it, while automatic strategy construction remains a longer-term goal.
A token budget hides an allocation decision
Katelyn Lesse, who leads platform engineering at Anthropic, and Angela Jiang, who leads platform product, start with the familiar way teams improve an agent: give it more tokens or more expensive tokens. That treats budget as the main lever and assumes every token makes an interchangeable contribution to the result.
A conventional agent receives a task and spends its allowance continuing the work. “Tokens should have jobs” reframes that allowance as compute that can serve different functions. Some tokens can execute the task, while others check the approach, judge an attempt, or preserve a lesson for later. A particular allocation of those roles becomes a strategy.
This does not make individual tokens intrinsically different. The difference comes from the prompts, context, agent roles, and control flow that determine what the model does with them. The practical question changes from “How large is the budget?” to “What work should this budget buy?”
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Advice, grading, and dreaming improve different moments
The three strategies intervene at different points in an agent’s lifecycle:
- Advising — during execution: An executor can ask a separate adviser whether its next step makes sense, then use that response to adjust its work. A sales agent, for example, might use advice while checking overdue follow-ups and stalled deals.
- Grading — after an attempt: A grader compares completed work with an explicit rubric. A passing attempt finishes; a failing attempt returns to the executor for another iteration.
- Dreaming — between runs: A dreamer inspects the executor’s work and transcript, extracts findings, and writes them to memory for the next run.
Grading works best when “good” can be stated clearly. Consider a customer asking a store for a refund. The executor drafts or selects a response; the rubric encodes the store’s refund criteria; the grader checks the proposed outcome against those criteria; and a failed check triggers another attempt. The loop makes policy the reference point, although its reliability still depends on whether the rubric captures the right rules and the grader applies them correctly.
Dreaming targets repeated work rather than the current answer. In the recruiting example, feedback about candidate fit accumulates across interactions. A dreamer turns findings from those interactions into persistent memory, and the next execution reads that memory before proceeding. The described improvement comes from changing future context, not from updating model weights. The talk does not specify how memories are selected, corrected, or retired, so those remain implementation problems.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
The first benchmark confounds strategy with spending
The experiment moves these strategies into a benchmark of financial-analysis tasks intended to resemble work performed by an expert analyst. Execution alone serves as the control; advising, grading, and dreaming are alternative ways to organize the work.
The first one-shot comparison cannot isolate the value of those jobs. Execution scores 15% while using 39,000 tokens, whereas more complex strategies spend more and perform better. Dreaming reaches roughly 600,000 tokens. Strategy and compute have both changed, so the result cannot tell whether the improvement came from orchestration, additional spending, or both.
The next comparison fixes a common maximum allowance at roughly 600,000 tokens. More test-time compute improves every strategy: execution rises from 15 to 76, while advising and grading move from scores in the 60s toward the 90s. The revealing comparison is within the shared allowance: execution scores 76 and advising scores 89. The same budget produces a different outcome when some of it funds consultation rather than continuation.
That reported gap motivates strategy design, but its scope is uncertain: the presentation does not provide the number of tasks, model configurations, repeated-run variance, or statistical uncertainty needed to establish how broadly advising will outperform execution. The supported result is narrower—on this financial-analysis benchmark, allocation mattered under the fixed allowance.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
An 80% accurate P&L is still unfinished work
The benchmark score still leaves a product question unanswered: can an analyst use the result? Suppose the agent produces a profit-and-loss statement that is 80% accurate. That sounds respectable as partial credit, but the analyst cannot safely accept invented or incorrect income and cost figures. They must recompute the statement or pay for another run. Observable progress on the benchmark has not produced a finished P&L.
The experiment is therefore rescored around usable completion. A perfectly scored task passes; anything below 100% fails. This converts the metric from average partial accuracy into the probability that one run returns a fully correct answer. The threshold fits the financial tasks being discussed and should not be assumed for applications where partial results retain value.
Under that criterion, execution passes about 42% of runs, while the more complex strategies reach as high as 75%. A failed run now has an explicit business consequence: its tokens count toward the cost of obtaining the answer, even though its output cannot be used.
The talk rounds execution’s pass rate to about 40% and budgets approximately three attempts. Assuming 600,000 tokens per attempt, that produces the headline estimate of 1.8 million tokens for a perfect answer: 3 × 600,000 = 1,800,000. This is a practical three-run estimate, not a guarantee or a precise expected-value calculation.
Applying retry cost across the strategies changes the decision. Advising and grading are described as comparatively token efficient because extra work inside a run can reduce the need to repeat the whole task. If cost per usable result is the priority, the recommendation for this domain is advising. If single-run reliability matters more, grading or dreaming becomes more attractive. Token efficiency and per-run reliability are related, but they are not the same objective.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
A meta-harness composes the jobs into one workflow
The implementation separates two layers. An individual agent harness supports execution. Above it, a meta-harness coordinates the executor, adviser, grader, dreamer, and any other participating roles. The lower layer makes one agent capable of doing work; the upper layer decides how multiple kinds of work fit together.
The concrete lower layer in the presentation is Claude Managed Agents, Anthropic’s managed-agent offering on the Claude platform. Jiang says its architecture supplies the individual-agent harness, while orchestration above it coordinates the roles in a strategy. She also says some capabilities—including dreaming—are available there out of the box. The talk does not detail the harness’s internal components or provide an implementation guide for those managed capabilities.
What does the composed control flow look like? A task enters the executor, which can consult an adviser while working. The result goes to a grader. A failure returns to execution; a pass proceeds to dreaming, where findings are written to memory for the next run.
The relationship visible in this workflow is temporal: advice changes the current attempt, grading governs whether that attempt may finish, and dreaming changes future attempts. Combining them can cover all three moments, but the talk presents this composition as an architectural construction rather than a separately benchmarked result.
These roles are primitives, not a closed taxonomy. Builders can introduce new jobs and more complex coordination patterns for their own tasks. The longer-term goal is for models and the platform to construct strategies dynamically as work proceeds. That automatic construction is presented as future direction; meanwhile, builders can combine the available primitives and evaluate each strategy against the outcome, reliability target, and cost that matter for the application.
The requested outcome enters the strategy.
Advice affects work in progress, grading controls retries, and dreaming carries findings into the next run.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Further reading
A longer conversation with Lesse and Jiang that places strategies and meta-harnesses within Anthropic’s knowledge, execution, and coordination layers.
- Evolving Claude APIs for AgentsReference
Lesse explains the lower-level agent primitives—reasoning budgets, tools, memory, context editing, and sandboxed execution—on which higher-level strategies can be built.
Related talks
- Claude for long-horizon tasks
A related topic for readers considering how these strategies might apply to longer tasks.
- How We Build Effective Agents
A companion topic for readers exploring agent design beyond token allocation.
- 12-Factor Agents: Patterns of reliable LLM applications
A related discussion of reliability patterns, complementing this talk’s distinction between efficiency and single-run success.
Read the complete timestamped transcript
- 0:19
Good morning. We're super excited to be
- 0:22
here at AI Engineer with all of you. I'm
- 0:25
Caitlyn and I lead platform engineering
- 0:26
at Anthropic.
- 0:28
>> And I'm Angela. I lead platform product
- 0:29
at Anthropic. And today we want to talk
- 0:31
to you about a concept that we've been
- 0:33
spending a lot of time thinking about
- 0:35
and working on with our team, which is
- 0:37
this idea that we think that tokens
- 0:39
should have jobs.
- 0:41
So if you're building an agentic system
- 0:43
and you're trying to accomplish some
- 0:45
specific outcome, you're trying to get
- 0:46
something done with agents, there's one
- 0:49
lever that everybody pulls in order to
- 0:50
get a better outcome, and that's usually
- 0:53
increasing your budget, which means you
- 0:54
spend more tokens or you spend more
- 0:56
expensive tokens.
- 0:59
But we've been wondering is that all
- 1:01
there is underlying this assumption of
- 1:05
uh using the budget is this kind of
- 1:06
implicit perspective that every single
- 1:09
token is basically fungeable. And we've
- 1:11
been wondering is that actually true?
- 1:13
Are all these tokens actually fungeible?
- 1:15
And to test that, we've been thinking,
- 1:17
what if we gave tokens jobs?
- 1:20
So, if you think about the way that you
- 1:21
would normally set up an agent to go
- 1:23
accomplish a task, you give it that
- 1:24
task, you give it this token budget, and
- 1:26
then all the tokens that are being spent
- 1:28
are basically indiscriminate in the
- 1:30
sense that they're all doing one job.
- 1:31
They're just executing.
- 1:34
But what if you take some of those
- 1:35
tokens and they're not just executing,
- 1:37
they're doing some other job. So, for
- 1:40
example, maybe you take some of your
- 1:42
tokens and they're advising the tokens
- 1:44
that are executing. Or maybe the tokens
- 1:47
that are executing try to get something
- 1:49
done well and you take some other tokens
- 1:51
and you actually grade how well the
- 1:52
executor is doing so that it can iterate
- 1:54
and try again. Or maybe you have tokens
- 1:58
that are dreaming. They're reflecting
- 1:59
back on the job that other executors
- 2:02
have done and writing learnings to
- 2:04
memory so that they can do it again. And
- 2:07
what we call each of these if you take
- 2:09
some tokens that are executing and some
- 2:10
tokens that are doing some other job.
- 2:12
Let's call this a strategy.
- 2:16
So let's go take a look at the first
- 2:17
strategy, the advising strategy. Here
- 2:19
we're splitting up an executor and an
- 2:21
adviser. The executor obviously
- 2:23
executes, but crucially they can call
- 2:25
out to an adviser for advice. And then
- 2:27
they can take this advice and figure out
- 2:28
if they're doing the next step
- 2:29
correctly.
- 2:31
This is really helpful in use cases. For
- 2:33
example, if you're building a sales
- 2:34
agent, in an ideal world, you'd have
- 2:36
that sales agent be able to actually
- 2:37
help the sales rep flag when a follow-up
- 2:39
is overdue or deal is stalling. In this
- 2:42
construct, having an adviser to be able
- 2:44
to kind of make sure that all the
- 2:45
different pieces are actually working is
- 2:47
really helpful.
- 2:49
So another example is grading. Let's say
- 2:51
you're executing and you kind of know
- 2:54
exactly what good really does look like.
- 2:56
You can define this in a rubric and then
- 2:58
each time an executor tries to
- 3:01
accomplish that outcome, you can have a
- 3:03
grader provisioned that grades how well
- 3:05
the executor did while looking at that
- 3:07
rubric. And if the executor did a good
- 3:10
job, then great, it can be done. But if
- 3:12
it didn't do such a great job, you can
- 3:13
iterate again until you get that good
- 3:15
outcome.
- 3:17
So an example in practice of when you
- 3:19
might want to use this is let's say you
- 3:20
have a customer service agent and you're
- 3:22
running a store and your customers are
- 3:24
writing in and they're saying, "H, I
- 3:26
should get a refund for this thing." And
- 3:27
your customer service agent needs to be
- 3:29
able to respond. You probably have some
- 3:31
like pretty specific criteria on when
- 3:33
you would give somebody a refund. And so
- 3:36
what you can do is define a rubric that
- 3:38
uses that criteria. You can have a
- 3:40
grader that goes and looks at the work
- 3:42
that the customer service agent is doing
- 3:44
and decide is it getting it right and is
- 3:46
it coming to the right outcome.
- 3:49
And the last strategy we have is
- 3:50
dreaming. So in dreaming there's an
- 3:52
executor who naturally executes and then
- 3:54
there's a dreamer. The dreamer is
- 3:56
actually able to inspect the work and
- 3:58
the transcripts of the executor and then
- 4:00
it takes any of the findings that it has
- 4:02
and it writes them to memory. This
- 4:03
memory is repicked up by the executor
- 4:05
for the next round. So ideally would
- 4:07
have improved.
- 4:09
A great use case for this is if you're
- 4:10
building a recruiting agent. Now
- 4:12
recruiting requires a lot of interaction
- 4:14
with feedback on whether or not a
- 4:15
candidate does or doesn't make sense and
- 4:17
if it's a good fit between both parties.
- 4:19
And so by taking all this type of data,
- 4:20
if you build a dreaming type of strategy
- 4:22
on this agent, it's actually able to
- 4:24
kind of sharpen the next round so that
- 4:25
it's more and more increasingly useful.
- 4:28
So let's make this concrete with some
- 4:30
experiments. So what we did was we
- 4:33
created a bench of a bunch of tasks
- 4:35
related to financial an analysis. And
- 4:37
what we were doing with each of these
- 4:38
tasks is trying to replicate in the real
- 4:40
world a expert human financial analyst.
- 4:43
How well would they do on each of these
- 4:45
various tasks? And so what we did was we
- 4:47
start with a control that's just
- 4:49
executing. Let's try each of these tasks
- 4:51
and we'll eval them when we're literally
- 4:52
just executing. But then we can
- 4:54
experiment with each of our strategies
- 4:56
and see how well we perform.
- 5:00
So, we start with a super basic
- 5:01
experiment. Let's just oneshot it. Let's
- 5:03
take each of our strategies and we'll go
- 5:05
and just make an attempt to accomplish
- 5:07
these tasks and we'll see how accurate
- 5:09
we are. And so, you can see here with
- 5:11
executing um it didn't do so well. 15%
- 5:14
accuracy, but because it was just a
- 5:16
oneshot, the strategy got to choose how
- 5:18
many tokens it would actually spend on
- 5:20
its own. And so, you can actually see
- 5:22
that execute decided not to spend that
- 5:24
many tokens, only 39,000. And as we go
- 5:27
into our larger strategies, our more
- 5:29
complex strategies, we did choose to
- 5:31
spend more tokens, but we did a better
- 5:32
job. So, this isn't really telling us
- 5:34
much because sure, Drain did really,
- 5:36
really well, but it used a whopping
- 5:38
600,000 tokens to get there. That's
- 5:40
right. So, in order to actually figure
- 5:42
out if varying the jobs produces any
- 5:45
alpha, what we need to do is hold the
- 5:46
budget constant. And to do this, we're
- 5:48
going to take Dreaming's budget, that
- 5:49
600,000 or so, as the maximum budget
- 5:51
that is fixed across the board. And we
- 5:53
give every single strategy this budget
- 5:55
in order to analyze how well it's
- 5:57
performing. And as expected again, if
- 6:00
you give a lot of strategies more
- 6:01
budget, you are going to see performance
- 6:03
increase across the board. So execute
- 6:04
went from 0.15 to 76. Advise and grade
- 6:08
went from the 60s to closer to the 90s.
- 6:10
And that's again expected given the fact
- 6:12
that if you give things more test time
- 6:14
compute, they should generally perform
- 6:16
better. But if that was the only thing
- 6:18
that mattered, we should actually expect
- 6:20
to see execute, advise, grade, dream
- 6:22
actually all be at the exact same level
- 6:24
given the exact same token budget. But
- 6:26
what we're actually seeing is that there
- 6:28
is an alpha or there is a difference and
- 6:30
therefore an alpha for us to exploit. If
- 6:32
you look at execute at this exact same
- 6:34
budget level, it gets to 76 but advise
- 6:37
is at 89. So while a minimal, it does
- 6:40
exist and so there is alpha for us to
- 6:41
take a look at.
- 6:44
Now, we decided to take a look at this
- 6:45
analysis from a completely different
- 6:47
lens. And as Kayla mentioned, you know,
- 6:49
we're doing this bench for a very
- 6:51
complex set of financial tasks in the
- 6:53
real world. And we wanted to analyze the
- 6:56
usage of agents with actual experts. So,
- 7:00
if we look at a financial analyst
- 7:02
expert, right, the kind of task that
- 7:03
they need to do with an agent is that
- 7:05
they're giving it something very
- 7:07
concrete like let's say make a P&L and
- 7:09
then they're getting the result back.
- 7:11
Now if that result is 80% accurate on a
- 7:14
bench that sounds great but in reality
- 7:16
what that means for that expert is they
- 7:18
have to go back and recomputee that P&L
- 7:20
themselves or alter or alternatively
- 7:22
send it through another run and that's
- 7:24
because in this kind of domain for this
- 7:25
kind of task if you're not 100% accurate
- 7:28
it's actually not useful.
- 7:31
You cannot make up an income number or
- 7:32
you can't make up a cost number right
- 7:34
you have to make sure that it's 100%
- 7:35
accurate. So with this lens of the real
- 7:37
world consequence associated with this
- 7:39
domain, we needed to recomputee our
- 7:40
experiments and score them a bit
- 7:42
differently. Crucially, we needed to
- 7:44
make sure that our experiment had this
- 7:46
kind of construct where if it was scored
- 7:48
perfectly, we'd actually give it a pass.
- 7:50
And if it scored anything less than 100%
- 7:52
on that kind of task, we would actually
- 7:53
mark it as a failure.
- 7:56
So let's look at a different cut of our
- 7:58
data from our experiments with this lens
- 8:00
where we're looking for this perfect run
- 8:02
100% accuracy pass. And let's look at
- 8:05
what percent of the time each of these
- 8:07
strategies was able to achieve a pass.
- 8:09
Um so we we've got executes um down at
- 8:12
42% and we've got our more complex
- 8:14
strategies doing a bit better up to 75%
- 8:17
accuracy. Um and again this doesn't
- 8:20
necessarily tell us a ton because um you
- 8:22
know each of these strategies might um
- 8:25
choose to use different budgets over
- 8:26
time, right? So what we did here was we
- 8:28
fixed the budget and we said within a
- 8:30
fixed budget, how well do each of these
- 8:32
strategies perform?
- 8:35
And so what really matters to us
- 8:37
actually is if you're trying to get this
- 8:39
perfect answer and you're in the real
- 8:41
world, you're running a business, what
- 8:42
matters to you is the cost to you to get
- 8:45
to that perfect answer. And so one way
- 8:47
we can think about this is we had our
- 8:49
execute strategy for example. The
- 8:50
execute strategy around 40% of the time
- 8:53
will give you that perfect answer. So on
- 8:55
average, you can expect to have to run
- 8:56
it three times and you should hopefully
- 8:58
sometime in those three runs get a
- 9:00
perfect answer. And as we talked about
- 9:02
earlier, we fixed our budget to that
- 9:04
highest token budget strategy, which was
- 9:06
600,000 tokens. So if you spend 600,000
- 9:09
tokens in each individual run, you have
- 9:11
to run approximately three times. You
- 9:13
can expect on average to have to spend
- 9:15
1.8 million tokens with the execution
- 9:17
strategy to get to your perfect answer.
- 9:21
And so if we take this analysis and run
- 9:23
it across the board against all these
- 9:24
strategies, this is actually the true
- 9:26
cost it took in this domain for that
- 9:29
agent to be useful for that strategy. So
- 9:31
as Caitlyn mentioned for execute, which
- 9:33
is our baseline, this is going to be 1.8
- 9:35
million true total token cost for you.
- 9:38
But advise, grade, and dream are showing
- 9:40
us a bit of difference. Crucially,
- 9:42
advise and grade are actually quite
- 9:44
token efficient when you think about the
- 9:46
actual usage of the end output of each
- 9:48
of these agents.
- 9:51
So what does this mean for you as a
- 9:52
business? Well, it actually really
- 9:54
depends on what kind of thing you want
- 9:56
to optimize for and it's going to vary,
- 9:58
right? There's going to be businesses
- 10:00
who say, "Actually, for me, the most
- 10:01
important thing is to be really token
- 10:03
efficient. In that case, you should
- 10:05
probably pick the advised type of
- 10:06
strategy in order to solve for that
- 10:08
particular domain in which you want to
- 10:09
optimize that." There's going to be
- 10:10
other areas or other businesses where
- 10:12
you're going to say, I'm not going to
- 10:14
care so much about token efficiency
- 10:15
because what I really care about is
- 10:16
reliability of that answer and so I need
- 10:19
to maximize the percentage of runs in
- 10:20
which I get that perfect answer. In
- 10:22
which case, you would actually pick
- 10:23
completely different strategies. You
- 10:24
probably lean towards grade or dream.
- 10:28
So if you take away one thing, the thing
- 10:30
we want everyone to think about is this
- 10:31
idea that tokens are not fungeible. You
- 10:34
can use your tokens to execute. You can
- 10:36
brute force your way through your task
- 10:37
and you can throw more budget at it. But
- 10:39
if you get really smart about having
- 10:41
your tokens do these different jobs and
- 10:43
try these different strategies, you're
- 10:45
very very likely to be able to get a
- 10:46
better outcome for the task at hand
- 10:49
within a fixed budget.
- 10:52
And so let's talk a little bit about how
- 10:53
we actually build strategies and how we
- 10:55
bring this to life. Um so we've done a
- 10:57
lot of work to create a really excellent
- 10:58
harness for individual agents. Um and if
- 11:01
you see this uh picture at the bottom
- 11:03
here, this is actually the architecture
- 11:05
that we've used for cloud managed agents
- 11:07
um which is our Aentic solution that we
- 11:09
give to you within the cloud platform.
- 11:12
And what we do on top of this is we
- 11:14
start to get into the meta harness level
- 11:15
like the multi- aent orchestration and
- 11:17
execution level where this strategy can
- 11:20
go and be um coordinated between our
- 11:23
executor and our adviser or the other
- 11:25
agents within our strategy. And some of
- 11:28
these um like dreaming and outcomes we
- 11:30
actually give to you out of the box
- 11:31
within cloud manage agents.
- 11:34
So with those set of primitives it's
- 11:36
actually relatively trivial for us to
- 11:38
construct this kind of you know
- 11:40
architecture where we're able to combine
- 11:41
these different types of strategies and
- 11:43
figure out how to they should work
- 11:45
together. So for example it's relatively
- 11:47
trivial for us to say okay now with this
- 11:49
I can take a task and I should be able
- 11:51
to execute it but also allow it to
- 11:53
advise and fable is back online. So we
- 11:55
could actually say Fable is the one
- 11:57
that's actually advising uh the
- 11:58
executor. And then I can take all these
- 12:00
results and say send them to a greater
- 12:02
so that I can make sure that this is
- 12:03
verifying in a loop that makes sense.
- 12:05
And if it passes, that's awesome. I want
- 12:07
to send all of that stuff to Dreaming
- 12:09
and make sure that my next run is better
- 12:10
than ever.
- 12:12
And of course, you don't have to stop
- 12:14
there, right? If the right primitives
- 12:15
are there and the right coordination is
- 12:17
there, then you can actually construct
- 12:19
really complex setups that fit for all
- 12:21
the different types of dynamic problems
- 12:23
that you have. You can invent these
- 12:25
kinds of large-scale architectures,
- 12:27
again, very triv. And you could also
- 12:29
invent completely new jobs, not just the
- 12:31
ones of the pieces that Caitlyn and I
- 12:33
have presented in this conversation.
- 12:36
So, a big goal that we have over time is
- 12:38
to get our models better and better and
- 12:40
our platform better and better at
- 12:42
dynamically constructing these
- 12:43
strategies for you as you're doing work.
- 12:45
But in the meantime, as we're working
- 12:47
our way there, we would love for you to
- 12:49
continue to think about this idea that
- 12:50
you should give your tokens jobs and you
- 12:52
should use different novel strategies by
- 12:54
combining these primitives in order to
- 12:56
get the outcomes that you want for your
- 12:58
tasks.
- 12:59
>> Thanks for joining us.