Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori
Read the talk
Mousepower: Agents That Can’t Be Measured Can’t Be Managed
Maximillian Piras explains why token counts fail to show whether agents create value, why faster generation moves the bottleneck to verification, and how task uncertainty can help identify work worth delegating.
From a talk by Maximillian Piras
At a glance
Ideas worth remembering
Token volume measures system activity, not customer value. Trace spend to accepted outcomes such as bugs closed or support requests resolved, then connect those outcomes to an objective.
Faster generation moves the bottleneck to verification. Agent output creates value only after its quality can be judged and the work can be accepted or deployed.
Mousepower is not a literal cursor-speed metric. It is the requirement to give customers an understandable rubric for deciding whether agent work was worth its cost.
Use deterministic scripts when task steps are predictable, and avoid delegation when checking requires doing the work again. The best agent candidates combine moderate execution uncertainty with relatively clear acceptance criteria.
When verification follows a repeatable pattern, an agent can help check another agent—but the checking rubric must still reflect the customer’s definition of good work.
Parallel agents are fun until the bill arrives
Maximillian Piras, founding designer at Yutori, starts with a familiar agent workflow: keep human attention on the main task while several agents pursue peripheral work in the background. For this talk, the concrete job is improving the slides. A design system and written guidance constrain the work; multiple agents then explore typography and layout in parallel. The observable result is more design variants produced without taking over his active attention. 0:42
Parallelism feels like free leverage because research and design exploration rarely have an obvious stopping point. But each additional run consumes tokens, and the bill eventually forces a question the workflow postponed: which explorations were useful enough to justify their cost? A pile of generated variants is visible output. It is not yet evidence of value. 1:42
That question is especially relevant to Yutori’s computer-use models. These models operate software as a person would when an API or MCP connection is unavailable. They can reach information trapped behind ordinary interfaces, but clicking through a rendered application is less efficient than calling a structured API. Computer use is therefore a fallback for the long tail of inaccessible systems, not the preferred interface when a direct integration exists. 2:42
Customer conversations repeatedly return to the same unresolved decisions: which jobs suit an agent, what tradeoff to make between token cost and value, and how to tell whether the result is good. Excitement is widespread, but so is the feeling of merely “scratching the surface.” The missing muscle is measurement. 3:42
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Horsepower made an unfamiliar machine legible
Experienced agent users can object that they already know how to run a fleet. Piras accepts that experience while questioning whether it generalizes. Early adopters enjoy experimenting, tolerate rough interfaces, and may willingly spend tokens to discover what works. People who have never used an agent—and may still work through copy-and-paste interactions—do not share that calibration. A product cannot explain its value to the wider market by pointing to the habits of its most enthusiastic users. 4:13
The historical analogy is James Watt selling steam engines to people who understood power through horses. A horse gin connected an animal to a rotary arm; walking in a circle supplied mechanical power to a mill. Replacing that familiar arrangement with an intimidating machine required more than claiming that steam was better. Buyers needed a comparison expressed in terms they already understood. 5:43
Watt studied horse gins and developed horsepower as a baseline for expressing relative performance. In Piras’s telling, the early figure was neither especially scientific nor necessarily accurate. Its commercial utility came from legibility: a buyer could translate machine output into a familiar unit and estimate the gain from trying the new technology. The measure reduced the cognitive distance between the old system and the new one. 7:12
The comparison also had to overcome attachment to the incumbent technology—horses, as Piras puts it, had “great vibes.” Efficiency alone does not create adoption when the new system feels unfamiliar or threatening. A comprehensible return on investment gives someone a reason to cross the threshold from curiosity to use. 8:12
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Tokens are cost accounting, not value accounting
The agent industry has not found an equivalent translation layer. Organizations can exhaust an annual token budget in a quarter or turn consumption into a leaderboard, but neither behavior establishes that the spending improved the business. When incentives reward usage rather than useful work, teams can become very good at consuming the resource they meant to evaluate. 9:12
Borrowing a term from Ramp, Piras describes the resulting pattern as “overspending and underusing”: teams maximize agent activity, hit austerity, retreat from the tools, and return when fear of missing out becomes strong enough. What does this cycle make visible? Enthusiasm without outcome measurement does not produce sustained adoption; it produces alternating bursts of consumption and withdrawal. 9:42
The diagram shows the missing link. Spending produces activity, but without an outcome measure the team cannot distinguish productive use from waste. Budget pressure therefore causes a broad retreat instead of a targeted adjustment to models, tasks, or workflows.
Choosing cheaper models by default and reserving frontier models for difficult work can separate token volume from total spend. That is useful cost control, but it still treats tokens as the central unit. Tokens are an internal output of the system. The business cares about what changed afterward. 10:12
The practical accounting chain should run from spend to a concrete outcome and then to an objective. For software work, ask how many bugs were closed with the purchased tokens. For support, ask how many requests were resolved. Return to the slide example: ten agents may create ten alternatives, but value appears only when review identifies a better treatment and that treatment improves the finished deck. Counting variants or tokens stops before the consequential step. 10:42
Run more agents and consume more tokens in pursuit of additional work.
Without outcome-based measurement, budget pressure interrupts adoption instead of teaching the team which agent work is valuable.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Generation got faster; review did not
Coding agents expose the deeper problem. They can generate changes quickly enough to leave teams “dying by a thousand pull requests,” while code review still consumes human attention. More generated code does not automatically create more accepted, deployed, useful software. When generation accelerates but review does not, the bottleneck moves downstream. 11:42
This creates a measurement problem as well as a workflow problem. A team can observe code volume immediately, but it cannot judge quality at the same speed. Until someone reviews the change and decides whether it is good enough to use, token spend has produced a proposal rather than established value. The original bottleneck may have been implementation; the new one is verification. 12:12
Code review nevertheless offers a useful clue because software teams have converged on shared assumptions about what acceptable work looks like. That convergence creates a rubric, even if its assumptions need revision for agent-generated code. Once people agree on what should be checked, parts of the judgment can happen consistently and at greater scale. 13:12
The product requirement follows: building an agent also means helping its user verify the output and connect accepted work to return on investment. Execution at the speed of compute is only half the system. Useful deployment needs measurement that can approach the same speed without pretending that generation itself proves success. 13:49
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Mousepower is a design requirement, not a unit
Mousepower is Piras’s proposed analogy for the agent era: establish a baseline for how people use computers, then express where an agent performs better on the task. The tempting implementation is literal—measure cursor distance or speed and compare human movement with agent movement. Piras even had a prototype measurement device generated for the joke. 14:19
The literal metric fails because computer work takes place in a high-dimensional information space. Faster pointer movement says little about whether the correct source was selected, the right judgment was applied, or the resulting action advanced the user’s goal. Two agents can traverse the same interface at similar speed and produce outcomes of radically different value. 15:19
Mousepower therefore names a product obligation rather than a standardized unit. If a product sells agent execution, it should also supply a rubric for determining whether that execution was good and whether its token cost was justified. Piras does not offer a universal formula for that rubric; it must fit the customer’s task and mental model. 15:49
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Choose tasks by two kinds of uncertainty
To narrow the search for suitable tasks, Piras borrows the idea of entropy as uncertainty. The proposed matrix has two dimensions: uncertainty in the steps required to complete a task, and uncertainty in the criteria used to accept the result. This is explicitly a thought starter still in development, not a validated task-selection model or a quantified use of information-theoretic entropy. 16:19
The first axis asks how predictable the route is. Booking a flight has required information and recognizable stages: departure, destination, and a seat selection of some kind. Painting a masterpiece has no comparable sequence that reliably yields success. The contrast concerns path structure, not whether either task can sometimes be completed by a model. 16:49
The second axis asks whether success has a clean grading rubric. An agent might be able to perform an open-ended task, yet still be a poor product if a person must redo the work to determine whether it was valuable. Verification cost belongs in the task-selection decision from the beginning. 17:49
What does the matrix help a builder decide? It places uncertainty in execution on the horizontal axis and uncertainty in acceptance on the vertical axis. Reading across distinguishes deterministic scripts from useful agent work and poorly modeled tasks. Reading upward shows how rising verification uncertainty makes delegation less economical. The promising region is moderate step uncertainty paired with relatively clear acceptance criteria.
The regions imply four practical decisions:
- Predictable steps — write a script. Deterministic automation avoids spending tokens on a path already known in advance.
- Extremely uncertain steps — reconsider the task. The work may fall outside useful training distribution and provide sparse reward signals.
- Unclear acceptance — avoid false delegation. If a person must effectively repeat the task to judge it, little work has been saved.
- Moderately uncertain steps with clear acceptance — consider an agent. The task needs judgment, but its result remains economical to check.
Piras informally compares the final region to an NP-style problem: solving is harder than checking. The important engineering property is the asymmetry, not a formal complexity classification. When acceptance follows a repeatable pattern, another agent can perform part of the verification. The resulting system contains an execution agent and a checking agent, both working against criteria the customer understands. That is how an agent begins to acquire its mousepower. 19:49
The path and acceptance criteria are predictable, so deterministic automation is cheaper.
Horizontal position represents uncertainty in the execution steps; vertical grouping represents uncertainty in acceptance criteria. The best agent candidates sit in the low-acceptance-uncertainty row and the moderate-step-uncertainty column.
Suggest correction
This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.
Resources
Related talks
- ReviewDebt: a practical framework for scoring every pull request — Sachin Gupta, eBay
Extends the review-bottleneck argument with a deterministic framework for estimating the human burden created by pull requests, especially AI-generated changes.
- Agent Evals: Finally, With The Map
Provides a complementary map for evaluating semantic quality, behavior, tool use, and the separate workflow that judges an agent.
- How We Build Effective Agents
Explains why coding is well suited to agents when outputs can be checked through unit tests and continuous integration, reinforcing the asymmetry between execution and verification.
Read the complete timestamped transcript
- 0:01
[music]
- 0:12
>> Um well, thanks all for your time.
- 0:14
Really appreciate you dropping by and
- 0:16
it's always a a great honor to speak at
- 0:18
the world's fair. So, I'll do my best to
- 0:21
uh
- 0:21
give you guys some valuable insights and
- 0:23
um
- 0:24
yeah, hopefully make it worth your time.
- 0:26
So, my name's Maximilian Piros and today
- 0:28
I'll be talking about mouse power.
- 0:31
And this is a talk about measuring
- 0:33
agents through mental models.
- 0:35
But before I get into talking about
- 0:37
measuring agents, I'm going to talk
- 0:39
through a bit about how I use them every
- 0:40
day. And it might seem familiar to you,
- 0:43
but just to level set, we'll go through
- 0:44
it. So, uh I tend to background them
- 0:46
like I'm sure a lot of you people are as
- 0:48
well.
- 0:49
Um so, while my active attention is
- 0:51
focusing on one thing, like perhaps
- 0:53
giving this talk to you, I still want to
- 0:55
make some progress on peripheral tasks.
- 0:57
So, I'll keep my attention focused on
- 1:00
giving this talk, while my agents can
- 1:03
help me explore some designs in the
- 1:04
background cuz I think that uh my slides
- 1:07
need a bit of work. So, you know, I've
- 1:08
got my design system already set up.
- 1:10
I've got some guidance uh given to my
- 1:12
agents. So, I'll kick off an agent to
- 1:14
just try to explore some different
- 1:16
directions on the type type two
- 1:17
treatment and the layout. And um you
- 1:20
know, just try to get as many
- 1:21
explorations as possible.
- 1:23
But of course, uh one agent's never
- 1:24
enough. So, I like to kick off a bunch
- 1:26
in parallel. You know, I've got a lot of
- 1:28
slides to get through. So, I need all of
- 1:30
my agents on and exploring it in
- 1:32
different directions and
- 1:34
uh hopefully I can get some interesting
- 1:35
things to make my slides a bit better.
- 1:38
And uh hopefully they can finish the job
- 1:40
pretty soon cuz we're obviously kind of
- 1:42
up against uh the timeline the the uh
- 1:44
deadline here.
- 1:46
So, um
- 1:47
this is generally how I work. I'm sure
- 1:48
it's probably familiar to a lot of you
- 1:50
where we're just trying to kick off
- 1:51
agents for as much as possible in
- 1:53
parallel cuz it always feels like
- 1:55
there's just way more research to do. We
- 1:57
want it to be as thorough as possible.
- 1:59
There's way more design explorations to
- 2:01
do. So, whenever our main focus is on
- 2:03
one thing, why not kick a bunch of
- 2:05
agents off in parallel and just try to
- 2:07
maximize your time. And it's a lot of
- 2:10
fun, of course, until you get the bill.
- 2:13
And then you start to wonder, was it all
- 2:15
worth it? Right? Did
- 2:17
did you vibe code too hard?
- 2:20
Were you token maxing too much? Like,
- 2:22
could you have been more efficient in
- 2:24
how you approached
- 2:26
your sequencing your agents?
- 2:28
And so, this is what I'm going to get
- 2:29
into today. It's It's how do we evaluate
- 2:31
the token cost? And specifically, how do
- 2:33
we help our customers value it?
- 2:36
So, for the past year and a half, I've
- 2:38
had the pleasure of working as a
- 2:39
founding designer at a company called
- 2:40
території and we focus on computer use
- 2:43
models. These are models that learn to
- 2:45
use a computer like a human would and
- 2:47
the use case for them is when you can't
- 2:49
get information from an API or an MCP,
- 2:52
why not just send an agent out to use
- 2:54
the computer like a human would and then
- 2:56
we can extract all types of data and
- 2:58
manipulate it in ways that let us access
- 3:01
all the stuff that wasn't accessible
- 3:03
previously. So, obviously less efficient
- 3:04
than APIs and MCPs, but as a last
- 3:07
resort, just have an agent go use the
- 3:09
computer and try to get the information.
- 3:12
Here's the U території agent using the U
- 3:14
території website.
- 3:16
Um checking out its own benchmark. So,
- 3:18
kind of it's admire itself in a way. So,
- 3:21
yeah, it gets it gets a bit weird. Like,
- 3:23
and a lot of what I do as a founding
- 3:25
designer there is talk to customers, try
- 3:27
to understand how can we make agents as
- 3:29
intuitive as possible, how do we figure
- 3:31
out the mental models they're using to
- 3:33
value the use cases they want to send
- 3:35
out agents for.
- 3:36
And a lot of them do seem pretty
- 3:39
confused so far. A lot of people are
- 3:41
excited about agents, but the phrase
- 3:43
that comes up quite often is that they
- 3:45
feel like they're just scratching the
- 3:46
surface. Um
- 3:47
it seems like it's not quite intuitive
- 3:48
how we can best use them yet. And so, in
- 3:51
a lot of my uh customer discussions,
- 3:53
it's it always comes down to a question
- 3:55
of like, what is the best way to use use
- 3:57
uh to use agents? What are the best use
- 3:58
cases for them? And how do I think about
- 4:00
the trade-offs with regards to token
- 4:02
costs relative to value? So, I think
- 4:04
we're still kind of building this muscle
- 4:06
today. And this leads me to the thesis
- 4:08
of the talk, which is that I think
- 4:10
agents have a measurement problem.
- 4:13
And uh as an example, here's me at work
- 4:15
trying to measure some agents, and one
- 4:17
of my co-workers took this photo and
- 4:19
told me I looked like
- 4:20
I was uh trying to solve the mystery of
- 4:22
Pepe Silvia. So,
- 4:23
uh
- 4:24
as you can see, it's it's
- 4:26
it's not a it's not an easy task to to
- 4:28
measure agents. But, I'm sure some of
- 4:30
you are saying, "Hold on a sec. Like,
- 4:32
what is this guy talking about? I've got
- 4:34
a fleet of agents working for me right
- 4:35
now. We're building our next
- 4:37
million-dollar app as we speak, and I'm
- 4:39
having a a totally fine time uh
- 4:41
measuring my agents." Uh to which I will
- 4:43
agree with you, uh but then I will point
- 4:45
you to the mandatory Upton Sinclair
- 4:48
quote to remind us all that everybody in
- 4:50
this room is very biased, and we're
- 4:52
early adopters, and we're very excited
- 4:55
to explore this new technology, but it
- 4:57
doesn't mean that we represent the
- 4:58
people that ultimately we're going to be
- 5:00
trying to help adopt this technology.
- 5:02
And so, you know, I think it's important
- 5:04
to remind ourselves that in some way or
- 5:06
another, we probably are selling tokens,
- 5:08
whether it's indirectly or directly. And
- 5:11
so, when we think about our own token
- 5:12
usage, uh is it really representative of
- 5:14
all the people out there who have never
- 5:15
touched an agent yet? Uh some people are
- 5:18
still copy and pasting into ChatGPT.
- 5:20
I may be [REDACTED] to one of these people,
- 5:22
and despite how much I tried to get her
- 5:24
to try out agents, she's not let me uh
- 5:26
set her up with it yet. And so, um as a
- 5:29
reminder, uh when we think about helping
- 5:31
people adopt agents, you know, all the
- 5:32
people across the world that we think
- 5:34
could get as much uh excitement and
- 5:36
value as as we do when we run off
- 5:37
parallel agents, let's uh just remember
- 5:39
this quote.
- 5:40
>> [snorts]
- 5:41
>> And um
- 5:42
so, it really boils down to the age-old
- 5:44
problem of a new technology.
- 5:47
And of course, there's tons of history
- 5:49
we can go to to study how people solved
- 5:51
this in the past. We have this really
- 5:53
exciting new thing, but we haven't quite
- 5:55
uh figured out the right ways to
- 5:56
communicate it.
- 5:57
And so for this talk, I'll go back to
- 5:59
the 1700s and we can take some notes
- 6:01
from when James Watt was trying to sell
- 6:04
steam engines.
- 6:06
And at the time he decided that uh a
- 6:08
great use case for his steam engines was
- 6:09
trying to replace a horse gin. And these
- 6:13
are the was the power source of a mill
- 6:15
at the time. So when you're
- 6:16
for uh let's say a brewery and you you
- 6:18
need some power source to to grind your
- 6:21
barley or whatever. I don't know. I'm
- 6:22
not I'm not like a big brewery guy, so I
- 6:23
don't know exactly what how it's made,
- 6:25
but you need a power source and the
- 6:27
power source at the time that was common
- 6:29
was you hooked a horse up to a rotary
- 6:30
arm and the horse walked in a circle and
- 6:32
that's how you generated your power.
- 6:34
Uh and seems crazy today maybe, but um
- 6:37
at the time was commonplace and Watt
- 6:38
thought, you know, it would be much
- 6:40
better than a horse is like a very
- 6:41
efficient machine.
- 6:43
Although he um rightfully acknowledged
- 6:45
that one of the big barriers to adopting
- 6:47
it would be this cognitive dissonance of
- 6:50
trying to tell people who kind of think
- 6:52
in horses, how do you adapt to this to
- 6:54
this uh cold machine that's kind of
- 6:56
intimidating and scary and perhaps uh
- 7:00
somebody's going to say it's going to
- 7:00
solve all your problems, but you you
- 7:02
can't quite see the vision yet. So uh
- 7:04
perhaps that sounds familiar to any of
- 7:06
us working in agents today.
- 7:08
And Watt's solution was that he needed
- 7:10
to understand um the mental model of
- 7:12
these people and specifically to create
- 7:13
a metric that would help him uh give
- 7:16
some baseline of the relative
- 7:17
improvement in efficiency.
- 7:19
And so he literally studied uh horse
- 7:22
gins and tried to get some kind of
- 7:25
armchair measurements of of how is uh
- 7:28
like where the mechanics and the average
- 7:29
um performance of it and eventually came
- 7:32
to a metric called horsepower,
- 7:35
which may sound familiar.
- 7:37
And uh he used this measure to, you
- 7:39
know, this was to quantify the the
- 7:41
general power that the horses were were
- 7:43
um creating at the time and then he
- 7:45
could use that as a basis to show the
- 7:46
multiplier of efficiency that a steam
- 7:48
engine could provide.
- 7:49
And this metric was not very scientific
- 7:52
at the time. It was not necessarily even
- 7:54
accurate, you could say. But the main
- 7:57
thing it did was it communicated an
- 7:59
increase in value and so this led people
- 8:01
who love horses
- 8:03
let them kind of calibrate their
- 8:07
the the efficiency gains that they could
- 8:08
get by by attempting to adopt a steam
- 8:11
engine. So not even necessarily um
- 8:14
what you would get when you use it, but
- 8:16
what would get you over the limit of
- 8:17
trying it out in the first place.
- 8:20
And you know, it's it's it's a pretty
- 8:21
big feat because like although um
- 8:24
he had efficiency on his side with
- 8:26
regards to this metric
- 8:27
you know, let's be honest regardless of
- 8:29
how efficient this was
- 8:31
horses just have
- 8:33
great vibes. So like it's kind of hard
- 8:36
to beat the vibes of horses and so he
- 8:37
knew he had to kind of overcome the
- 8:39
emotion
- 8:40
and actually speak to to something that
- 8:42
gave them an ability to calculate the
- 8:44
ROI.
- 8:46
And oh, sorry. Skipped something.
- 8:48
And so yeah, the the lesson being if
- 8:50
we're not able to give something that is
- 8:52
a tangible ROI for our customers, then
- 8:55
it's very hard for us to communicate
- 8:57
value.
- 8:58
And I think we only need look to our own
- 9:00
industry to see all the examples where
- 9:03
other people in the technology sector
- 9:05
are failing to calculate good ROIs as
- 9:07
well.
- 9:08
And so we might in this room think this
- 9:09
is some sort of solved problem.
- 9:11
But if you look to the other engineers
- 9:13
in the world who are perhaps not as AI
- 9:15
pilled they're theoretically very smart
- 9:18
and should be able to figure out how to
- 9:20
calculate this pretty well, but then you
- 9:21
get these scenarios where people are
- 9:23
blowing through their entire
- 9:25
token budget for a year and they're
- 9:27
blowing through it in in a quarter or
- 9:29
they're like dealing with token
- 9:30
leaderboards and such and so obviously
- 9:33
the incentives haven't quite aligned and
- 9:34
we haven't perhaps got the right measure
- 9:36
of value in terms of the technology
- 9:38
sector itself. And so how then do we end
- 9:40
up scaling past past that and talk to
- 9:42
people who have no idea what we're
- 9:44
talking about, but still try to provide
- 9:46
them um a measure of the like increase
- 9:49
efficiency with agents.
- 9:51
And so right now I think we're kind of
- 9:52
in this doom loop where we're
- 9:54
we're overspending and we're underusing.
- 9:56
Uh this is a term I borrowed from Ramp
- 9:58
um and they have a great blog post on
- 9:59
this. And so it's kind of this vicious
- 10:01
cycle where we're just token maxing and
- 10:04
ourselves into austerity
- 10:06
kind of dropping out of the loop until
- 10:07
we get more FOMO to to get activated
- 10:10
enough to try it again.
- 10:12
And so I think we have to break this
- 10:13
loop and I think the way we do that is
- 10:14
by getting better measures of that will
- 10:16
communicate value.
- 10:19
Some people are obviously on the right
- 10:20
track. There was this chart floating
- 10:22
around on X recently that the Coinbase
- 10:24
Coinbase CEO posted where they hadn't
- 10:27
really started changing the defaults of
- 10:29
what models they will start with and
- 10:31
trying to only save the frontier models
- 10:33
for the hardest tasks and as a result
- 10:35
saw some good
- 10:36
saw AI spend start to diverge from token
- 10:39
usage.
- 10:40
And uh this is a good start. Uh Ramp
- 10:42
also, as I mentioned, has a great blog
- 10:43
post about this. Uh but I think the
- 10:44
problem is still that it's too focused
- 10:47
on tokens.
- 10:48
And tokens are of course uh useful as a
- 10:52
measurement of an internal system, but
- 10:54
at at the end of the day they're just an
- 10:56
output. And so the tokens need to then
- 10:59
be traced very cleanly to an outcome.
- 11:01
So how many
- 11:02
uh bug how many bugs did the tokens uh
- 11:05
how sorry how many um
- 11:07
bugs squashed did the tokens that we
- 11:09
bought um sorry, totally butchered that.
- 11:12
Um how many uh bugs got squashed with
- 11:14
the to with our token spend? How many uh
- 11:16
support requests got closed, etc. So
- 11:18
clean outcomes and then cleanly tying
- 11:20
those to
- 11:22
to progress on our objectives. And so
- 11:24
without uh a very tight measure of ROI,
- 11:26
this becomes very hard to do.
- 11:29
And I think I'll take this further and
- 11:31
um say that it need not even be the the
- 11:33
broader um, technology industry where
- 11:35
it's encountering this problem, but also
- 11:38
many of us in this room perhaps are.
- 11:40
And uh, although we're all probably
- 11:42
enjoying uh, coding with uh, with
- 11:44
various agents and feeling like it it's
- 11:46
it does feel like there's something
- 11:47
there in terms of the increase in
- 11:49
ability and efficiency.
- 11:51
Um, the problem of course is that we're
- 11:52
all kind of dying by a thousand pull
- 11:54
requests. And so, uh, even Anthropic who
- 11:57
has uh, some people on the team have
- 11:59
claimed have solved coding, uh, they
- 12:01
have also admitted that they've not
- 12:02
solved code review. And so, as a result,
- 12:05
um, the the bottleneck is now shifted to
- 12:08
the human review where the efficiency
- 12:10
gains from coding agents aren't quite
- 12:11
are aren't seen yet because we spend
- 12:13
most of the time reviewing the code and
- 12:15
we've not figured out how to scale that
- 12:16
in tandem with uh, the generation of the
- 12:18
code itself.
- 12:20
And so, the bottleneck ends up shifting
- 12:22
to the verification side and thus we
- 12:24
don't have a way to uh, measure value at
- 12:27
scale and uh, to judge quality at at the
- 12:29
same speed. And so, again going back to
- 12:31
the ROI calculations, we generate all
- 12:33
this code, but how do we know uh, we
- 12:36
don't know that enough of it is good to
- 12:37
justify the spend.
- 12:40
And of course, maybe uh, code review was
- 12:42
always flawed, uh, but it's just that
- 12:44
agents are now exposing it for
- 12:46
uh, the are exposing the actual problem.
- 12:49
Um, but I like this quote by Noah Hein
- 12:51
who from a a post about how to solve
- 12:53
code review where he's mentioning
- 12:54
specifically that the assumptions
- 12:56
underneath code review are what's now
- 12:57
being uh, what needs to be revisited.
- 12:59
So, you have to uh, check our priors to
- 13:02
try to figure out a new basis for um,
- 13:04
how we can code review in the age of
- 13:06
agents.
- 13:07
And I'm not going to go into how to
- 13:09
solve code review. I think that's
- 13:09
definitely better uh, a talk that's
- 13:11
better given by somebody else and um, is
- 13:14
totally different subject, but uh, what
- 13:16
I think is important for this talk is
- 13:17
why does code review feel like it is
- 13:19
solvable? And I think that Noah is
- 13:20
hitting on something important here
- 13:22
which is that as a as a culture uh, code
- 13:24
review has a a good uh,
- 13:26
convergence on shared assumptions and
- 13:28
that lets you
- 13:30
um
- 13:31
that lets you measure things at scale
- 13:33
when we can all kind of converge on the
- 13:34
measurement and it becomes somewhat of
- 13:38
clear rubric and so the task at hand now
- 13:41
is we have to adopt we have to adapt
- 13:44
those assumptions for the agentic age.
- 13:49
And so
- 13:50
we can we need to go if we're able to do
- 13:52
that then we can go from execution at
- 13:54
the speed of computer to measurement at
- 13:56
the speed of computer and of course the
- 13:58
measurements need to fit the mental
- 13:59
models of the customers using it.
- 14:01
And
- 14:03
I think the lesson here being that if
- 14:05
you're going to
- 14:06
think of how to build an agent for
- 14:08
something you also have to think about
- 14:10
how do you help the customers build
- 14:12
build or at least create a method for
- 14:14
verifying that the output is good and so
- 14:17
it's not enough to build it we also have
- 14:18
to help them
- 14:20
we also we also also have to help them
- 14:21
get to clear ROI calculations to justify
- 14:24
their spend.
- 14:26
And so this brings me to the idea of
- 14:29
mouse power which could be the
- 14:31
equivalent of horsepower for the agentic
- 14:33
age just as James Watt was able to show
- 14:35
a measure of efficiency relative to the
- 14:37
horses in in the gins in the horse gins
- 14:40
that were the source of power at the
- 14:41
time we perhaps can also figure out how
- 14:44
do we create a baseline of efficiency
- 14:46
for the way we use computers today and
- 14:49
can then demonstrate how much better or
- 14:51
perhaps more performant on certain
- 14:53
vectors an agent could be at that task.
- 14:56
Um and
- 14:57
of course it's not as easy perhaps as
- 14:59
easy a task as he had back then where he
- 15:01
could just study the horse gin cuz it's
- 15:03
not as if we can create some method to
- 15:05
measure cursor movements and like figure
- 15:07
out the delta of how much more efficient
- 15:09
an agent could move them and thus we can
- 15:11
say yeah agents are this much more
- 15:13
performant than humans at these tasks.
- 15:15
Trust me I've I've tried I had Claude
- 15:18
vibe code me this measurement device and
- 15:20
I thought maybe if I can figure out the
- 15:22
movement like the potential movement
- 15:24
across the screen and measure how fast
- 15:26
it went, I could get some clean measure
- 15:28
of mouse power. Uh but of course this is
- 15:30
only joking. Um this is of course um
- 15:33
like a fool's errand because information
- 15:35
space is just way too high dimensional
- 15:37
and so I think mouse power is is never
- 15:40
going to be a metric of course, but it's
- 15:41
more so an idea. Which the idea being if
- 15:44
you're going to sell somebody an agent,
- 15:45
you also have to help them with the with
- 15:47
the rubric of how do we actually verify
- 15:50
that this agent is doing good work and
- 15:52
thus we can uh have a good measure of
- 15:54
saying that these tokens are worth it.
- 15:56
Um so how to do that of course is is
- 15:58
really up to you and I won't be able to
- 16:00
tell you how do you I don't have any
- 16:02
good frameworks for how do you figure
- 16:04
out the right measurements to to help
- 16:05
provide anybody you're building an agent
- 16:07
for. Uh but what I can do is give a
- 16:09
principle uh give an idea that I've been
- 16:11
kicking around which is based um
- 16:13
in information theory. So going back to
- 16:16
Claude Shannon's ideas about measuring
- 16:18
entropy and information.
- 16:20
Uh entropy being uh the uncertainty of a
- 16:22
probability distribution and of course
- 16:25
very much the basis of how we train
- 16:27
agents today.
- 16:28
Things like cross entropy and such uh
- 16:30
being a big factor in determining how
- 16:31
capable an agent is.
- 16:33
Um I think that entropy's an interesting
- 16:35
idea to think through with regards to
- 16:36
not just the performance of an agent,
- 16:39
but also the task that we're setting
- 16:40
them out to to perform on.
- 16:42
And so uh I put together this matrix
- 16:45
which uh it maps on the x-axis axis the
- 16:49
uncertainty in the steps it takes to
- 16:50
perform a task. And so when we're
- 16:52
thinking of building an agent, I think
- 16:54
it's not enough to just think what would
- 16:56
be a valuable task for the agent to do,
- 16:58
but also thinking about how um how much
- 17:01
uncertainty are in the steps to perform
- 17:02
that task itself. So an example would be
- 17:05
uh booking a flight has uh much less
- 17:08
uncertainty than let's say painting a
- 17:09
masterpiece, right? Because you know
- 17:11
there's certain information that has to
- 17:13
be that has to happen in the flight
- 17:14
purchase. There has to be a departing
- 17:16
destination, arriving destination.
- 17:18
There's going to be a seat chosen. It
- 17:19
might be by the person. It might just be
- 17:21
random.
- 17:22
But these things have to happen for that
- 17:24
task to be completed. And on the other
- 17:25
hand, there is the task of like painting
- 17:28
a masterpiece, right? And who knows what
- 17:30
the steps are to that? Maybe you can get
- 17:32
an agent to do it, but it would be very
- 17:34
hard to figure out how we can actually
- 17:36
create a a relatively predictable
- 17:39
pathway to that.
- 17:40
But then on the other axis is the the
- 17:43
uncertainty in the acceptance criteria
- 17:45
itself. So not just can the agent
- 17:47
perform the task, but can we help
- 17:49
somebody actually or is is there
- 17:51
actually a a clean rubric for how it's
- 17:53
graded? And so thinking about ideas on
- 17:56
on these two axes and where they
- 17:57
intersect, perhaps gives us a better
- 17:59
guide for how to build agents and we can
- 18:01
run through a few examples. So if we
- 18:05
look at the at the left side, your right
- 18:07
side.
- 18:09
Yes.
- 18:10
No, you're left as well.
- 18:12
Then
- 18:13
No, last speaker was also a confused
- 18:15
about.
- 18:16
So uh
- 18:17
Yeah, on the left side when uncertainty
- 18:20
in the task steps are low, then it's a
- 18:22
very it's a very predictable outcome or
- 18:25
it's a very predictable pathway to
- 18:26
achieve that goal. And so then, you
- 18:28
know, why would you waste tokens? Just
- 18:30
write a script. On the other side, when
- 18:33
the
- 18:34
the steps to do perform the task are
- 18:36
very high in in uncertainty, then you
- 18:39
you have very unpredictable information.
- 18:40
And so it's probably at risk of being
- 18:43
out of distribution in pre-training and
- 18:45
probably has very sparse rewards for
- 18:46
reinforcement learning. And so perhaps
- 18:48
it's not a a good task for an agent
- 18:50
because it's just much harder to figure
- 18:52
out how to actually model that data.
- 18:54
And so obviously in the middle is is um
- 18:58
is so I'm I think I'm out of time, but
- 19:00
I'm not getting kicked off yet.
- 19:02
So I'll just finish this up quickly. Um
- 19:04
so yeah, in the middle is is probably
- 19:05
the sweet spot, but then on the other
- 19:06
axis,
- 19:07
what's the uncertainty in verifying that
- 19:09
this is actually valuable? So when you
- 19:11
have high uncertainty in the acceptance
- 19:13
criteria, you pretty much are in a spot
- 19:14
where verification is indistinguishable
- 19:16
from execution. So, why would you build
- 19:18
an agent for something that to verify
- 19:21
was useful, a person pretty much has to
- 19:22
do the work again. So, like waste of
- 19:25
tokens obviously.
- 19:27
And then it leaves that that middle area
- 19:29
where you have this interesting
- 19:30
intersection of tasks that are um
- 19:33
they're not too uncertain in that
- 19:36
they or they they have a degree of
- 19:38
uncertainty where they're not great
- 19:41
they're not just a a script or they're
- 19:43
not out of distribution for training,
- 19:45
but they have enough uncertainty to be
- 19:46
interesting, but at the at the same time
- 19:49
they also have a property of being
- 19:50
relatively easy to
- 19:52
validate the
- 19:54
to check the value of them. And so they
- 19:56
become in this place where they kind of
- 19:58
become the shape of an NP-style problem,
- 20:00
which means they're easier to verify
- 20:01
than to execute. And the reason I say
- 20:03
that is because if you can figure out a
- 20:05
pretty repeatable pattern for verifying
- 20:07
their their work, you can actually just
- 20:09
throw agents to that problem as well.
- 20:11
And so of course you don't just build
- 20:13
the agent, you perhaps build the agent
- 20:15
that verifies the work of the agent.
- 20:17
Um and so yeah, this is perhaps this is
- 20:21
a thought starter mostly kind of have
- 20:23
kind of still in the works, so
- 20:25
I'm happy to hear any thoughts on it,
- 20:26
but if with this guidance I hope when
- 20:29
you're building your next agent you can
- 20:30
also figure out how to also build its
- 20:32
mouse power.
- 20:34
And thanks very much.