Mousepower: agents that can’t be measured, can’t be managed. — Maximillian Piras, Yutori

Read the talk

Mousepower: Agents That Can’t Be Measured Can’t Be Managed

Maximillian Piras explains why token counts fail to show whether agents create value, why faster generation moves the bottleneck to verification, and how task uncertainty can help identify work worth delegating.

From a talk by Maximillian Piras

At a glance

Ideas worth remembering

  • Token volume measures system activity, not customer value. Trace spend to accepted outcomes such as bugs closed or support requests resolved, then connect those outcomes to an objective.

  • Faster generation moves the bottleneck to verification. Agent output creates value only after its quality can be judged and the work can be accepted or deployed.

  • Mousepower is not a literal cursor-speed metric. It is the requirement to give customers an understandable rubric for deciding whether agent work was worth its cost.

  • Use deterministic scripts when task steps are predictable, and avoid delegation when checking requires doing the work again. The best agent candidates combine moderate execution uncertainty with relatively clear acceptance criteria.

  • When verification follows a repeatable pattern, an agent can help check another agent—but the checking rubric must still reflect the customer’s definition of good work.

Parallel agents are fun until the bill arrives

Maximillian Piras, founding designer at Yutori, starts with a familiar agent workflow: keep human attention on the main task while several agents pursue peripheral work in the background. For this talk, the concrete job is improving the slides. A design system and written guidance constrain the work; multiple agents then explore typography and layout in parallel. The observable result is more design variants produced without taking over his active attention. 0:42

Recording frame at 130 seconds
Recording frame at 130 seconds

Parallelism feels like free leverage because research and design exploration rarely have an obvious stopping point. But each additional run consumes tokens, and the bill eventually forces a question the workflow postponed: which explorations were useful enough to justify their cost? A pile of generated variants is visible output. It is not yet evidence of value. 1:42

That question is especially relevant to Yutori’s computer-use models. These models operate software as a person would when an API or MCP connection is unavailable. They can reach information trapped behind ordinary interfaces, but clicking through a rendered application is less efficient than calling a structured API. Computer use is therefore a fallback for the long tail of inaccessible systems, not the preferred interface when a direct integration exists. 2:42

Customer conversations repeatedly return to the same unresolved decisions: which jobs suit an agent, what tradeoff to make between token cost and value, and how to tell whether the result is good. Excitement is widespread, but so is the feeling of merely “scratching the surface.” The missing muscle is measurement. 3:42

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Horsepower made an unfamiliar machine legible

Experienced agent users can object that they already know how to run a fleet. Piras accepts that experience while questioning whether it generalizes. Early adopters enjoy experimenting, tolerate rough interfaces, and may willingly spend tokens to discover what works. People who have never used an agent—and may still work through copy-and-paste interactions—do not share that calibration. A product cannot explain its value to the wider market by pointing to the habits of its most enthusiastic users. 4:13

Recording frame at 374 seconds
Recording frame at 374 seconds

The historical analogy is James Watt selling steam engines to people who understood power through horses. A horse gin connected an animal to a rotary arm; walking in a circle supplied mechanical power to a mill. Replacing that familiar arrangement with an intimidating machine required more than claiming that steam was better. Buyers needed a comparison expressed in terms they already understood. 5:43

Watt studied horse gins and developed horsepower as a baseline for expressing relative performance. In Piras’s telling, the early figure was neither especially scientific nor necessarily accurate. Its commercial utility came from legibility: a buyer could translate machine output into a familiar unit and estimate the gain from trying the new technology. The measure reduced the cognitive distance between the old system and the new one. 7:12

The comparison also had to overcome attachment to the incumbent technology—horses, as Piras puts it, had “great vibes.” Efficiency alone does not create adoption when the new system feels unfamiliar or threatening. A comprehensible return on investment gives someone a reason to cross the threshold from curiosity to use. 8:12

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

4:13 · section reference included

Tokens are cost accounting, not value accounting

The agent industry has not found an equivalent translation layer. Organizations can exhaust an annual token budget in a quarter or turn consumption into a leaderboard, but neither behavior establishes that the spending improved the business. When incentives reward usage rather than useful work, teams can become very good at consuming the resource they meant to evaluate. 9:12

Recording frame at 676 seconds
Recording frame at 676 seconds

Borrowing a term from Ramp, Piras describes the resulting pattern as “overspending and underusing”: teams maximize agent activity, hit austerity, retreat from the tools, and return when fear of missing out becomes strong enough. What does this cycle make visible? Enthusiasm without outcome measurement does not produce sustained adoption; it produces alternating bursts of consumption and withdrawal. 9:42

The diagram shows the missing link. Spending produces activity, but without an outcome measure the team cannot distinguish productive use from waste. Budget pressure therefore causes a broad retreat instead of a targeted adjustment to models, tasks, or workflows.

Choosing cheaper models by default and reserving frontier models for difficult work can separate token volume from total spend. That is useful cost control, but it still treats tokens as the central unit. Tokens are an internal output of the system. The business cares about what changed afterward. 10:12

The practical accounting chain should run from spend to a concrete outcome and then to an objective. For software work, ask how many bugs were closed with the purchased tokens. For support, ask how many requests were resolved. Return to the slide example: ten agents may create ten alternatives, but value appears only when review identifies a better treatment and that treatment improves the finished deck. Counting variants or tokens stops before the consequential step. 10:42

How it fits togetherThe overspending-and-underusing loop

Run more agents and consume more tokens in pursuit of additional work.

Without outcome-based measurement, budget pressure interrupts adoption instead of teaching the team which agent work is valuable.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:12 · section reference included

Generation got faster; review did not

Coding agents expose the deeper problem. They can generate changes quickly enough to leave teams “dying by a thousand pull requests,” while code review still consumes human attention. More generated code does not automatically create more accepted, deployed, useful software. When generation accelerates but review does not, the bottleneck moves downstream. 11:42

Recording frame at 818 seconds
Recording frame at 818 seconds

This creates a measurement problem as well as a workflow problem. A team can observe code volume immediately, but it cannot judge quality at the same speed. Until someone reviews the change and decides whether it is good enough to use, token spend has produced a proposal rather than established value. The original bottleneck may have been implementation; the new one is verification. 12:12

Code review nevertheless offers a useful clue because software teams have converged on shared assumptions about what acceptable work looks like. That convergence creates a rubric, even if its assumptions need revision for agent-generated code. Once people agree on what should be checked, parts of the judgment can happen consistently and at greater scale. 13:12

The product requirement follows: building an agent also means helping its user verify the output and connect accepted work to return on investment. Execution at the speed of compute is only half the system. Useful deployment needs measurement that can approach the same speed without pretending that generation itself proves success. 13:49

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

11:42 · section reference included

Mousepower is a design requirement, not a unit

Mousepower is Piras’s proposed analogy for the agent era: establish a baseline for how people use computers, then express where an agent performs better on the task. The tempting implementation is literal—measure cursor distance or speed and compare human movement with agent movement. Piras even had a prototype measurement device generated for the joke. 14:19

The literal metric fails because computer work takes place in a high-dimensional information space. Faster pointer movement says little about whether the correct source was selected, the right judgment was applied, or the resulting action advanced the user’s goal. Two agents can traverse the same interface at similar speed and produce outcomes of radically different value. 15:19

Mousepower therefore names a product obligation rather than a standardized unit. If a product sells agent execution, it should also supply a rubric for determining whether that execution was good and whether its token cost was justified. Piras does not offer a universal formula for that rubric; it must fit the customer’s task and mental model. 15:49

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

14:19 · section reference included

Choose tasks by two kinds of uncertainty

To narrow the search for suitable tasks, Piras borrows the idea of entropy as uncertainty. The proposed matrix has two dimensions: uncertainty in the steps required to complete a task, and uncertainty in the criteria used to accept the result. This is explicitly a thought starter still in development, not a validated task-selection model or a quantified use of information-theoretic entropy. 16:19

Recording frame at 997 seconds
Recording frame at 997 seconds

The first axis asks how predictable the route is. Booking a flight has required information and recognizable stages: departure, destination, and a seat selection of some kind. Painting a masterpiece has no comparable sequence that reliably yields success. The contrast concerns path structure, not whether either task can sometimes be completed by a model. 16:49

The second axis asks whether success has a clean grading rubric. An agent might be able to perform an open-ended task, yet still be a poor product if a person must redo the work to determine whether it was valuable. Verification cost belongs in the task-selection decision from the beginning. 17:49

What does the matrix help a builder decide? It places uncertainty in execution on the horizontal axis and uncertainty in acceptance on the vertical axis. Reading across distinguishes deterministic scripts from useful agent work and poorly modeled tasks. Reading upward shows how rising verification uncertainty makes delegation less economical. The promising region is moderate step uncertainty paired with relatively clear acceptance criteria.

The regions imply four practical decisions:

  • Predictable steps — write a script. Deterministic automation avoids spending tokens on a path already known in advance.
  • Extremely uncertain steps — reconsider the task. The work may fall outside useful training distribution and provide sparse reward signals.
  • Unclear acceptance — avoid false delegation. If a person must effectively repeat the task to judge it, little work has been saved.
  • Moderately uncertain steps with clear acceptance — consider an agent. The task needs judgment, but its result remains economical to check.

Piras informally compares the final region to an NP-style problem: solving is harder than checking. The important engineering property is the asymmetry, not a formal complexity classification. When acceptance follows a repeatable pattern, another agent can perform part of the verification. The resulting system contains an execution agent and a checking agent, both working against criteria the customer understands. That is how an agent begins to acquire its mousepower. 19:49

Compare the ideasThe two-axis task uncertainty matrix

The path and acceptance criteria are predictable, so deterministic automation is cheaper.

Horizontal position represents uncertainty in the execution steps; vertical grouping represents uncertainty in acceptance criteria. The best agent candidates sit in the low-acceptance-uncertainty row and the moderate-step-uncertainty column.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

16:19 · section reference included

Resources

Read the complete timestamped transcript
  1. 0:01

    [music]

  2. 0:12

    >> Um well, thanks all for your time.

  3. 0:14

    Really appreciate you dropping by and

  4. 0:16

    it's always a a great honor to speak at

  5. 0:18

    the world's fair. So, I'll do my best to

  6. 0:21

    uh

  7. 0:21

    give you guys some valuable insights and

  8. 0:23

    um

  9. 0:24

    yeah, hopefully make it worth your time.

  10. 0:26

    So, my name's Maximilian Piros and today

  11. 0:28

    I'll be talking about mouse power.

  12. 0:31

    And this is a talk about measuring

  13. 0:33

    agents through mental models.

  14. 0:35

    But before I get into talking about

  15. 0:37

    measuring agents, I'm going to talk

  16. 0:39

    through a bit about how I use them every

  17. 0:40

    day. And it might seem familiar to you,

  18. 0:43

    but just to level set, we'll go through

  19. 0:44

    it. So, uh I tend to background them

  20. 0:46

    like I'm sure a lot of you people are as

  21. 0:48

    well.

  22. 0:49

    Um so, while my active attention is

  23. 0:51

    focusing on one thing, like perhaps

  24. 0:53

    giving this talk to you, I still want to

  25. 0:55

    make some progress on peripheral tasks.

  26. 0:57

    So, I'll keep my attention focused on

  27. 1:00

    giving this talk, while my agents can

  28. 1:03

    help me explore some designs in the

  29. 1:04

    background cuz I think that uh my slides

  30. 1:07

    need a bit of work. So, you know, I've

  31. 1:08

    got my design system already set up.

  32. 1:10

    I've got some guidance uh given to my

  33. 1:12

    agents. So, I'll kick off an agent to

  34. 1:14

    just try to explore some different

  35. 1:16

    directions on the type type two

  36. 1:17

    treatment and the layout. And um you

  37. 1:20

    know, just try to get as many

  38. 1:21

    explorations as possible.

  39. 1:23

    But of course, uh one agent's never

  40. 1:24

    enough. So, I like to kick off a bunch

  41. 1:26

    in parallel. You know, I've got a lot of

  42. 1:28

    slides to get through. So, I need all of

  43. 1:30

    my agents on and exploring it in

  44. 1:32

    different directions and

  45. 1:34

    uh hopefully I can get some interesting

  46. 1:35

    things to make my slides a bit better.

  47. 1:38

    And uh hopefully they can finish the job

  48. 1:40

    pretty soon cuz we're obviously kind of

  49. 1:42

    up against uh the timeline the the uh

  50. 1:44

    deadline here.

  51. 1:46

    So, um

  52. 1:47

    this is generally how I work. I'm sure

  53. 1:48

    it's probably familiar to a lot of you

  54. 1:50

    where we're just trying to kick off

  55. 1:51

    agents for as much as possible in

  56. 1:53

    parallel cuz it always feels like

  57. 1:55

    there's just way more research to do. We

  58. 1:57

    want it to be as thorough as possible.

  59. 1:59

    There's way more design explorations to

  60. 2:01

    do. So, whenever our main focus is on

  61. 2:03

    one thing, why not kick a bunch of

  62. 2:05

    agents off in parallel and just try to

  63. 2:07

    maximize your time. And it's a lot of

  64. 2:10

    fun, of course, until you get the bill.

  65. 2:13

    And then you start to wonder, was it all

  66. 2:15

    worth it? Right? Did

  67. 2:17

    did you vibe code too hard?

  68. 2:20

    Were you token maxing too much? Like,

  69. 2:22

    could you have been more efficient in

  70. 2:24

    how you approached

  71. 2:26

    your sequencing your agents?

  72. 2:28

    And so, this is what I'm going to get

  73. 2:29

    into today. It's It's how do we evaluate

  74. 2:31

    the token cost? And specifically, how do

  75. 2:33

    we help our customers value it?

  76. 2:36

    So, for the past year and a half, I've

  77. 2:38

    had the pleasure of working as a

  78. 2:39

    founding designer at a company called

  79. 2:40

    території and we focus on computer use

  80. 2:43

    models. These are models that learn to

  81. 2:45

    use a computer like a human would and

  82. 2:47

    the use case for them is when you can't

  83. 2:49

    get information from an API or an MCP,

  84. 2:52

    why not just send an agent out to use

  85. 2:54

    the computer like a human would and then

  86. 2:56

    we can extract all types of data and

  87. 2:58

    manipulate it in ways that let us access

  88. 3:01

    all the stuff that wasn't accessible

  89. 3:03

    previously. So, obviously less efficient

  90. 3:04

    than APIs and MCPs, but as a last

  91. 3:07

    resort, just have an agent go use the

  92. 3:09

    computer and try to get the information.

  93. 3:12

    Here's the U території agent using the U

  94. 3:14

    території website.

  95. 3:16

    Um checking out its own benchmark. So,

  96. 3:18

    kind of it's admire itself in a way. So,

  97. 3:21

    yeah, it gets it gets a bit weird. Like,

  98. 3:23

    and a lot of what I do as a founding

  99. 3:25

    designer there is talk to customers, try

  100. 3:27

    to understand how can we make agents as

  101. 3:29

    intuitive as possible, how do we figure

  102. 3:31

    out the mental models they're using to

  103. 3:33

    value the use cases they want to send

  104. 3:35

    out agents for.

  105. 3:36

    And a lot of them do seem pretty

  106. 3:39

    confused so far. A lot of people are

  107. 3:41

    excited about agents, but the phrase

  108. 3:43

    that comes up quite often is that they

  109. 3:45

    feel like they're just scratching the

  110. 3:46

    surface. Um

  111. 3:47

    it seems like it's not quite intuitive

  112. 3:48

    how we can best use them yet. And so, in

  113. 3:51

    a lot of my uh customer discussions,

  114. 3:53

    it's it always comes down to a question

  115. 3:55

    of like, what is the best way to use use

  116. 3:57

    uh to use agents? What are the best use

  117. 3:58

    cases for them? And how do I think about

  118. 4:00

    the trade-offs with regards to token

  119. 4:02

    costs relative to value? So, I think

  120. 4:04

    we're still kind of building this muscle

  121. 4:06

    today. And this leads me to the thesis

  122. 4:08

    of the talk, which is that I think

  123. 4:10

    agents have a measurement problem.

  124. 4:13

    And uh as an example, here's me at work

  125. 4:15

    trying to measure some agents, and one

  126. 4:17

    of my co-workers took this photo and

  127. 4:19

    told me I looked like

  128. 4:20

    I was uh trying to solve the mystery of

  129. 4:22

    Pepe Silvia. So,

  130. 4:23

    uh

  131. 4:24

    as you can see, it's it's

  132. 4:26

    it's not a it's not an easy task to to

  133. 4:28

    measure agents. But, I'm sure some of

  134. 4:30

    you are saying, "Hold on a sec. Like,

  135. 4:32

    what is this guy talking about? I've got

  136. 4:34

    a fleet of agents working for me right

  137. 4:35

    now. We're building our next

  138. 4:37

    million-dollar app as we speak, and I'm

  139. 4:39

    having a a totally fine time uh

  140. 4:41

    measuring my agents." Uh to which I will

  141. 4:43

    agree with you, uh but then I will point

  142. 4:45

    you to the mandatory Upton Sinclair

  143. 4:48

    quote to remind us all that everybody in

  144. 4:50

    this room is very biased, and we're

  145. 4:52

    early adopters, and we're very excited

  146. 4:55

    to explore this new technology, but it

  147. 4:57

    doesn't mean that we represent the

  148. 4:58

    people that ultimately we're going to be

  149. 5:00

    trying to help adopt this technology.

  150. 5:02

    And so, you know, I think it's important

  151. 5:04

    to remind ourselves that in some way or

  152. 5:06

    another, we probably are selling tokens,

  153. 5:08

    whether it's indirectly or directly. And

  154. 5:11

    so, when we think about our own token

  155. 5:12

    usage, uh is it really representative of

  156. 5:14

    all the people out there who have never

  157. 5:15

    touched an agent yet? Uh some people are

  158. 5:18

    still copy and pasting into ChatGPT.

  159. 5:20

    I may be [REDACTED] to one of these people,

  160. 5:22

    and despite how much I tried to get her

  161. 5:24

    to try out agents, she's not let me uh

  162. 5:26

    set her up with it yet. And so, um as a

  163. 5:29

    reminder, uh when we think about helping

  164. 5:31

    people adopt agents, you know, all the

  165. 5:32

    people across the world that we think

  166. 5:34

    could get as much uh excitement and

  167. 5:36

    value as as we do when we run off

  168. 5:37

    parallel agents, let's uh just remember

  169. 5:39

    this quote.

  170. 5:40

    >> [snorts]

  171. 5:41

    >> And um

  172. 5:42

    so, it really boils down to the age-old

  173. 5:44

    problem of a new technology.

  174. 5:47

    And of course, there's tons of history

  175. 5:49

    we can go to to study how people solved

  176. 5:51

    this in the past. We have this really

  177. 5:53

    exciting new thing, but we haven't quite

  178. 5:55

    uh figured out the right ways to

  179. 5:56

    communicate it.

  180. 5:57

    And so for this talk, I'll go back to

  181. 5:59

    the 1700s and we can take some notes

  182. 6:01

    from when James Watt was trying to sell

  183. 6:04

    steam engines.

  184. 6:06

    And at the time he decided that uh a

  185. 6:08

    great use case for his steam engines was

  186. 6:09

    trying to replace a horse gin. And these

  187. 6:13

    are the was the power source of a mill

  188. 6:15

    at the time. So when you're

  189. 6:16

    for uh let's say a brewery and you you

  190. 6:18

    need some power source to to grind your

  191. 6:21

    barley or whatever. I don't know. I'm

  192. 6:22

    not I'm not like a big brewery guy, so I

  193. 6:23

    don't know exactly what how it's made,

  194. 6:25

    but you need a power source and the

  195. 6:27

    power source at the time that was common

  196. 6:29

    was you hooked a horse up to a rotary

  197. 6:30

    arm and the horse walked in a circle and

  198. 6:32

    that's how you generated your power.

  199. 6:34

    Uh and seems crazy today maybe, but um

  200. 6:37

    at the time was commonplace and Watt

  201. 6:38

    thought, you know, it would be much

  202. 6:40

    better than a horse is like a very

  203. 6:41

    efficient machine.

  204. 6:43

    Although he um rightfully acknowledged

  205. 6:45

    that one of the big barriers to adopting

  206. 6:47

    it would be this cognitive dissonance of

  207. 6:50

    trying to tell people who kind of think

  208. 6:52

    in horses, how do you adapt to this to

  209. 6:54

    this uh cold machine that's kind of

  210. 6:56

    intimidating and scary and perhaps uh

  211. 7:00

    somebody's going to say it's going to

  212. 7:00

    solve all your problems, but you you

  213. 7:02

    can't quite see the vision yet. So uh

  214. 7:04

    perhaps that sounds familiar to any of

  215. 7:06

    us working in agents today.

  216. 7:08

    And Watt's solution was that he needed

  217. 7:10

    to understand um the mental model of

  218. 7:12

    these people and specifically to create

  219. 7:13

    a metric that would help him uh give

  220. 7:16

    some baseline of the relative

  221. 7:17

    improvement in efficiency.

  222. 7:19

    And so he literally studied uh horse

  223. 7:22

    gins and tried to get some kind of

  224. 7:25

    armchair measurements of of how is uh

  225. 7:28

    like where the mechanics and the average

  226. 7:29

    um performance of it and eventually came

  227. 7:32

    to a metric called horsepower,

  228. 7:35

    which may sound familiar.

  229. 7:37

    And uh he used this measure to, you

  230. 7:39

    know, this was to quantify the the

  231. 7:41

    general power that the horses were were

  232. 7:43

    um creating at the time and then he

  233. 7:45

    could use that as a basis to show the

  234. 7:46

    multiplier of efficiency that a steam

  235. 7:48

    engine could provide.

  236. 7:49

    And this metric was not very scientific

  237. 7:52

    at the time. It was not necessarily even

  238. 7:54

    accurate, you could say. But the main

  239. 7:57

    thing it did was it communicated an

  240. 7:59

    increase in value and so this led people

  241. 8:01

    who love horses

  242. 8:03

    let them kind of calibrate their

  243. 8:07

    the the efficiency gains that they could

  244. 8:08

    get by by attempting to adopt a steam

  245. 8:11

    engine. So not even necessarily um

  246. 8:14

    what you would get when you use it, but

  247. 8:16

    what would get you over the limit of

  248. 8:17

    trying it out in the first place.

  249. 8:20

    And you know, it's it's it's a pretty

  250. 8:21

    big feat because like although um

  251. 8:24

    he had efficiency on his side with

  252. 8:26

    regards to this metric

  253. 8:27

    you know, let's be honest regardless of

  254. 8:29

    how efficient this was

  255. 8:31

    horses just have

  256. 8:33

    great vibes. So like it's kind of hard

  257. 8:36

    to beat the vibes of horses and so he

  258. 8:37

    knew he had to kind of overcome the

  259. 8:39

    emotion

  260. 8:40

    and actually speak to to something that

  261. 8:42

    gave them an ability to calculate the

  262. 8:44

    ROI.

  263. 8:46

    And oh, sorry. Skipped something.

  264. 8:48

    And so yeah, the the lesson being if

  265. 8:50

    we're not able to give something that is

  266. 8:52

    a tangible ROI for our customers, then

  267. 8:55

    it's very hard for us to communicate

  268. 8:57

    value.

  269. 8:58

    And I think we only need look to our own

  270. 9:00

    industry to see all the examples where

  271. 9:03

    other people in the technology sector

  272. 9:05

    are failing to calculate good ROIs as

  273. 9:07

    well.

  274. 9:08

    And so we might in this room think this

  275. 9:09

    is some sort of solved problem.

  276. 9:11

    But if you look to the other engineers

  277. 9:13

    in the world who are perhaps not as AI

  278. 9:15

    pilled they're theoretically very smart

  279. 9:18

    and should be able to figure out how to

  280. 9:20

    calculate this pretty well, but then you

  281. 9:21

    get these scenarios where people are

  282. 9:23

    blowing through their entire

  283. 9:25

    token budget for a year and they're

  284. 9:27

    blowing through it in in a quarter or

  285. 9:29

    they're like dealing with token

  286. 9:30

    leaderboards and such and so obviously

  287. 9:33

    the incentives haven't quite aligned and

  288. 9:34

    we haven't perhaps got the right measure

  289. 9:36

    of value in terms of the technology

  290. 9:38

    sector itself. And so how then do we end

  291. 9:40

    up scaling past past that and talk to

  292. 9:42

    people who have no idea what we're

  293. 9:44

    talking about, but still try to provide

  294. 9:46

    them um a measure of the like increase

  295. 9:49

    efficiency with agents.

  296. 9:51

    And so right now I think we're kind of

  297. 9:52

    in this doom loop where we're

  298. 9:54

    we're overspending and we're underusing.

  299. 9:56

    Uh this is a term I borrowed from Ramp

  300. 9:58

    um and they have a great blog post on

  301. 9:59

    this. And so it's kind of this vicious

  302. 10:01

    cycle where we're just token maxing and

  303. 10:04

    ourselves into austerity

  304. 10:06

    kind of dropping out of the loop until

  305. 10:07

    we get more FOMO to to get activated

  306. 10:10

    enough to try it again.

  307. 10:12

    And so I think we have to break this

  308. 10:13

    loop and I think the way we do that is

  309. 10:14

    by getting better measures of that will

  310. 10:16

    communicate value.

  311. 10:19

    Some people are obviously on the right

  312. 10:20

    track. There was this chart floating

  313. 10:22

    around on X recently that the Coinbase

  314. 10:24

    Coinbase CEO posted where they hadn't

  315. 10:27

    really started changing the defaults of

  316. 10:29

    what models they will start with and

  317. 10:31

    trying to only save the frontier models

  318. 10:33

    for the hardest tasks and as a result

  319. 10:35

    saw some good

  320. 10:36

    saw AI spend start to diverge from token

  321. 10:39

    usage.

  322. 10:40

    And uh this is a good start. Uh Ramp

  323. 10:42

    also, as I mentioned, has a great blog

  324. 10:43

    post about this. Uh but I think the

  325. 10:44

    problem is still that it's too focused

  326. 10:47

    on tokens.

  327. 10:48

    And tokens are of course uh useful as a

  328. 10:52

    measurement of an internal system, but

  329. 10:54

    at at the end of the day they're just an

  330. 10:56

    output. And so the tokens need to then

  331. 10:59

    be traced very cleanly to an outcome.

  332. 11:01

    So how many

  333. 11:02

    uh bug how many bugs did the tokens uh

  334. 11:05

    how sorry how many um

  335. 11:07

    bugs squashed did the tokens that we

  336. 11:09

    bought um sorry, totally butchered that.

  337. 11:12

    Um how many uh bugs got squashed with

  338. 11:14

    the to with our token spend? How many uh

  339. 11:16

    support requests got closed, etc. So

  340. 11:18

    clean outcomes and then cleanly tying

  341. 11:20

    those to

  342. 11:22

    to progress on our objectives. And so

  343. 11:24

    without uh a very tight measure of ROI,

  344. 11:26

    this becomes very hard to do.

  345. 11:29

    And I think I'll take this further and

  346. 11:31

    um say that it need not even be the the

  347. 11:33

    broader um, technology industry where

  348. 11:35

    it's encountering this problem, but also

  349. 11:38

    many of us in this room perhaps are.

  350. 11:40

    And uh, although we're all probably

  351. 11:42

    enjoying uh, coding with uh, with

  352. 11:44

    various agents and feeling like it it's

  353. 11:46

    it does feel like there's something

  354. 11:47

    there in terms of the increase in

  355. 11:49

    ability and efficiency.

  356. 11:51

    Um, the problem of course is that we're

  357. 11:52

    all kind of dying by a thousand pull

  358. 11:54

    requests. And so, uh, even Anthropic who

  359. 11:57

    has uh, some people on the team have

  360. 11:59

    claimed have solved coding, uh, they

  361. 12:01

    have also admitted that they've not

  362. 12:02

    solved code review. And so, as a result,

  363. 12:05

    um, the the bottleneck is now shifted to

  364. 12:08

    the human review where the efficiency

  365. 12:10

    gains from coding agents aren't quite

  366. 12:11

    are aren't seen yet because we spend

  367. 12:13

    most of the time reviewing the code and

  368. 12:15

    we've not figured out how to scale that

  369. 12:16

    in tandem with uh, the generation of the

  370. 12:18

    code itself.

  371. 12:20

    And so, the bottleneck ends up shifting

  372. 12:22

    to the verification side and thus we

  373. 12:24

    don't have a way to uh, measure value at

  374. 12:27

    scale and uh, to judge quality at at the

  375. 12:29

    same speed. And so, again going back to

  376. 12:31

    the ROI calculations, we generate all

  377. 12:33

    this code, but how do we know uh, we

  378. 12:36

    don't know that enough of it is good to

  379. 12:37

    justify the spend.

  380. 12:40

    And of course, maybe uh, code review was

  381. 12:42

    always flawed, uh, but it's just that

  382. 12:44

    agents are now exposing it for

  383. 12:46

    uh, the are exposing the actual problem.

  384. 12:49

    Um, but I like this quote by Noah Hein

  385. 12:51

    who from a a post about how to solve

  386. 12:53

    code review where he's mentioning

  387. 12:54

    specifically that the assumptions

  388. 12:56

    underneath code review are what's now

  389. 12:57

    being uh, what needs to be revisited.

  390. 12:59

    So, you have to uh, check our priors to

  391. 13:02

    try to figure out a new basis for um,

  392. 13:04

    how we can code review in the age of

  393. 13:06

    agents.

  394. 13:07

    And I'm not going to go into how to

  395. 13:09

    solve code review. I think that's

  396. 13:09

    definitely better uh, a talk that's

  397. 13:11

    better given by somebody else and um, is

  398. 13:14

    totally different subject, but uh, what

  399. 13:16

    I think is important for this talk is

  400. 13:17

    why does code review feel like it is

  401. 13:19

    solvable? And I think that Noah is

  402. 13:20

    hitting on something important here

  403. 13:22

    which is that as a as a culture uh, code

  404. 13:24

    review has a a good uh,

  405. 13:26

    convergence on shared assumptions and

  406. 13:28

    that lets you

  407. 13:30

    um

  408. 13:31

    that lets you measure things at scale

  409. 13:33

    when we can all kind of converge on the

  410. 13:34

    measurement and it becomes somewhat of

  411. 13:38

    clear rubric and so the task at hand now

  412. 13:41

    is we have to adopt we have to adapt

  413. 13:44

    those assumptions for the agentic age.

  414. 13:49

    And so

  415. 13:50

    we can we need to go if we're able to do

  416. 13:52

    that then we can go from execution at

  417. 13:54

    the speed of computer to measurement at

  418. 13:56

    the speed of computer and of course the

  419. 13:58

    measurements need to fit the mental

  420. 13:59

    models of the customers using it.

  421. 14:01

    And

  422. 14:03

    I think the lesson here being that if

  423. 14:05

    you're going to

  424. 14:06

    think of how to build an agent for

  425. 14:08

    something you also have to think about

  426. 14:10

    how do you help the customers build

  427. 14:12

    build or at least create a method for

  428. 14:14

    verifying that the output is good and so

  429. 14:17

    it's not enough to build it we also have

  430. 14:18

    to help them

  431. 14:20

    we also we also also have to help them

  432. 14:21

    get to clear ROI calculations to justify

  433. 14:24

    their spend.

  434. 14:26

    And so this brings me to the idea of

  435. 14:29

    mouse power which could be the

  436. 14:31

    equivalent of horsepower for the agentic

  437. 14:33

    age just as James Watt was able to show

  438. 14:35

    a measure of efficiency relative to the

  439. 14:37

    horses in in the gins in the horse gins

  440. 14:40

    that were the source of power at the

  441. 14:41

    time we perhaps can also figure out how

  442. 14:44

    do we create a baseline of efficiency

  443. 14:46

    for the way we use computers today and

  444. 14:49

    can then demonstrate how much better or

  445. 14:51

    perhaps more performant on certain

  446. 14:53

    vectors an agent could be at that task.

  447. 14:56

    Um and

  448. 14:57

    of course it's not as easy perhaps as

  449. 14:59

    easy a task as he had back then where he

  450. 15:01

    could just study the horse gin cuz it's

  451. 15:03

    not as if we can create some method to

  452. 15:05

    measure cursor movements and like figure

  453. 15:07

    out the delta of how much more efficient

  454. 15:09

    an agent could move them and thus we can

  455. 15:11

    say yeah agents are this much more

  456. 15:13

    performant than humans at these tasks.

  457. 15:15

    Trust me I've I've tried I had Claude

  458. 15:18

    vibe code me this measurement device and

  459. 15:20

    I thought maybe if I can figure out the

  460. 15:22

    movement like the potential movement

  461. 15:24

    across the screen and measure how fast

  462. 15:26

    it went, I could get some clean measure

  463. 15:28

    of mouse power. Uh but of course this is

  464. 15:30

    only joking. Um this is of course um

  465. 15:33

    like a fool's errand because information

  466. 15:35

    space is just way too high dimensional

  467. 15:37

    and so I think mouse power is is never

  468. 15:40

    going to be a metric of course, but it's

  469. 15:41

    more so an idea. Which the idea being if

  470. 15:44

    you're going to sell somebody an agent,

  471. 15:45

    you also have to help them with the with

  472. 15:47

    the rubric of how do we actually verify

  473. 15:50

    that this agent is doing good work and

  474. 15:52

    thus we can uh have a good measure of

  475. 15:54

    saying that these tokens are worth it.

  476. 15:56

    Um so how to do that of course is is

  477. 15:58

    really up to you and I won't be able to

  478. 16:00

    tell you how do you I don't have any

  479. 16:02

    good frameworks for how do you figure

  480. 16:04

    out the right measurements to to help

  481. 16:05

    provide anybody you're building an agent

  482. 16:07

    for. Uh but what I can do is give a

  483. 16:09

    principle uh give an idea that I've been

  484. 16:11

    kicking around which is based um

  485. 16:13

    in information theory. So going back to

  486. 16:16

    Claude Shannon's ideas about measuring

  487. 16:18

    entropy and information.

  488. 16:20

    Uh entropy being uh the uncertainty of a

  489. 16:22

    probability distribution and of course

  490. 16:25

    very much the basis of how we train

  491. 16:27

    agents today.

  492. 16:28

    Things like cross entropy and such uh

  493. 16:30

    being a big factor in determining how

  494. 16:31

    capable an agent is.

  495. 16:33

    Um I think that entropy's an interesting

  496. 16:35

    idea to think through with regards to

  497. 16:36

    not just the performance of an agent,

  498. 16:39

    but also the task that we're setting

  499. 16:40

    them out to to perform on.

  500. 16:42

    And so uh I put together this matrix

  501. 16:45

    which uh it maps on the x-axis axis the

  502. 16:49

    uncertainty in the steps it takes to

  503. 16:50

    perform a task. And so when we're

  504. 16:52

    thinking of building an agent, I think

  505. 16:54

    it's not enough to just think what would

  506. 16:56

    be a valuable task for the agent to do,

  507. 16:58

    but also thinking about how um how much

  508. 17:01

    uncertainty are in the steps to perform

  509. 17:02

    that task itself. So an example would be

  510. 17:05

    uh booking a flight has uh much less

  511. 17:08

    uncertainty than let's say painting a

  512. 17:09

    masterpiece, right? Because you know

  513. 17:11

    there's certain information that has to

  514. 17:13

    be that has to happen in the flight

  515. 17:14

    purchase. There has to be a departing

  516. 17:16

    destination, arriving destination.

  517. 17:18

    There's going to be a seat chosen. It

  518. 17:19

    might be by the person. It might just be

  519. 17:21

    random.

  520. 17:22

    But these things have to happen for that

  521. 17:24

    task to be completed. And on the other

  522. 17:25

    hand, there is the task of like painting

  523. 17:28

    a masterpiece, right? And who knows what

  524. 17:30

    the steps are to that? Maybe you can get

  525. 17:32

    an agent to do it, but it would be very

  526. 17:34

    hard to figure out how we can actually

  527. 17:36

    create a a relatively predictable

  528. 17:39

    pathway to that.

  529. 17:40

    But then on the other axis is the the

  530. 17:43

    uncertainty in the acceptance criteria

  531. 17:45

    itself. So not just can the agent

  532. 17:47

    perform the task, but can we help

  533. 17:49

    somebody actually or is is there

  534. 17:51

    actually a a clean rubric for how it's

  535. 17:53

    graded? And so thinking about ideas on

  536. 17:56

    on these two axes and where they

  537. 17:57

    intersect, perhaps gives us a better

  538. 17:59

    guide for how to build agents and we can

  539. 18:01

    run through a few examples. So if we

  540. 18:05

    look at the at the left side, your right

  541. 18:07

    side.

  542. 18:09

    Yes.

  543. 18:10

    No, you're left as well.

  544. 18:12

    Then

  545. 18:13

    No, last speaker was also a confused

  546. 18:15

    about.

  547. 18:16

    So uh

  548. 18:17

    Yeah, on the left side when uncertainty

  549. 18:20

    in the task steps are low, then it's a

  550. 18:22

    very it's a very predictable outcome or

  551. 18:25

    it's a very predictable pathway to

  552. 18:26

    achieve that goal. And so then, you

  553. 18:28

    know, why would you waste tokens? Just

  554. 18:30

    write a script. On the other side, when

  555. 18:33

    the

  556. 18:34

    the steps to do perform the task are

  557. 18:36

    very high in in uncertainty, then you

  558. 18:39

    you have very unpredictable information.

  559. 18:40

    And so it's probably at risk of being

  560. 18:43

    out of distribution in pre-training and

  561. 18:45

    probably has very sparse rewards for

  562. 18:46

    reinforcement learning. And so perhaps

  563. 18:48

    it's not a a good task for an agent

  564. 18:50

    because it's just much harder to figure

  565. 18:52

    out how to actually model that data.

  566. 18:54

    And so obviously in the middle is is um

  567. 18:58

    is so I'm I think I'm out of time, but

  568. 19:00

    I'm not getting kicked off yet.

  569. 19:02

    So I'll just finish this up quickly. Um

  570. 19:04

    so yeah, in the middle is is probably

  571. 19:05

    the sweet spot, but then on the other

  572. 19:06

    axis,

  573. 19:07

    what's the uncertainty in verifying that

  574. 19:09

    this is actually valuable? So when you

  575. 19:11

    have high uncertainty in the acceptance

  576. 19:13

    criteria, you pretty much are in a spot

  577. 19:14

    where verification is indistinguishable

  578. 19:16

    from execution. So, why would you build

  579. 19:18

    an agent for something that to verify

  580. 19:21

    was useful, a person pretty much has to

  581. 19:22

    do the work again. So, like waste of

  582. 19:25

    tokens obviously.

  583. 19:27

    And then it leaves that that middle area

  584. 19:29

    where you have this interesting

  585. 19:30

    intersection of tasks that are um

  586. 19:33

    they're not too uncertain in that

  587. 19:36

    they or they they have a degree of

  588. 19:38

    uncertainty where they're not great

  589. 19:41

    they're not just a a script or they're

  590. 19:43

    not out of distribution for training,

  591. 19:45

    but they have enough uncertainty to be

  592. 19:46

    interesting, but at the at the same time

  593. 19:49

    they also have a property of being

  594. 19:50

    relatively easy to

  595. 19:52

    validate the

  596. 19:54

    to check the value of them. And so they

  597. 19:56

    become in this place where they kind of

  598. 19:58

    become the shape of an NP-style problem,

  599. 20:00

    which means they're easier to verify

  600. 20:01

    than to execute. And the reason I say

  601. 20:03

    that is because if you can figure out a

  602. 20:05

    pretty repeatable pattern for verifying

  603. 20:07

    their their work, you can actually just

  604. 20:09

    throw agents to that problem as well.

  605. 20:11

    And so of course you don't just build

  606. 20:13

    the agent, you perhaps build the agent

  607. 20:15

    that verifies the work of the agent.

  608. 20:17

    Um and so yeah, this is perhaps this is

  609. 20:21

    a thought starter mostly kind of have

  610. 20:23

    kind of still in the works, so

  611. 20:25

    I'm happy to hear any thoughts on it,

  612. 20:26

    but if with this guidance I hope when

  613. 20:29

    you're building your next agent you can

  614. 20:30

    also figure out how to also build its

  615. 20:32

    mouse power.

  616. 20:34

    And thanks very much.