← All speakers

Bio, Work & Ideas

Joel Becker

Conference affiliation: METR · 2025

Joel Becker is the founder and chief executive of Qally’s and an AI evaluation researcher whose work tests whether increasingly capable models make skilled people more productive. Formerly on the technical staff at METR, he helped develop human-calibrated measures of autonomous AI capabilities and co-led a randomized experiment that found experienced open-source developers worked more slowly with AI assistance.

Becker studied economics and econometrics at the University of Bristol, worked as a predoctoral fellow at Harvard and the National Bureau of Economic Research, pursued graduate economics research at New York University, and was a visiting researcher at Oxford’s Global Priorities Institute. He also co-owned a statistics consultancy serving professional soccer teams.

His research extended into genomics: he was first author of a 2021 study introducing the Polygenic Index Repository and examining bias in genetic prediction. He subsequently founded Qally’s and worked on evaluation methods at METR, applying economic ideas about incentives, counterfactuals, and measurement to advanced AI systems.

  • Human-calibrated AI time horizons. Becker coauthored research measuring autonomous task completion and contributed to human baselines and data collection. The resulting metric estimates how long a task would take a skilled human if an AI system can complete it with a specified probability. A 50 percent horizon makes capability easier to interpret, but does not establish that models can handle organizational context or achieve the reliability required in production.
  • The developer-productivity paradox. In a randomized study of 16 experienced open-source developers completing 246 tasks, participants took 19 percent longer when early-2025 AI tools were allowed, despite believing those tools accelerated their work. Becker attributes the gap to factors including verification costs, unreliable outputs, complex codebases, and developers’ extensive repository knowledge. The finding describes a specific population and period, not all programmers or subsequent models.
  • Mergeability over benchmark passing. Becker contributed to research on agent-written pull requests indicating that approximately half of benchmark-passing changes would not have been accepted by maintainers. Automated tests cannot fully capture architectural fit, maintainability, or the judgment exercised during human review.

His research on compute slowdowns further examines how constraints on computing investment could delay projected AI capability milestones. Across these projects, Becker focuses on whether measured performance survives the practical demands of real work: ambiguous goals, accumulated context, coordination, quality control, and deciding which tasks are actually worth doing.

Read the topics behind these talks

2 conference talks

Key ideas

Scroll to read ↓

Compute trends suggest longer AI task horizons, but expert development, messy data, and physical production expose the work that capability measurements still miss.

  • What happens when compute growth bends?
    0:23 ↗
  • A task horizon measures human work
    3:24 ↗
  • Learning a tool can cost time before it saves time
    6:35 ↗
  • Which repository, and how much experience?
    13:04 ↗
  • An unusual population can expose an important bottleneck
    18:16 ↗
  • Writing SQL is easier than knowing what the data means
    25:15 ↗
  • The hidden joins behind a deployment metric
    30:23 ↗
  • Expertise, institutions, and adoption
    33:08 ↗
  • Completion time and speed are different scales
    36:42 ↗
  • Generating a change is only part of submitting a clean PR
    40:34 ↗
  • Measuring capability under oversight
    46:21 ↗
  • Measure work people can build on
    51:04 ↗
  • What happens when the goal is not benchmark-shaped?
    55:58 ↗
  • Who does the task reformulation?
    1:01:22 ↗
  • Similar slopes can hide different starting points
    1:03:34 ↗
  • Closing the loop requires physical production
    1:06:24 ↗
  • Compute headroom, mask design, and yield
    1:12:25 ↗

Key ideas

Scroll to read ↓

AI benchmarks show rapid gains, yet a randomized study found experienced developers took longer with AI access. Understanding both results requires examining what each measurement rewards.

  • What does an impressive benchmark score tell us?
    0:21 ↗
  • Use human task duration as the difficulty scale
    3:13 ↗
  • The time horizon is a success threshold, not a runtime
    4:55 ↗
  • What changes when the measurement moves into real work?
    7:16 ↗
  • Randomize AI access on developers’ own issues
    10:36 ↗
  • Where the extra time can enter
    14:14 ↗
  • Could the experiment itself explain the gap?
    16:45 ↗
  • Reliability, mergeability and the cost of losing context
    18:16 ↗
  • Extend the tasks and the experiments
    20:25 ↗

References