← All speakers
  • SWE Hiring is Cooked ↗nickheiner.com
  • Nick Heiner's Substack | Substack ↗

    Independent benchmarks, essays on the future of work, and dispatches from someone building AI products and testing AI agents every day at Surge AI. Click to read Nick Heiner's Substack, a Substack publication. Nick Heiner's Substack

    nickheiner.com

Nick Heiner leads reinforcement-learning environments at Surge AI, developing realistic simulations that test whether AI agents can complete demanding professional work. A former Netflix engineer and founding engineer at Fixie, he focuses on the widening gap between impressive benchmark scores and practical usefulness.

Heiner studied at Cornell University and worked at Opower and the United States Digital Service before joining Netflix, where he worked on user-interface platforms. After ChatGPT accelerated his interest in AI, he joined Fixie as a founding engineer and subsequently moved to Surge. The company has also identified him as its vice president of product in an assessment of Claude’s coding behavior.

He coauthored Corecraft, a simulated customer-support workplace that measures agents against multistep tasks, realistic operational tools, and expert-authored grading criteria. The research found that leading models struggled when required to satisfy every criterion, while training in high-fidelity environments improved performance on held-out tasks and other evaluations.

  • Benchmark contamination obscures actual capability. Public test material can enter training data, making memorization resemble progress. In his AI Engineer World’s Fair talk, Heiner examines SWE-bench Verified and argues for private holdout sets, contamination disclosures, and scrutiny of what models have already encountered.
  • Prompt-verifier alignment makes evaluations trustworthy. Graders should check everything a task requests, and tasks should disclose everything graders enforce. Otherwise, contradictory instructions, rigid formatting requirements, and Unicode-based reward hacking can produce misleading scores regardless of an agent’s real competence.
  • Expert judgment is essential for sophisticated work. Heiner helped develop Hemingway-bench, which uses blind comparisons by professional writers to evaluate qualities such as originality, coherence, and taste. He argues that automated judges and mechanical checklists cannot reliably capture nuanced writing quality.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

A leaderboard can reward progress that users never feel. Nick Heiner traces how benchmark design, verification shortcuts, and lab incentives create that gap—and what better evaluation requires.

  • The release chart meets actual use
    0:19 ↗
  • Why weak benchmarks remain influential
    1:31 ↗
  • Good tasks are expensive to create and replace
    3:19 ↗
  • Public tests can become remembered answers
    4:41 ↗
  • The verifier can reward the wrong thing
    5:54 ↗
  • Instruction following needs product judgment
    7:32 ↗
  • Quality control is part of measurement
    9:35 ↗
  • Optimizing past what people prefer
    10:51 ↗
  • Anonymous voting and selective disclosure
    11:55 ↗
  • Build the benchmark around deployment
    13:09 ↗
  • Align the prompt, verifier, and apparent ceiling
    14:26 ↗
  • Paying for the judgment the task requires
    15:38 ↗

References