← All speakers

Bio, Work & Ideas

Alex Duffy

Conference affiliation: Every · 2025

Alex Duffy is the co-founder and chief executive of Good Start Labs, which uses games to evaluate and train artificial-intelligence models. His work tests whether interactive environments can develop abilities that conventional benchmarks struggle to measure: negotiation, trust, deception, and strategic judgment.

Duffy co-founded AI Camp, where students built working AI products, and served as head of product. He later became vice president of AI at Salt AI, leading researchers working on drug discovery and developing tools for emerging AI workflows. His 2024 introduction to Salt’s platform described combining language, images, video, and other media into collaborative software.

At Every, he led AI training and consulting for organizations in journalism, finance, construction, and technology. With Tyler Marques, he developed AI Diplomacy, an experiment placing language models inside a strategy game where success requires private negotiation, coalition-building, and anticipating betrayal. Their work became Good Start Labs, which spun out of Every with $3.6 million in funding in October 2025.

  • Benchmarks shape model behavior. Once an evaluation becomes influential, developers optimize for it; poorly designed tests can therefore reward shallow proxies or harmful behavior. Duffy identifies conversational sycophancy as one consequence of feedback systems favoring agreement over judgment, and advocates human-centered model evaluation grounded in practical situations.
  • Games expose social and strategic reasoning. AI Diplomacy makes models navigate incomplete information and conflicting incentives. In one observed game, Gemini 2.5 Pro gained an early advantage before OpenAI o3 prevailed through shifting alliances, while Claude Opus proved vulnerable to manipulation. These results describe particular game conditions, not fixed model personalities. Duffy develops this perspective in his AI Engineer World’s Fair talk.
  • Training through play can test real-world transfer. Good Start Labs investigates whether reinforcement-learning data generated in games improves practical work. Duffy’s research on the railroad-finance game 1830 reported apparent improvements in research involving corporate filings—an early investigation into transferable strategic reasoning, not proof of universal transfer.

Duffy also argues that teachers, journalists, financial professionals, and other domain specialists should help define worthwhile AI behavior. Their judgments can make evaluation more relevant to the people expected to trust and use these systems.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

A benchmark can turn a personal curiosity into a development target. AI Diplomacy shows what becomes visible when the test includes negotiation, deception and human judgment.

  • How a personal test becomes a public target
    0:00 ↗
  • From an interesting failure to a saturated benchmark
    1:51 ↗
  • What a positive feedback signal can reward
    5:16 ↗
  • Design for strategies, learning and deeper challenges
    6:19 ↗
  • A game where language changes the strategic situation
    7:36 ↗
  • How o3 broke Gemini’s winning alliance
    8:32 ↗
  • Social competence does not follow a simple ranking
    10:18 ↗
  • Measure what subject-matter experts care about
    12:04 ↗
  • Define the goal, inspect the attempt, give feedback
    12:48 ↗
  • A yoga teacher’s benchmark
    14:16 ↗

References