← All speakers

Bio, Work & Ideas

Brooke Hopkins

Conference affiliation: Coval · 2025

Brooke Hopkins is the founder and chief executive of Coval, which builds simulation and evaluation infrastructure for autonomous voice and chat agents. She applies lessons from self-driving cars to a central challenge of conversational AI: making systems reliable when every interaction changes what happens next.

Hopkins studied computer science and mathematics at New York University and began her career developing voice technology for Google Assistant. At Waymo, she built systems for managing simulation datasets and led the evaluation-job infrastructure team, creating developer tools for running autonomous-driving simulations.

She founded Coval in 2024. The company joined Y Combinator’s Summer 2024 batch and raised $3.3 million led by MaC Venture Capital, with participation from General Catalyst and Y Combinator.

How Hopkins approaches reliable AI agents

  • Probabilistic evaluation: Instead of forcing agents into rigid decision trees, Hopkins simulates many plausible conversations and measures how consistently they complete tasks. Her reference-free evaluation judges overall outcomes without prescribing every conversational turn. Repeating failures distinguishes rare anomalies from systematic weaknesses.
  • Purpose-built simulation fidelity: Text-based tests can expose problems with workflows, tool calls, and instruction following; voice interactions reveal interruptions and latency; realistic accents, background noise, and degraded audio matter when those specific conditions are being evaluated. Hopkins treats reliability metrics as product-specific: appointment booking, refunds, and outbound sales require different standards.
  • Human-calibrated automated judgment: Coval’s Metric Studio aligns automated language-model evaluations with human-labeled conversations before applying those judgments at scale. Hopkins also coauthored research on scripted comparative evaluation, which fixes one side of a conversation to isolate model behavior. Coval’s open voice-model benchmarks measure speech-recognition and speech-generation latency and accuracy.

For Hopkins, dependable autonomy depends on continuously measuring real-world uncertainty, translating failures into better tests, and preserving the flexibility that makes conversational agents useful.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Voice agents need more than a convincing demo: they need simulations, reusable metrics, and continuous evaluation that preserve autonomy while exposing failures.

  • Production trust without sacrificing autonomy
    0:16 ↗
  • From road observations to aggregate simulation
    2:59 ↗
  • Conversations need a responsive environment
    4:12 ↗
  • Measure behavior without prescribing every answer
    5:27 ↗
  • A local fix must survive broader regression tests
    6:44 ↗
  • Build metrics by reading conversations
    8:05 ↗
  • Match simulation realism to the question
    9:18 ↗
  • Denoise a failure by rerunning its scenario
    11:26 ↗
  • Choose metrics from the product's purpose
    12:18 ↗
  • Calibrate automated judges with human feedback
    13:44 ↗
  • Benchmark components, then test the full system
    15:16 ↗
  • Give production failures an owner and a test set
    16:53 ↗
  • Voice as an expected enterprise interface
    17:38 ↗

References