← All speakers

Bio, Work & Ideas

Julia Neagu

Conference affiliation: Quotient · 2025

Julia Neagu is a member of technical staff at Databricks and the co-founder and former chief executive of Quotient AI, the agent-evaluation company Databricks acquired in March 2026. She builds systems that identify failures in deployed AI agents and turn those failures into improvements.

Neagu studied physics at Princeton and Harvard, earning an AB, MA, and PhD. She developed quantitative models at Aon, led analytics at Tamr, and subsequently directed data and evaluation work for GitHub Copilot. In 2023, she and former GitHub colleague Freddie Vargus founded Quotient to help developers test AI applications against actual product requirements. Her announcement introducing Quotient emphasized realistic, domain-specific experimentation over subjective judgments and generic benchmarks.

  • Production evaluation over static benchmarks. Search agents confront changing websites, unpredictable questions, retrieval failures, hallucinations, and reasoning errors simultaneously. Neagu argues that fixed datasets and conventional monitoring cannot capture these shifting, interconnected production conditions, as she outlined in a collaborative AI-search evaluation session.
  • Reference-free failure detection. Quotient developed evaluators that identify problems in live agent interactions without waiting for labeled answers or human feedback. Its systems examine production traces for failures involving grounding, retrieval, reasoning, and tool use.
  • Continuous agent improvement. Neagu envisions agents recognizing unreliable sources, outdated information, and emerging hallucinations, then using those signals to improve subsequent behavior. Databricks acquired Quotient to bring production-derived evaluation datasets and reward signals into enterprise products including Genie, Genie Code, and Agent Bricks.
  • Practical adoption of open models. Her analysis of open and proprietary language models weighs customization, infrastructure, operating costs, and domain-specific performance; fine-tuning on specialized data can favor open models, while proprietary systems can simplify deployment.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Dynamic datasets, reference-free metrics, and access to retrieved evidence turn search evaluation from a provider leaderboard into a way to diagnose failures.

  • Two inputs you cannot control
    0:16 ↗
  • Correct for which source, time, and user?
    2:30 ↗
  • Keep static benchmarks, refresh the questions
    4:05 ↗
  • Generate questions with an evidence trail
    6:34 ↗
  • Changing the benchmark changes the ranking
    10:21 ↗
  • The expected answer is not the whole response
    12:43 ↗
  • Three checks without a reference answer
    13:42 ↗
  • Completeness tracks performance; documents explain it
    14:58 ↗
  • Relevant evidence does not prevent unsupported claims
    17:04 ↗
  • Use the metric combination to choose an investigation
    18:40 ↗
  • Let evaluation inform the next decision
    19:26 ↗

References