← All speakers

Bio, Work & Ideas

Doug

Conference affiliation: Braintrust · 2025

Doug participated in an AI Engineer workshop. The workshop covered AI evaluation workflows, automated scoring methods, and human review processes.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

A changelog generator provides a practical path through prompt comparisons, code-defined evals, production tracing, and human feedback that improves both the application and its judges.

  • How do you know a change makes an AI application better?
    0:48 ↗
  • Task, dataset, scorer
    7:20 ↗
  • Expand the task without losing the test boundary
    11:34 ↗
  • Connect the changelog application
    21:08 ↗
  • Evaluate accuracy, completeness, and formatting separately
    27:15 ↗
  • Publish resources, then run evaluations
    35:08 ↗
  • Treat the judge as something to evaluate
    40:46 ↗
  • Inspect rationales and investigate disagreement
    47:32 ↗
  • Keep experiments as the historical record
    53:00 ↗
  • Trace the application and score live traffic
    55:19 ↗
  • Turn scored logs into regression cases
    1:04:29 ↗
  • Use human feedback to improve the application and the judge
    1:10:14 ↗
  • Keep the evaluation connected to the changing application
    1:15:44 ↗

References