▶ Watch ↗AI Engineer World's Fair 20251:25:08
[Evals Workshop] Mastering AI Evaluation: From Playground to Production
Read the full talk →Key ideas
Scroll to read ↓A changelog generator provides a practical path through prompt comparisons, code-defined evals, production tracing, and human feedback that improves both the application and its judges.
- How do you know a change makes an AI application better?0:48 ↗
- Task, dataset, scorer7:20 ↗
- Expand the task without losing the test boundary11:34 ↗
- Connect the changelog application21:08 ↗
- Evaluate accuracy, completeness, and formatting separately27:15 ↗
- Publish resources, then run evaluations35:08 ↗
- Treat the judge as something to evaluate40:46 ↗
- Inspect rationales and investigate disagreement47:32 ↗
- Keep experiments as the historical record53:00 ↗
- Trace the application and score live traffic55:19 ↗
- Turn scored logs into regression cases1:04:29 ↗
- Use human feedback to improve the application and the judge1:10:14 ↗
- Keep the evaluation connected to the changing application1:15:44 ↗