▶ Watch ↗AI Engineer Europe 202640:51
Judge the Judge: Building LLM Evaluators That Actually Work with GEPA — Mahmoud Mabrouk, Agenta AI
Read the full talk →Key ideas
Scroll to read ↓A customer-support evaluator learns airline policy from annotated failures, revealing how seed prompts, reflection, candidate diversity and data quality shape reliable evaluation.
- The monitor says everything is fine0:00 ↗
- Evaluation determines how fast the application can improve1:18 ↗
- A cancellation can succeed and still violate policy5:47 ↗
- Derive the metrics from observed errors8:10 ↗
- The explanation supplies the missing policy knowledge11:46 ↗
- Generate candidates, test them, preserve complementary strengths14:32 ↗
- The evaluator is the optimizer’s feedback channel19:44 ↗
- Keep related tasks out of the validation split22:26 ↗
- Start with a judge that does not invent violations26:06 ↗
- Teach reflection to extract rules from failures29:09 ↗
- Better accuracy does not mean the merge problem is solved32:46 ↗
- Inspect one iteration before paying for a large search34:51 ↗
- Supplying the policy and learning the policy are different starting points37:04 ↗
- Spend the optimization budget where it can reduce recurring cost38:16 ↗