▶ Watch ↗AI Engineer World's Fair 202648:17
The State of Model Routing — NVIDIA, Cognition, OpenRouter
Read the full talk →Key ideas
Scroll to read ↓Routing an agent means deciding who plans, who executes, what context they share, and when to switch—not merely choosing the cheapest model for a prompt.
- Which tasks justify the expensive model?0:21 ↗
- Keep frontier judgment, delegate the implementation3:15 ↗
- The task changes while the agent is working8:48 ↗
- Cheap tokens can produce an expensive task13:53 ↗
- A sidekick keeps its own running context18:23 ↗
- Routing can happen inside a model artifact20:12 ↗
- Remember where the evidence lives21:59 ↗
- Cache-aware routing meets the always-running agent24:43 ↗
- Local capacity and shorter context change the calculation29:53 ↗
- Use supervision opportunities to detect trouble32:43 ↗
- Measure confusion, then inspect the trace37:48 ↗
- Learn from the routing decisions users correct41:16 ↗
- What remains when models become better collaborators?43:04 ↗