▶ Watch ↗AI Engineer World's Fair 20261:03:26
Context Engineering in 2026: Compaction, Memory & Cost
Read the full talk →Key ideas
Scroll to read ↓An AI tutor’s compaction experiments show how cached history, retrieval, and hardware limits change what belongs in the context window—and when removing it becomes necessary.
- The instruction was there. Why did the agent ignore it?0:34 ↗
- One window, several kinds of context5:14 ↗
- Start with inexpensive compaction8:51 ↗
- Keep the source outside the window11:55 ↗
- A shorter prompt can cost more15:32 ↗
- Choose what changes, then measure it18:55 ↗
- One agent, middleware, and two tools21:56 ↗
- Retrieve a few useful chunks from a large corpus24:43 ↗
- Give browsing a bounded filesystem27:07 ↗
- Turn an intuitive configuration into an observable system32:17 ↗
- Test answers, retained facts, and whether the policy ran36:23 ↗
- Full history beats the initial defaults42:25 ↗
- Cheaper cached inference preserves the same advantage47:10 ↗
- Long-context recall and service cost are different questions53:24 ↗
- Caching cannot make an oversized input fit55:36 ↗
- Combine semantic retrieval with keyword search58:31 ↗
- Name the constraint before choosing compaction1:00:33 ↗