Stop Chunking Like It's 2022 — Yuval Belfer, AI21 Labs
AI Engineer World's Fair 2026 · 18:00
Enterprise AI agents and foundation models
AI21 Labs develops AI systems for enterprises, with Maestro helping teams optimize agents for production. Its tools adjust execution strategies, optimize the combination of models and agent harnesses, and dynamically route calls among models to balance cost and output quality. The company’s work spans agent orchestration and foundation models, although its 2026 restructuring concentrated the business around Maestro.
Founded in 2017 by Ori Goshen, Yoav Shoham and Amnon Shashua, AI21 is led by co-CEOs Goshen and Shoham. Its lasting technical contributions include Jamba, which combines Mamba state-space models, Transformer attention and mixture-of-experts layers to address memory use and throughput in long-context inference. The original model supported a 256,000-token context window. AI21 announced a halt to Jamba development in May 2026; Shashua subsequently expressed his intention to leave the company and board in July.
The May 2026 restructuring reduced AI21’s workforce from 180 to about 70 employees. Its Nebius relationship is a technology licensing partnership; both companies denied characterizations of a sale or acquihire. AI21 said it had signed Maestro contracts worth tens of millions of dollars with international customers, including Nebius. Earlier, its completed $208 million Series C in 2023 brought total funding to $336 million at a $1.4 billion financing valuation.
This AI21 Labs archive currently contains one supplied recording. Use it as a guide to query-dependent retrieval and multi-scale indexing; the findings below are claims reported in the recording, not verified statements about today’s products or performance.
Start with Stop Chunking Like It’s 2022 — Yuval Belfer for the tradeoff between focused detail and broader context. Contrasting Seinfeld questions and experiments with six chunk sizes illustrate why one indexing-time choice can lose information needed by later queries. Belfer reports an answer-informed oracle gap of roughly 20–40 percent; the supplied summary does not establish whether this means relative improvement or percentage points.
Return to Belfer’s talk for a concrete retrieval approach: index multiple window sizes, query the indexes, map retrieved chunks to whole documents, and combine document rankings with Reciprocal Rank Fusion. He reports matching or beating the best fixed-size approach across QMSum, NarrativeQA, Seinfeld, and FinanceBench, with roughly two to five times the memory and little added latency through parallel retrieval. Treat these as reported experimental results. Window selection and better fusion remain open questions in the talk.
AI Engineer World's Fair 2026 · 18:00
AI Engineer World's Fair 2026 · 18:06
AI Engineer World's Fair 2025 · 10:58
Affiliations reflect their AIE appearances, not necessarily current employment.