▶ Watch ↗AI Engineer World's Fair 202553:15
Daria Soboleva is head research scientist at Cerebras, where she develops efficient language models across training data, sparse architectures, optimization, and specialized hardware. Her contributions include SlimPajama, an open 627-billion-token pretraining dataset, and BTLM-3B-8K, a compact model designed to deliver larger-model performance with substantially less computation.
Earlier in her career, Soboleva developed YATI for Yandex Search, then worked on automatic speech recognition for Google Assistant and models deployed in Google Captions and Gboard. At Cerebras, she shifted from production machine-learning applications toward the data, architecture, and infrastructure underlying foundation models.
In 2023, she co-led SlimPajama’s release, reducing the 1.21-trillion-token RedPajama corpus to 627 billion tokens by removing duplicates and low-quality documents. The project also released distributed preprocessing tools and evaluation splits decontaminated against the training data, making dataset quality a practical lever for reducing wasted computation.
She subsequently co-developed BTLM-3B-8K, a three-billion-parameter model trained on SlimPajama with contexts of up to 8,192 tokens. It delivered competitive results against some seven-billion-parameter models while using approximately 2.5 times less inference computation, expanding the possibilities for deployment under tighter memory constraints.
▶ Watch ↗AI Engineer World's Fair 202553:15