▶ Watch ↗AI Engineer Europe 202616:30
Ibragim Badertdinov is a London-based research engineer at Nebius and the lead author of SWE-rebench, an open initiative for evaluating and training coding agents on real software-engineering tasks. He builds the benchmarks, executable environments and reinforcement-learning systems needed to distinguish genuine programming ability from memorized answers, unreliable infrastructure and misleading test results.
Badertdinov trained in dentistry from 2013 to 2018, moved into healthcare management, and began working in machine learning and natural-language processing around 2019. His transition from dentistry to AI research shaped his attention to the consequences of systems failing in practice. By 2024, he was working in research and open source; projects such as his machine-learning design primer also reflect his interest in accessible technical education.
In 2025, he led the original SWE-rebench research, which introduced more than 21,000 interactive Python software-engineering environments. Its continuously updated leaderboard evaluates coding agents against recently created GitHub issues, limiting the advantage models gain when established benchmark answers enter training data.
Badertdinov also coauthored research on guided search for software-engineering agents and multi-turn reinforcement learning; the latter reported improving an open-weight coding agent’s SWE-bench Verified result from 20% to 39%. In 2026, his SWE-rebench V2 expanded the dataset beyond Python to more than 32,000 executable tasks across 20 programming languages and over 3,600 repositories.
curl. Such failures demand scrutiny of network permissions, sandbox boundaries and execution trajectories, not merely passing patches. His AI Engineer Europe talk connects these examples to practical evaluation safeguards.