← All speakers

Bio, Work & Ideas

Jost Tobias Springenberg

Conference affiliation: Physical Intelligence · 2025

Jost Tobias Springenberg, known as Tobi, is a robotics and machine-learning researcher at Physical Intelligence developing general-purpose robot policies that can manipulate unfamiliar objects, adapt to new environments, and complete extended physical tasks. His work spans computer vision, reinforcement learning, and the data and model architectures needed to bring capable robots out of controlled laboratories.

Springenberg studied cognitive science at the University of Osnabrück and earned a master’s degree in computer science at the University of Freiburg. His early research included the all-convolutional network, which replaced conventional pooling operations with strided convolutions and introduced a method for visualizing learned image features.

At Google DeepMind, he worked on machine learning for science and robotics. He contributed to Gato, helping deploy the multimodal generalist agent on a physical robot, and coauthored RoboCat, a robotic agent designed to improve across manipulation tasks and hardware configurations. In 2024, he led research on scaling offline actor-critic reinforcement learning, showing how large models can learn continuous-control tasks from both expert demonstrations and imperfect experience.

At Physical Intelligence, Springenberg coauthored research on π0.5, a robot foundation model designed to generalize to unfamiliar homes. His approach brings together several practical convictions:

  • A robotics data engine is foundational. Physical interaction has no ready-made equivalent to the public internet’s training corpus. Useful robot policies require teleoperated demonstrations, careful annotation, and deliberately varied tasks, environments, and hardware.
  • Diverse environments improve generalization. Training across many homes can help robots adapt to unseen settings, including multistep cleaning and clothing-handling tasks.
  • Reasoning and movement operate at different speeds. Vision-language backbones can interpret scenes and decompose instructions, while action-specialist components generate rapid, continuous robot movements; knowledge-insulated architectures help retain pretrained understanding.
  • Long-horizon manipulation requires memory and refinement. His work on multi-scale embodied memory addresses extended tasks, while online reinforcement learning for precise manipulation targets the accuracy and dexterity that demonstrations alone may not provide.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Vision-language-action models connect visual understanding to robot control, but dexterity and generalization depend just as much on how demonstrations are collected, annotated, and diversified.

  • What changes when the environment stops being predictable?
    0:16 ↗
  • From answering questions to controlling a robot
    2:24 ↗
  • A data engine built around human demonstrations
    5:09 ↗
  • More hours, then more kinds of environments
    7:40 ↗
  • From early object selection to dexterous control
    9:04 ↗
  • π0.5 couples task decomposition with continuous actions
    11:18 ↗
  • Testing what transfers to an unfamiliar home
    12:47 ↗
  • From unfamiliar homes to a partner’s robot
    15:03 ↗

References