← All speakers

Bio, Work & Ideas

Shane Gu

Conference affiliation: Google DeepMind · 2026

On this page

Shixiang Shane Gu researches Gemini Thinking and post-training, drawing on a career in generative modeling, robot learning, and language-model reasoning. His homepage identifies him as a senior staff research scientist at Google DeepMind, focusing on the Gemini Thinking team. At AI Engineer World’s Fair 2026, the generative media panel introduced his work on Omni Thinking and Gemini reinforcement learning. His contributions include co-developing Gumbel-Softmax and co-authoring research on eliciting and improving reasoning in pretrained language models.

From deep learning and robotics to Gemini Thinking

Gu studied engineering science at the University of Toronto under Geoffrey Hinton’s supervision, before pursuing doctoral research in machine learning at the University of Cambridge and the Max Planck Institute for Intelligent Systems. His doctoral supervisors were Richard E. Turner, Zoubin Ghahramani, and Bernhard Schölkopf. His homepage documents this education and his subsequent research roles.

At Google Brain, Gu became a founding member of the robotics team. In the generative media panel, he also described leading a dexterity moonshot. His homepage credits him with developing sample-efficient deep reinforcement-learning algorithms and lists his research with Ethan Holly, Timothy Lillicrap, and Sergey Levine on robotic manipulation with asynchronous off-policy updates. His later publications on offline reinforcement learning address learning from existing experience rather than relying solely on newly collected interactions.

Gu subsequently shifted his main research attention toward language models. In the panel, he explained that this move reflected his expectation that symbolic intelligence would advance faster than physical intelligence. His homepage records his work as a senior researcher on the ChatGPT team and his leadership of Japan market entry at OpenAI, followed by leadership of multilingual Gemini post-training at Google DeepMind and his focus on Gemini Thinking. Having grown up in Japan and studied and researched internationally, his career spans both research communities and language contexts.

Contributions across modeling, learning, and reasoning

Several contributions explain how his research developed across these fields:

  • Learning through discrete choices. With Eric Jang and Ben Poole, Gu co-authored the Gumbel-Softmax research listed on his homepage. The homepage identifies him as a co-inventor of Gumbel-Softmax and places this work among his contributions to generative modeling.
  • A simpler approach to offline reinforcement learning. With Scott Fujimoto, Gu co-authored A Minimalist Approach to Offline Reinforcement Learning, listed on his homepage as a NeurIPS 2021 publication. This work forms part of his broader research on sample-efficient reinforcement learning and learning from existing datasets.
  • Eliciting reasoning without worked examples. Gu co-authored Large Language Models are Zero-Shot Reasoners with Takeshi Kojima, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. His homepage credits this work with zero-shot chain-of-thought prompting using the instruction to think step by step. In the panel, he explained why natural-language reasoning can be useful: intermediate steps can draw directly on knowledge acquired during pretraining, while extracting capabilities through reinforcement learning can be computationally expensive.
  • Turning generated reasoning into supervision. His homepage credits Gu with demonstrating how language models can self-improve through self-generated reasoning, and he recalled this work in the panel. This connects his research on eliciting reasoning with the broader question of how generated reasoning can contribute to further model improvement; the supplied evidence does not describe a specific selection or training procedure.
  • Testing video models beyond media generation. In the generative media panel, Gu described evaluation work on video models as zero-shot learners and reasoners. He discussed classical vision tasks, visual quizzes, and physical intuition, while emphasizing substantial room for improvement. These experiments give concrete substance to his interest in reasoning in space and time: video models may complement symbolic reasoning by representing motion, spatial relationships, and physical consequences.

World models, physical prediction, and useful generation

Gu’s discussion in the World’s Fair 2026 panel reconnects this work with his robotics background. He grounds world models in model-based reinforcement learning, using the term for the model in that framework. His ambition is to bring symbolic reasoning and physical prediction closer together, with video models contributing capabilities that language alone may not capture. He presents cooperation between video and language models as a research direction, with agentic workflows offering an incremental way to explore it.

At the World’s Fair 2026 generative media panel, Gu joined Dumitru Erhan and Nicole Brichtova in a discussion moderated by swyx. The panel examined how language, perception, and generation interact, and why attractive outputs can still fail to be realistic or useful. Its discussion of expert feedback and customer workflows places the evaluation problem alongside the modeling problem: broad preference judgments can miss details that determine whether generated media works for a particular task.

1 conference talk

Key ideas

Scroll to read ↓

Google DeepMind’s generative media team discusses how images, video, audio and language fit together—and why attractive outputs, human preference scores and real creative workflows can point toward different models.

  • References carry scene, voice and style information that users may struggle to express in language; instructions identify what should change and what should remain.
    4:59 ↗
  • Joint audiovisual generation models moving lips and audible speech as consequences of one event, addressing synchronization inside generation rather than repairing it afterward.
    25:20 ↗
  • Human preference can favor sharpness, saturation and flattering skin tones without establishing realism or task success. Expert judgment and instruction following help reveal and control those differences.
    34:43 ↗
  • Media evaluation combines objective checks, thousands of human-evaluated items, live experiments and feedback from real workflows; free-form editing makes coverage especially difficult.
    42:17 ↗
  • The useful missing data includes creative trajectories: revisions, selections and the reasons behind them. FDEs can help turn customer failures into improvements upstream in modeling.
    48:18 ↗

References