← All speakers

Bio, Work & Ideas

Paige Bailey

Conference affiliation: Google DeepMind · 2026

Paige Bailey is an engineering lead for developer relations at Google DeepMind and a contributor to the research behind PaLM 2, Gemini, and Gemma. Her career spans geophysics, scientific Python, machine-learning infrastructure, developer tools, and frontier models, with a consistent emphasis on making advanced capabilities accessible to working developers.

Bailey began working with machine learning around 2009, contributing to open-source scientific-computing projects including NumPy, SciPy, Matplotlib, and scikit-learn. She studied geophysics and applied mathematics, pursued graduate work in computer science and carbonate geology, and worked at Chevron on subsurface geoscience, velocity modeling, drilling optimization, distributed computing, and GPU-based workloads.

She contributed to TensorFlow, worked on machine-learning developer experiences at Microsoft and Google, and joined GitHub, where her work included VS Code and early GitHub Copilot user-experience testing. Returning to Google, she contributed to the PaLM 2, Gemini, and Gemma model programs before leading developer-relations engineering at Google DeepMind.

Her independent projects include thinking-in-data, a VS Code extension pack for exploring and visualizing data; signals-and-systems, interactive visualizations for an open-source engineering textbook; and Gemini and Gemma examples, practical notebooks for contemporary AI models.

  • Open models as practical independence. Bailey advocates models developers can download, adapt, and run on their own infrastructure, including offline or privacy-sensitive environments. Her characteristically irreverent case for open models emphasizes affordability, stability, and freedom from unexpected platform changes.
  • Model selection as an engineering decision. She evaluates models against cost, latency, tools, and actual task performance. Her hands-on Gemini demonstration shows a smaller model using sandboxed Python to analyze images quickly; her assessment of computer-use benchmarks for CAD argues that affordable systems approaching frontier performance could broaden access to physical-design tools.
  • Multimodal applications with persistent state. Bailey turns model capabilities into complete software: one example identifies books from a photograph, supplements missing details through search, authenticates users, and stores their libraries in Firebase. She also distinguishes Project Genie’s dynamically generated video environments from exportable three-dimensional game assets.
  • Generative media with creative control and safeguards. Her work with Veo addresses camera control, visual consistency, synchronized dialogue, and SynthID watermarking. A commercial-reconstruction example contrasts manually stitching together generated video, speech, and music with a more integrated Veo 3 workflow.
  • Embodied AI with separated control layers. In robotics, Bailey distinguishes multimodal perception and high-level planning from specialized software executing physical movement locally—a practical architecture for applying advanced models to hardware without confusing reasoning with motor control.

Read the topics behind these talks

3 conference talks

Key ideas

Scroll to read ↓

From timestamped dinosaur videos to a persistent bookshelf catalog, Paige Bailey demonstrates how models, grounding, executable tools and app infrastructure work together.

  • Which model belongs in the application?
    1:35 ↗
  • Turn a dinosaur video into a cited table
    5:48 ↗
  • Give smaller models executable vision tools
    13:06 ↗
  • Supply recent sources, then choose a model tier
    16:20 ↗
  • Keep screen, speech and camera in one conversation
    19:34 ↗
  • Ask for a bookshelf catalog with authentication and storage
    24:51 ↗
  • Navigate a world generated as pixels
    29:17 ↗
  • Ask for a vector representation of the Lego image
    38:14 ↗
  • Combine reference images with a searched object
    41:03 ↗
  • Inspect the SVG, then add code execution
    44:22 ↗
  • Generate footage from an expanded creative brief
    47:05 ↗
  • Evaluate an open model before running it locally
    50:34 ↗
  • Give the food truck a Spanish theme song
    52:55 ↗
  • Test authentication and persistence before sharing
    56:30 ↗
  • Separate planning from action in the systems around you
    59:25 ↗

Key ideas

Scroll to read ↓

Follow video analysis, a bookshelf app, a book-to-media pipeline, and local Gemma agents through working results, failed assumptions, and live repairs.

  • From scientific Python to model infrastructure
    0:15 ↗
  • Which constraints will survive the next model?
    4:42 ↗
  • A toolbox of inputs, outputs, and application surfaces
    8:07 ↗
  • Turn a dinosaur video into a grounded table
    11:51 ↗
  • Bounding boxes versus a verification loop
    17:42 ↗
  • Specify an account-linked bookshelf catalog
    21:04 ↗
  • Change language while sharing a screen
    23:53 ↗
  • Explore a world generated as pixels
    28:56 ↗
  • Recognition succeeds before persistence does
    35:07 ↗
  • Choose the media model for the job
    43:11 ↗
  • Prepare a book-to-media pipeline
    51:27 ↗
  • Select the API surface and retain useful context
    54:20 ↗
  • Replace implicit image history with explicit references
    58:09 ↗
  • Animate the scene, then score the chapter
    1:03:01 ↗
  • Create distinct characters with two configured voices
    1:07:53 ↗
  • Convert the notebook without hiding the workflow
    1:13:40 ↗
  • Composition limits, service tiers, and capacity
    1:18:40 ↗
  • Fit the model to hardware you control
    1:24:26 ↗
  • Run skills and small coding loops on a phone
    1:29:42 ↗
  • Serve one local model to an agent workload
    1:32:47 ↗
  • Connect OpenCode, then test the generated game
    1:36:54 ↗
  • A game inside a game, and a screenshot into HTML
    1:44:43 ↗
  • Map language onto robot and development tools
    1:51:00 ↗
  • The final test is an interaction
    1:52:45 ↗

Key ideas

Scroll to read ↓

A tour of reference images, scene editing, native audio and music tools leads to a practical comparison: rebuilding a commercial with separate generators versus one Veo 3 prompt.

  • What should it take to create an ad?
    0:27 ↗
  • Keep the character, change the environment
    1:57 ↗
  • Extend a scene, edit its contents, direct its motion
    4:04 ↗
  • Generate sound and pictures together
    6:16 ↗
  • Image detail includes typography
    8:58 ↗
  • Music generation as a workspace for exploration
    9:51 ↗
  • One raccoon prompt across model generations
    11:57 ↗
  • Animate a still and expand a short prompt
    13:20 ↗
  • Music, multilingual dialogue and visual oddities
    14:49 ↗
  • Access routes and a small generation request
    16:32 ↗
  • Rebuild a commercial with separate generators
    17:55 ↗
  • One prompt, then an animated farewell
    19:33 ↗

References