Nicole Brichtova develops generative media products that let people create and revise images through conversation and visual references. At AI Engineer World’s Fair 2026, she represented Google DeepMind in a generative media panel. Her work connects model capabilities to practical creative decisions: what an edit should change, what it should preserve, and how users can refine the result.
From image prompting to conversational editing
In December 2024, she co-authored the announcement for Whisk, a Google Labs experiment built around image-based prompting. Users provide pictures for a subject, scene and style; Gemini converts those references into captions, and Imagen 3 generates an image from them. The workflow gives people a way to express visual ideas they may struggle to describe in words. It favors exploration and remixing, capturing the essence of a reference rather than reproducing it precisely.
In March 2025, Brichtova and fellow product manager Kat Kampf introduced Gemini 2.0 Flash’s experimental native image generation to developers. Combining text and image output within one model supported conversational editing: users could revise an image over successive turns, or generate an illustrated story and then change its drawings through feedback. The evolving image became something users could keep working on rather than a result they had to accept or discard.
Preserving identity and refining edits in Nano Banana
The August 2025 Nano Banana launch announcement, co-authored with David Sharon, identified Brichtova as Gemini Image Product Lead and made identity-preserving editing a central promise. Users could change a person’s outfit or surroundings while retaining their likeness, combine photographs, or alter individual parts of a room. The room-editing example showed why continuity matters: someone could paint the walls, add furniture and keep refining the same image without replacing everything else. Her guidance for everyday creators explains how to use these capabilities in practical editing tasks.
Evaluating likeness and keeping editing useful
Likeness evaluation: Brichtova’s account of Nano Banana’s development connects that product promise to evaluation. People notice subtle errors in their own faces that strangers may miss, so her team incorporated evaluations in which members judged generated images of themselves. A picture can look convincing to an unfamiliar viewer while failing to preserve the person it is supposed to depict.
Editing speed and control: She also treats editing speed as part of the interaction design. Minute-long waits interrupt the back-and-forth that makes conversational editing useful: users need to inspect a change, decide what still needs work and try again. Her priorities distinguish easier consumer interfaces, which should require less elaborate prompting, from professional tools that need reproducibility and precise control over what changes. Both depend on helping users communicate intent and preserve the parts of an image they already value.
Evaluating generative media in real workflows
At World’s Fair 2026, Brichtova joined Dumitru Erhan and Shane Gu in a generative media panel moderated by swyx. Their shared discussion examined editing workflows and the difficulty of evaluating generated images, video and audio. The panel distinguished visually appealing outputs from realistic or useful ones: human preferences can reward unwanted aesthetics, while expert feedback and real customer workflows reveal failures that general evaluations miss. That discussion places her product work in a broader practical challenge—making generated media useful for the task a person is actually trying to complete.
Google DeepMind’s generative media team discusses how images, video, audio and language fit together—and why attractive outputs, human preference scores and real creative workflows can point toward different models.
References carry scene, voice and style information that users may struggle to express in language; instructions identify what should change and what should remain.
Joint audiovisual generation models moving lips and audible speech as consequences of one event, addressing synchronization inside generation rather than repairing it afterward.
Human preference can favor sharpness, saturation and flattering skin tones without establishing realism or task success. Expert judgment and instruction following help reveal and control those differences.
Media evaluation combines objective checks, thousands of human-evaluated items, live experiments and feedback from real workflows; free-form editing makes coverage especially difficult.
The useful missing data includes creative trajectories: revisions, selections and the reasons behind them. FDEs can help turn customer failures into improvements upstream in modeling.