← All speakers

Bio, Work & Ideas

Brian Johnson

Conference affiliation: Tavus · 2025

Brian Johnson is a staff applied machine-learning engineer at Tavus and founder of aerbits.ai. He develops the real-time systems that help digital humans recognize when to listen, respond, or ignore interruptions.

Johnson studied electrical engineering at the University of Texas at Austin, earned a law degree focused on intellectual property, and managed engineering teams at Zillow and Earnest. His startup aerbits.ai applied drones and computer vision to detecting illegal dumping, giving municipalities another way to identify urban problems.

At Tavus, Johnson works on conversational timing and selective attention for video agents. His contributions include:

  • Sparrow-0 and semantic turn detection: Sparrow-0 uses linguistic context to distinguish completed thoughts from ordinary pauses, reducing interruptions without introducing unnecessary delays.
  • Sparrow-1 and conversational floor ownership: Sparrow-1 incorporates rhythm, hesitation, intonation, overlapping speech, and individual pacing to determine who should speak and when. Response timing can shift with emotional context: quick exchanges demand speed; sensitive moments benefit from a pause.
  • Real-time speaker identification: His work on speaker embeddings and speaker focus helps digital humans distinguish their conversation partner from background voices and other interruptions.

Johnson also advocates integrating Tavus models with established conversational infrastructure. At AI Engineer World’s Fair 2025, he explained how Tavus rendering, turn-taking, and perception can connect with Pipecat, the separate open-source orchestration framework developed by Daily. Building these systems has also sharpened his attention to how silence, hesitation, and attentive listening shape human conversation.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

A responsive video replica needs more than fast rendering: it needs streaming orchestration, sensible turn timing, synchronized media, and infrastructure that connects each user to a bot.

  • What would make the robot concierge work?
    0:19 ↗
  • From speech pipelines to video replicas
    1:04 ↗
  • Understanding why the bot behaves that way
    3:41 ↗
  • Frames, processors, and pipelines
    6:13 ↗
  • From incoming audio to a completed turn
    9:25 ↗
  • Turning synthesized speech into synchronized video
    11:37 ↗
  • Running analysis beside the conversation
    12:25 ↗
  • Moving model capabilities into the orchestration layer
    14:06 ↗
  • Finishing a turn and choosing when to answer
    15:20 ↗
  • Starting a bot and connecting its media
    17:11 ↗

References