← All speakers

Bio, Work & Ideas

Paul

Conference affiliation: ElevenLabs · 2025

Paul participated in an AI Engineer World's Fair 2025 workshop. The workshop covered multilingual voice agents, speech synthesis customization, and retrieval-augmented orchestration.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Follow a voice agent from speech recognition to language switching, tool calls and spoken replies, then examine what breaks with slow systems, mixed languages and domain vocabulary.

  • Which languages—and which accents—should the agent speak?
    0:22 ↗
  • Text to Bark: generated sound is not translation
    5:08 ↗
  • Speech, text, intelligence, speech
    8:26 ↗
  • A transcript can carry timing, speakers and events
    10:51 ↗
  • Forward a voice message, get readable text
    12:39 ↗
  • Choose the intelligence layer and the speaking voice
    17:05 ↗
  • Configure a conference agent
    20:31 ↗
  • Switch languages without restarting the conversation
    23:38 ↗
  • Language recognition selects a configured voice
    28:49 ↗
  • Give the agent appointment-setting tools
    30:58 ↗
  • Balance response time, session cost and task complexity
    33:01 ↗
  • Keep the caller informed while tools run
    39:06 ↗
  • A mixed-language question exposes a recognition boundary
    43:56 ↗
  • Voice generation needs misuse controls
    49:00 ↗
  • Test the languages people actually mix
    52:00 ↗
  • Generated voice is only one part of an avatar
    55:06 ↗
  • Pronouncing SAP and recognizing Joule are different problems
    57:49 ↗

References