← All speakers

Bio, Work & Ideas

Toki Sherbakov

Conference affiliation: OpenAI · 2025

Toki Sherbakov leads solutions architecture at OpenAI, helping enterprises translate frontier AI models into dependable products and internal systems. His work spans enterprise deployment, research on embeddings, and the emerging architecture of conversational voice agents.

Before joining OpenAI’s enterprise business, Sherbakov spent five years in go-to-market roles at Palantir Technologies and two years building and leading the deployment team at Peregrine Technologies, which developed software for state and local governments. At OpenAI, he has worked directly with customer engineering teams to move applications from initial design through evaluation, production rollout, and ongoing maintenance.

In 2022, he coauthored research on text and code embeddings examining vector representations for semantic search, classification, and code retrieval. OpenAI also credits him among the contributors to GPT-4; its GPT-4o system card includes him in demonstration-production and go-to-market credits. He is additionally among the authors of the GPT-5 system card.

  • Enterprise AI starts with business priorities. Sherbakov advocates identifying consequential use cases, defining success metrics early, and building internal capabilities as applications mature from workforce tools to automation and customer-facing products. His enterprise deployment work includes improving a Morgan Stanley knowledge assistant through embeddings, chunking, reranking, classification, prompting, and query expansion.
  • Voice-agent architecture depends on the application. Sherbakov distinguishes conventional transcription–language model–speech pipelines from speech-to-speech voice agents built with the Realtime API. Integrated systems can improve responsiveness, expressiveness, and conversational continuity; chained systems may offer greater control when accuracy and predictable behavior matter more.
  • Production constraints determine technical trade-offs. In his analysis of practical voice systems, consumer applications prioritize natural interaction and low latency, while customer-service deployments demand accurate actions, internal-system connectivity, and telephony integration. The appropriate architecture follows those requirements.

Read the topics behind these talks

2 conference talks

Key ideas

Scroll to read ↓

A voice agent must respond quickly, accept interruptions, and make reliable decisions. Those demands shape its architecture, prompts, tools, evaluations, and guardrails.

  • Can the agent change course while speaking?
    0:15 ↗
  • From three model stages to speech-to-speech
    4:08 ↗
  • Choose for the consequences of an error
    5:25 ↗
  • Keep the voice responsive while specialists decide
    7:27 ↗
  • Prompt the delivery and the conversation
    9:53 ↗
  • Keep tools focused and carry context forward
    11:37 ↗
  • Evaluate the decisions and the audio
    12:56 ↗
  • Use the gap between generation and playback
    15:06 ↗
  • Build the feedback loop early
    15:49 ↗

Key ideas

Scroll to read ↓

Start with business priorities, build measurable use cases, and let observed failures guide the move from a single agent to specialized networks with explicit safety gates.

  • Getting models into everyday work
    0:17 ↗
  • Three surfaces for enterprise adoption
    1:49 ↗
  • Business strategy before organizational scale
    3:02 ↗
  • Define success before development
    4:26 ↗
  • Dedicated teams, bounded roadmap visibility
    5:54 ↗
  • Morgan Stanley: improving the whole retrieval system
    6:50 ↗
  • The model controls an execution loop
    8:21 ↗
  • Understand the primitives before abstracting
    9:56 ↗
  • Learn from one focused agent in production
    11:43 ↗
  • Handoffs preserve context while changing specialists
    12:57 ↗
  • Run guardrails separately, gate consequential actions
    14:53 ↗

References