▶ Watch ↗AI Engineer Europe 202622:26
Samuel Humeau is an AI scientist at Mistral AI building systems that understand conversations and generate natural speech. His contributions to Voxtral Realtime and Voxtral TTS address both sides of a voice assistant’s job: recognizing speech as it arrives and answering quickly in an expressive synthetic voice.
Earlier in his career, Humeau worked on multimodal information extraction at Diffbot, combining text, images, and structured product data. At Facebook AI Research, he coauthored Image-Chat, which grounds dialogue in images and conversational style, and helped develop poly-encoders, transformer architectures that make ranking potential conversational responses more computationally efficient.
He subsequently became lead machine learning engineer at Nabla, applying speech recognition and language generation to medical consultations. His work encompassed transcription, speaker identification, clinical summaries, and specialized models. In guidance on evaluating clinical-documentation systems, he emphasized automated testing, clinician review, staged deployment, and feedback from actual medical practice.
At Mistral, Humeau contributed to Voxtral Realtime, a streaming speech-recognition model, and Voxtral TTS, a multilingual speech-generation model. His approach centers on four practical distinctions:
His AI Engineer Europe appearance connected these concerns to a concrete product goal: voice interfaces that retain human expressiveness while remaining responsive, modular, and resistant to casual impersonation.
▶ Watch ↗AI Engineer Europe 202622:26