Shijia Liao is the co-founder and chief scientist of Fish Audio, where he builds expressive, instruction-controlled voice synthesis. His work gives synthesized speech recognizable character, emotional range, and conversational context while making powerful voice models available to developers.
After studying at the University of Maryland, Liao contributed to multimodal AI research at NVIDIA, including LITA, which helps video-language models locate events in time, and Eagle, which combines visual encoders to improve multimodal understanding.
Liao was first author of the 2024 Fish Speech paper, which introduced a dual-autoregressive speech architecture supporting multilingual generation and voice cloning. At AI Engineer World’s Fair 2025, he demonstrated OpenAudio S1’s ability to control vocal emphasis and emotional delivery. As the first-listed core contributor to Fish Audio S2, he extended that work into natural-language direction, multi-speaker dialogue, and streaming inference, with publicly released model weights and fine-tuning tools.
His contributions center on three connected challenges:
Directed vocal performance: Giving creators precise control over character, emotion, emphasis, and delivery.
Open-source voice infrastructure: Releasing speech models, training tools, and serving infrastructure for practical experimentation and deployment.
Nine startup pitches trace the work around AI models: finding generated content, building conversational devices, structuring enterprise data, controlling speech and normalizing inference.
Krea: predicting cars is easier than predicting traffic