▶ Watch ↗AI Engineer World's Fair 202543:42
Yineng Zhang is senior director of inference at Together AI, a core maintainer of SGLang, and a co-creator of TokenSpeed. He builds the open-source systems that make advanced language models faster and more practical to run, from GPU kernels and attention mechanisms to distributed serving infrastructure.
Earlier in his career, Zhang optimized recommendation-ranking models and language-model inference at Meituan. After leaving, he joined the SGLang project, working with its creator, Lianmin Zheng, and contributing to FlashInfer, the attention and sampling infrastructure underpinning its runtime.
At Baseten, he helped deploy newly released open models, co-authoring Qwen 3 deployment benchmarks covering mixture-of-experts architectures, tensor and expert parallelism, FP8 quantization, and latency-throughput tradeoffs. He also contributed to SGLang’s DeepSeek-V3 and DeepSeek-R1 work. He subsequently joined Together AI as a principal AI researcher before becoming its inference leader.
Zhang also serves on the LightSeek Foundation’s governing board and helped create TokenSpeed, an inference engine developed with contributors from Together AI, NVIDIA, AMD, and other organizations.
▶ Watch ↗AI Engineer World's Fair 202543:42