← All organizations

Research University and AI Research

Stanford University

Stanford University provides undergraduate and graduate education and conducts research across seven schools, including engineering, medicine, business, and the humanities. Its offerings extend to online courses for the public and executive and policy education through the Stanford Institute for Human-Centered AI (HAI). HAI funds research, trains early-career scholars, and produces the annual AI Index, connecting technical work with questions about AI’s economic and societal effects.

Founded in 1885 by Leland Stanford and Jane Lathrop Stanford, the university opened in 1891 and is now led by President Jonathan Levin. Its AI research includes FlashAttention, an algorithm that accelerates Transformer attention without approximation. By computing in blocks and avoiding large intermediate matrices in GPU memory, it reduces memory traffic and makes attention more efficient for model training and inference.

Stanford reported enrollment of 7,289 undergraduates and 10,025 graduate students in autumn 2025. Its funding includes student income, sponsored research, endowment income, gifts, and health care services. In 2026, Stanford announced that Stanford HAI and Stanford Data Science would combine under the HAI name, led by James Landay. The planned structure brings together HAI’s research and policy network with Data Science’s Marlowe computing cluster and fellowship program, linking computational resources with research across disciplines.

Explore the recordings

The supplied Stanford University archive contains one recording: a technical talk on transport design for AI clusters. Use it to explore networking bottlenecks, Homa’s design, and the limits of the benchmark evidence presented.

Start with why small messages matter

In Homa: The End of TCP for AI Clusters, John Ousterhout argues that inference and agentic workloads make small coordination exchanges and tail latency increasingly important. Follow this thread for his explanation of how slow synchronization can leave GPUs idle, and how incast, delayed congestion feedback, and byte-stream head-of-line blocking contribute to delays.

Follow the transport design

Return to the same recording for Homa’s proposed response: independent RPC messages, shortest remaining processing time, receiver-issued packet grants, and switch priority queues. This path connects the bottlenecks described in the talk to mechanisms intended to favor short transfers.

Assess the evidence before planning adoption

In the benchmark and implementation discussion, Ousterhout reports roughly 13× lower short-message P99 latency than TCP and nearly 2× better latency for the longest messages in a mixed-message benchmark. These are recorded benchmark claims, not demonstrated end-to-end AI application speedups. His statements that a Linux kernel module is available on GitHub and upstreaming is in progress describe the time of the talk; the supplied catalog does not establish today’s implementation or upstream status.

Company sources · checked 2026-08-28