← All speakers

Bio, Work & Ideas

Dylan Patel

Conference affiliation: SemiAnalysis · 2026

On this page

Dylan Patel is the founder, chief executive, and chief analyst of SemiAnalysis, the semiconductor and AI-infrastructure research firm he started building in 2020. He analyzes the physical and financial foundations of advanced artificial intelligence: accelerator architectures, memory bandwidth, networking, chip manufacturing, electricity, and the economics of operating enormous computing clusters.

An unconventional route into chips

Patel grew up around his family’s small businesses and developed an early fascination with computer hardware after repairing an Xbox. He immersed himself in online communities focused on smartphones, graphics processors, and personal computers, then spent two years as a quantitative analyst at a risk-focused financial firm before turning his independent semiconductor research into SemiAnalysis. What began as a solo publication grew into a research and advisory business covering the entire chain from chip manufacturing to cloud infrastructure. His account of that career trajectory helps explain his close attention to both engineering constraints and commercial incentives.

The infrastructure questions that define his work

  • Prefill and decode demand different machines. Processing a prompt is compute-intensive; generating successive tokens depends heavily on memory bandwidth. Patel applies that distinction to frontier-model serving, including continuous batching, vLLM, TensorRT-LLM, and the practical cost of deploying large open models.
  • Disaggregated prefill changes inference economics. Assigning prompt processing and token generation to separate accelerators reduces contention between customers and enables hardware tailored to each workload. His research on Nvidia’s Rubin CPX connects specialized accelerators, memory choices, rack architecture, and operating costs.
  • Context caching makes long documents affordable. Legal and contract-review applications can waste money repeatedly processing identical source material. Reusing previously computed model state reduces that expense, although caches consume substantial memory and may need to move between accelerators, host systems, and storage.
  • Cluster performance depends on reliability and power. Optical failures, slower individual chips, software immaturity, and electricity availability can erase theoretical hardware advantages. Patel’s comparison of H100 and GB200 systems evaluates training performance alongside downtime, energy consumption, and total ownership costs.
  • AI competition is industrial and geopolitical. His research on Huawei Ascend production identifies high-bandwidth memory as a potential constraint on Chinese accelerator manufacturing. His analysis of global AI infrastructure extends to export controls, Middle Eastern data-center financing, competing rack-scale architectures, and American electricity shortages.

Patel advocates hardware-software co-design: optimizing models, kernels, memory, networking, and silicon together. As inference clusters increasingly support reinforcement-learning workloads, he treats computing capacity as a shared resource constrained by manufacturing, financing, reliability, and power.

Read the topics behind these talks

2 conference talks

Key ideas

Scroll to read ↓

Larger models demand more than faster GPUs: economical inference depends on batching, workload isolation, and cache reuse, while frontier training exposes failures and stragglers at enormous scale.

  • Why better models can still feel like stagnation
    0:00 ↗
  • One request, two resource bottlenecks
    1:27 ↗
  • Admit new requests while existing requests keep running
    3:56 ↗
  • Separate prompt processing from token generation
    5:41 ↗
  • Reuse a document instead of paying to process it repeatedly
    9:06 ↗
  • Beyond the GPT-4 training generation
    12:06 ↗
  • The physical footprint of frontier compute
    13:26 ↗
  • A large cluster must keep training while components fail
    14:08 ↗
  • A functioning GPU can still slow the whole job
    15:38 ↗
  • Build for the models these systems might enable
    17:11 ↗

Key ideas

Scroll to read ↓

Huawei’s optical accelerator systems, Gulf investment deals and America’s electricity constraints show why usable AI compute depends on much more than the chip.

  • Huawei trades power and optics for system scale
    0:36 ↗
  • Foreign inputs complicate export controls
    2:36 ↗
  • The claimed packaged-memory route
    3:43 ↗
  • A process node is not yet an accelerator supply
    4:38 ↗
  • Restrictions still remove substantial supply
    5:38 ↗
  • The UAE bargain: GPU access and US investment
    6:17 ↗
  • Inference capacity can feed training
    8:46 ↗
  • Saudi construction and the supplier ecosystem
    9:35 ↗
  • What export agreements must actually control
    11:33 ↗
  • Who pays before the rental revenue arrives?
    12:53 ↗
  • US electricity supply limits what capital can build
    14:30 ↗
  • Electricity changes the value of chip efficiency
    16:42 ↗
  • Imported tools also help develop domestic substitutes
    17:49 ↗

References