← All speakers

Bio, Work & Ideas

Paul Gilbert

Conference affiliation: Arista Networks · 2025

Paul Gilbert is a New York–based enterprise networking specialist who designs the infrastructure connecting large-scale AI systems. An Arista Networks technical lead in 2025, he focuses on keeping expensive GPU clusters productive when synchronized accelerator traffic overwhelms conventional data-center architecture.

Gilbert spent more than 25 years at Cisco Systems, becoming a Distinguished Systems Engineer and working on major data-center and financial-services networks. He left Cisco in 2022 for an AI startup, spent roughly a year there, and subsequently joined Arista, where he specialized in enterprise AI networking. His industry biography traces that progression from established enterprise infrastructure into accelerator-intensive computing.

How Gilbert approaches AI infrastructure

  • Dedicated GPU fabrics: Gilbert separates accelerator-to-accelerator traffic from the front-end networks supplying training data. He favors isolated leaf-spine topologies, straightforward BGP routing, and non-oversubscribed AI networks because synchronized GPUs can saturate links simultaneously and turn one slow device into a cluster-wide bottleneck.
  • Workload-aware network design: Training collectives, fine-tuning, and inference create distinct traffic patterns that conventional enterprise load balancing can mishandle. Gilbert advocates bandwidth-aware traffic distribution and closer coordination between networking teams and the engineers designing model workloads.
  • RoCEv2 congestion control: Explicit Congestion Notification tells senders to slow down as pressure builds; Priority Flow Control acts as an emergency stop when buffers fill. Gilbert also challenges rigid assumptions about lossless networking: occasional packet drops matter less than persistent congestion, unstable latency, and deteriorating job completion times.
  • GPU-to-switch observability: Gilbert combines RDMA error diagnostics, packet-header inspection, workload-tuned switch buffers, and communication between GPU software and Arista EOS to locate failures across accelerator and network boundaries. He also emphasizes upgrades that avoid interrupting active infrastructure.

His AI Engineer Summit session extends that operational focus to rack power, liquid cooling, and NIC-centric Ethernet architectures that shift coordination toward network interfaces while keeping switches focused on efficient packet forwarding.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Training infrastructure must handle synchronized bursts, fragile jobs and unfamiliar power demands. Paul Gilbert explains how those constraints shape the network from GPU ports to operational telemetry.

  • How much infrastructure does a model need?
    0:17 ↗
  • Separate GPU communication from storage
    2:37 ↗
  • Read the server before designing its network
    4:14 ↗
  • Collectives change the bandwidth calculation
    5:44 ↗
  • Enough bandwidth can still land on the wrong uplink
    8:35 ↗
  • Cables, power and cooling become job dependencies
    9:29 ↗
  • Two traffic directions, two congestion responses
    11:04 ↗
  • Keep the fabric simple and involve workload owners
    12:58 ↗
  • Control queues, then preserve evidence of failure
    15:12 ↗
  • Connect host state to switch state
    17:35 ↗
  • Turn the requirements into a deployable fabric
    19:07 ↗
  • Move more transport work into the NIC
    21:22 ↗
  • The final measure is job completion time
    22:05 ↗

References