AI Engineer World's Fair 2026

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

Read the talk

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes

Yuchen Fama and Ashish Kamra explain how llm-d finds reusable context, separates prompt processing from token generation, and sizes each pool for agentic workloads—with gains that depend on traffic, cache retention, and network capacity.

From a talk by Yuchen Fama and Ashish Kamra

At a glance

Ideas worth remembering

  • Agentic inference evaluations need repeated prefixes, context variation, and sub-agent fan-out. Average steady-state throughput misses pressures that determine interactive latency.

  • Cache-aware placement balances prefix reuse against running and queued work. Retention and eviction policy determine whether useful state remains available for the next turn.

  • Prefill is compute-intensive; decode needs memory bandwidth and predictable token intervals. Separate pools reduce their interference by transferring computed KV state to the decoder.

  • P/D gains depend on workload, concurrency, pool sizing, and network fabric. Long input-heavy contexts and strict streaming requirements favor it; a fabric unsuited to KV transfer favors aggregated serving.

  • Independent scaling makes prefill and decode capacity separate decisions. The closing H200 case study targets prefill capacity for an input-heavy workload, with scheduler tuning and agent-program orchestration continuing upstream.

Agent sessions break the steady-state benchmark

A steady-state inference benchmark can tell you how fast a server handles a controlled stream of requests. An agent session brings a different problem: the same context returns across many turns, its length changes, and sub-agents suddenly add concurrent work. Red Hat’s Ashish Kamra, senior manager of performance engineering, and Yuchen Fama, product manager working with vLLM and llm-d, begin with that mismatch. Their two main tools are KV cache-aware routing and prefill/decode disaggregation.

Red Hat’s surrounding inference stack includes tools for benchmarking, model quantization, and speculative decoding. This session focuses on coordinating the model engine with distributed serving once a client controls the structure and timing of its context.

The agentic workloads examined in the talk have several properties that change capacity planning:

  • Long sessions: Interactions range from a few turns to 3,000 turns. Each request belongs to a continuing sequence.
  • Repeated input: Reused system prompts and tool definitions produce cache hit rates often exceeding 90%. Much of the incoming context has already been processed.
  • Input-heavy work: Input-to-output token ratios often exceed 100:1. Counting output tokens alone hides most of the context the server handles.
  • Uneven demand: Context lengths vary, and sub-agent fan-out complicates scheduling. Distributions and P90 values matter when sizing capacity; an average smooths away the larger requests and bursts.

Trace replay brings those pressures into evaluation. The collaboration with Google and IBM adds a replay tool to inference-perf to study the request patterns that agent sessions create. Repeated context and sudden concurrency occur together, so testing either in isolation misses part of the scheduling problem.

0:120:42
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

0:12 · section reference included

Find the cached prefix—and keep it worth finding

The KV cache holds state computed while processing a prompt. Reusing that state avoids repeating work on an unchanged prefix. But the client decides when context grows, changes, or returns. Frequent evictions and rewrites make cache management volatile, so engine-level caching needs help from scheduling and routing.

That shifts the performance target toward interactive latency and makes cached throughput worth measuring separately. The Anthropic pricing example presented in the talk has a 10× difference between cached and uncached input-token costs. It illustrates the economic value of reuse under that pricing scheme; it does not measure a 10× reduction in the cost of running an llm-d cluster.

The llm-d router’s endpoint picker plugin, or EPP, scores candidate pods using both load and cache information. It probes running and waiting requests, KV cache utilization, and prefix-cache availability. Placement balances the chance of a cache hit against the work already waiting on a pod. A cached prefix offers less recomputation, while a shorter queue offers less waiting; the picker needs both signals to choose an endpoint.

Routing and retention address different parts of the problem:

  • Cache locality: The endpoint picker sends a request toward a pod that can reuse its prefix while accounting for load.
  • Offloading tiers: The described development work explores NVMe/SSD, filesystems such as Ceph, and KV-oriented stores such as Mooncake to retain hot, warm, and cold session state.
  • Session-aware eviction: Priority and session pinning aim to keep important context available when a session needs it again.

A router can only find reusable state that the cache system has retained. Additional tiers create room to preserve it; eviction policy decides which state remains available.

5:015:32
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

5:01 · section reference included

Four turns make cache locality visible

The routing demo follows one concrete change: whether the system prompt stays the same. The first request takes roughly three seconds and populates the cache. There is no hit because this is the opening turn. The second turn keeps the system prompt, returns to the same pod address, and reuses the cache. Its displayed duration falls to about one second.

The third turn changes the system prompt. It has no cache hit, lands on a different pod, and takes roughly three seconds again. The following turn changes only the user prompt while retaining this new system prompt. Reuse returns, and the duration falls to about one second. The causal steps are an unchanged prefix, available cached state, placement that reaches that state, and less repeated prompt processing. The pod address makes placement visible; the reported cache hit establishes reuse.

What changes across these four turns? The comparison below puts prompt continuity beside cache behavior and duration. It shows two cold-to-warm transitions. Changing the system prompt creates a new starting point, while changing the user prompt can leave the shared prefix reusable.

Prefix-aware routing primarily helps the wait before generation begins—time to first token, or TTFT—and can improve throughput by avoiding repeated work. Once output starts streaming, another measure matters: inter-token latency, or ITL, the time between generated tokens. Faster startup does not by itself prevent a stream from stalling.

Compare the ideasTwo cold-to-warm cache transitions

No cache hit; populates KV cache; roughly 3 seconds.

The demo reports roughly three seconds without reuse and one second with reuse. These are displayed request durations, rather than a controlled latency distribution.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

7:41 · section reference included

Separate prompt processing from the active stream

llm-d supplies the distributed control plane around model execution. Alongside the router, workload APIs such as LeaderWorkerSet and disaggregated sets organize multi-node execution. Autoscalers respond to capacity and traffic so pods can scale independently. This coordination becomes especially useful when prefill and decode run in separate pools.

Aggregated serving asks one pod to optimize both startup and streaming. The two execution phases have different needs:

  • Prefill: Processes the initial prompt and builds its KV cache. It is compute-intensive, bursty, and benefits from large-batch parallelism.
  • Decode: Generates one token at a time. It needs memory bandwidth, resident KV state, and predictable latency.

When both share a GPU, a long incoming prompt can interrupt ongoing token generation. This phase interference appears to a user as pauses and jitter in a stream that had already started.

In the P/D request flow, the gateway router evaluates cluster state through the endpoint picker and selects a prefill worker and a decode worker. It coordinates prompt processing with the prefill worker, which builds the initial cache and produces KV transfer metadata. The selected decoder uses that metadata to pull the computed cache across the network. Prefill and decode now run on independently scalable inference pods.

Where does the state move when execution is split? The flow below separates placement decisions from KV movement. The decoder pulls the result of prompt processing instead of rebuilding it. Separating the GPU work can protect streaming, but moving its state introduces a network requirement that aggregated serving avoids.

How it fits togetherPlacement first, then KV transfer

Arrives at the gateway router.

The router selects both workers. Prefill builds prompt state; decode pulls that state using transfer metadata and generates tokens.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

9:29 · section reference included

Smoother streaming, with gains that depend on load

The first latency comparison reports P99 ITL near 900 milliseconds for aggregated serving and about 100 milliseconds for disaggregated serving. P99 describes the slower tail of token intervals. The disaggregated curve is also smoother, so the improvement concerns both the length and variability of streaming pauses.

The following internal Red Hat experiment uses GPT-OSS 120B on 16 H100s. Aggregated serving has four replicas with tensor parallelism four; disaggregated serving has two prefill and two decode workers, each also using tensor parallelism four. The highly multi-turn workload has a 10,000-token prefix and 128 tokens per turn.

Its comparison separates three choices: aggregated serving with default Kubernetes scheduling, aggregated serving with llm-d cache-aware routing, and disaggregated serving. Routing alone produces gains. P/D is similar to aggregated serving at low and high concurrency, with its strongest advantage in the middle band. That band has no numeric threshold stated here; finding the favorable range requires a load curve for the deployment being sized.

A separate prefill-heavy test uses GPT-OSS 120B on 64 H100s: eight aggregated replicas at tensor parallelism eight, compared with three prefill and five decode workers at the same parallelism. Average input length is 5,000 tokens and output length is 500 tokens. Here, the presented P/D Pareto curve dominates the aggregated curve across the tested interactivity range. A Pareto comparison asks which configuration offers a better combination of throughput and responsiveness. The different result reinforces that workload and pool sizing determine the benefit.

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

12:56 · section reference included

Choose P/D for the bottleneck and fabric you have

P/D makes a phase-separation tradeoff. Its appeal grows when long contexts, high input-to-output ratios, and strict streaming requirements make prompt processing interfere with generation. Large models with useful model-parallel execution options and the favorable middle concurrency regime are further reasons to consider it.

The network belongs in that decision. The talk recommends a high-speed fabric such as RDMA or RoCE to support prefill-to-decode KV transfer. Without a suitable fabric, stay aggregated. Aggregated serving is also a reasonable choice for short or moderate contexts, low concurrency, or workloads whose main requirement is TTFT, which can be tuned within an aggregated deployment.

Once P/D is deployed, the prefill-to-decode ratio becomes another variable to manage. A static starting ratio needs to evolve as traffic changes. The scheduler considers service-level objectives, queue depths, cache locality, pool ratios, and network topology together to choose placement. Independent autoscaling and adjustments to tensor and data parallelism let the system add capacity where work is accumulating.

This connects the two techniques: routing reduces unnecessary prompt processing, while disaggregation lets necessary prompt processing run apart from active generation. Their combination still needs rate matching between the pools. The amount of prefill and decode capacity should follow the changing traffic mix and the latency objectives.

15:3716:07
Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

15:37 · section reference included

Bring the knobs together for GLM-5.2 on H200s

The closing case study starts with a hardware constraint: customers have H200s, even when impressive published results use B200s. Serving GLM-5.2 well on that existing hardware requires combining cache-aware routing, P/D separation, and parallelism. The companion technical article develops the workload and deployment pattern in more detail.

The design described in the recording supports up to three prefill workers optimized for throughput and one dedicated decode worker optimized for low latency. Each worker uses a LeaderWorkerSet group with tensor parallelism one, data parallelism eight, and expert parallelism eight. NIXL, identified in the companion article, handles KV movement between the pools. More prefill workers can be added without reconfiguring the decoder, so additional input-processing capacity does not require rebuilding both sides.

One tentative finding complicates the cache-format choice: Yuchen reports that BF16 KV cache was faster than another cache format for longer prefill. The comparison format is not established clearly enough to name, and the team was still investigating the result. For this case study, cache precision remained a performance choice to explore rather than a settled prescription.

In the closing case study, Yuchen reports preliminary results of 4× faster TTFT and 60% more requests with two prefill workers and one decode worker on an agentic dataset with a 45:1 input-to-output ratio, where prefill is the constraint. These are the speaker’s reported figures, not an established controlled P/D-versus-aggregate improvement: the comparison baseline is unspecified. The architecture supports up to three prefill workers, a design limit separate from the two-prefill, one-decode configuration reported for these results. The companion article’s separate experiments cannot supply the missing comparison baseline for these recording results.

The next steps follow the remaining bottleneck: tune the upper-layer scheduler to reduce TTFT and add more prefill replicas. The broader development direction includes session-graph orchestration, program-aware scheduling, state-reuse lifecycle management, and agentic benchmarks. These continuing efforts extend scheduling toward a connected agent program, whose turns and sub-agents share state and dependencies.

The closing invitation is to build that work upstream through llm-d and its special-interest groups, alongside collaborators including CoreWeave, Google, IBM, and Nvidia. The H200 serving guide provides a practical deployment path with configurable routing, offloading, and pool layouts. It gives the closing question something concrete to work with: how much reusable context and prefill capacity does the workload need before its next stream can begin?

Suggest correction

This note stays in this page until you copy or download it. Nothing is submitted; reloading clears the draft.

18:02 · section reference included

Resources

  • Expands the closing case study with workload distributions, cache capacity, offloading, pool-sizing experiments, and limits on performance comparisons.

  • Provides deployment manifests and composable options for prefix-aware routing, KV offloading, multi-token prediction, and prefill/decode layouts.

Read the complete timestamped transcript
  1. 0:12

    All right. Um, welcome everyone to yet

  2. 0:15

    another inference talk. I hope you have

  3. 0:18

    had a good conference so far. And u, so

  4. 0:22

    in this session, I mean I'm sure you

  5. 0:24

    people who have been in the room uh must

  6. 0:26

    have heard these terms many times by

  7. 0:28

    now. So we're going to do a little bit

  8. 0:30

    more deep dive into the challenges of

  9. 0:32

    LLM deployments for agentic workloads

  10. 0:35

    and uh in this session we'll focus

  11. 0:37

    specifically on KV cache away routing

  12. 0:39

    and uh PD disagregation

  13. 0:43

    um and also you know when you when you

  14. 0:45

    look at public inference uh benchmark

  15. 0:48

    results you are typically looking at

  16. 0:49

    very steady state isolated highly

  17. 0:51

    sanitized numbers and what those

  18. 0:53

    benchmarks actually don't show you u is

  19. 0:56

    the chaotic reality of multi-turn

  20. 0:59

    interactions, massive context

  21. 1:00

    fluctuations which are very typical of

  22. 1:03

    agentic workloads. So we'll also try to

  23. 1:06

    pull the curtain back on some of those

  24. 1:08

    complexities. Um by by way of

  25. 1:10

    introduction uh my name is Ashish Kamra.

  26. 1:13

    I'm a senior manager of performance

  27. 1:15

    engineering at Red Hat. And with me

  28. 1:18

    >> hi I'm Yuch Chen. I'm the product

  29. 1:20

    manager at Red Hat Inference working

  30. 1:21

    closely with VLM and AMD core

  31. 1:24

    maintainers. also a contributor myself.

  32. 1:28

    >> So here is the agenda for the next 20

  33. 1:30

    minutes or so. Um Euchen will start with

  34. 1:33

    an analysis of inference behavior in the

  35. 1:36

    agentic era and some of the core

  36. 1:38

    characteristics and challenges. Uh next

  37. 1:41

    we next you will walk us through the KV

  38. 1:44

    cache um utilization and management

  39. 1:47

    strategies.

  40. 1:48

    I will break down the mechanics of

  41. 1:50

    pre-fill decode disagregation and walk

  42. 1:52

    you through some some results and then

  43. 1:55

    Euchen will again bring it all back

  44. 1:57

    together with our ongoing case study on

  45. 1:59

    our favorite open coding model GLM 5.2.

  46. 2:03

    Um and just a couple of uh sources from

  47. 2:06

    our side if you are more interested in

  48. 2:08

    learning more about open source

  49. 2:10

    inference we have a free course free

  50. 2:12

    course on deep learning.ai AI uh by

  51. 2:15

    Cedric and with Andrew Ning. Um and the

  52. 2:17

    other is a series of blogs on the Red

  53. 2:19

    Hat developer portal on distributed

  54. 2:22

    inference concepts uh troubleshooting

  55. 2:24

    and deployment patterns.

  56. 2:27

    Uh and for those who may not be aware

  57. 2:30

    since Red Hat is better known as the

  58. 2:32

    Linux company for enterprise Linux and

  59. 2:36

    uh the Kubernetes company for Open Shift

  60. 2:39

    uh but more recently we are also a major

  61. 2:41

    player in open source AI inference with

  62. 2:45

    uh us being the top contributor in VLM

  63. 2:47

    LLMD and the case of projects and also

  64. 2:51

    uh having incubated guide LLM for

  65. 2:54

    benchmarking LLM compressor for model

  66. 2:56

    quantization and speculators for uh

  67. 3:00

    speculative uh decoding models and we

  68. 3:04

    also bring it bring all of that together

  69. 3:05

    in a optimized model hub on hugging

  70. 3:09

    phase under the Red Hat AI arc.

  71. 3:12

    Um and we are also building the platform

  72. 3:14

    for the next wave of agentic inference

  73. 3:16

    workloads and with that I will hand over

  74. 3:18

    to you to uh walk you through more of

  75. 3:21

    it.

  76. 3:25

    So we are currently um at this

  77. 3:27

    inflection point moving from the era of

  78. 3:29

    classic inference to the agentic era. So

  79. 3:33

    when we look at the real world agentic

  80. 3:36

    work workloads such as uh sweet bench

  81. 3:38

    and also watrices from real world cloud

  82. 3:41

    code sessions they fundamentally break

  83. 3:44

    many assumptions we made with classic LM

  84. 3:46

    serving. uh as you heard actually many

  85. 3:48

    times in previous sessions for example

  86. 3:50

    multi-turns and new standard we found

  87. 3:52

    from a few turns all the way to 3,000

  88. 3:55

    turns and also because agent frequently

  89. 3:57

    reuse the uh system prompt and the total

  90. 4:00

    definitions we usually see super high

  91. 4:02

    cash hit rate um oftentimes well

  92. 4:05

    exceeding 90%. Uh another thing is input

  93. 4:08

    output ratio is uh is massive oftentimes

  94. 4:12

    over a 100 ratio and even higher and in

  95. 4:15

    many cases and on top of that the

  96. 4:17

    context management is is incredibly

  97. 4:20

    complex due to this high variance

  98. 4:22

    because we can't just simply take the

  99. 4:24

    average and oftentimes we need to look

  100. 4:27

    at the distributions and the P90 numbers

  101. 4:29

    especially when you do uh capacity

  102. 4:31

    planning and also we observe really

  103. 4:34

    interesting patterns like sub Asian

  104. 4:36

    panel which is which further complex uh

  105. 4:39

    complicates scheduling. So to help

  106. 4:41

    communities study um this patterns we

  107. 4:44

    collaborate with Google thank you and

  108. 4:46

    also IBM our parent company to add uh a

  109. 4:49

    a trace replay tool in the inference

  110. 4:51

    perf you heard from earlier sessions u

  111. 4:54

    from Ashoken and Jason. Um so yeah feel

  112. 4:56

    free to check it out and the link is

  113. 4:58

    here.

  114. 5:01

    Uh next slide. Oh, so transition from

  115. 5:04

    the class uh the characteristics um we

  116. 5:06

    just saw for agentic workloads. We're no

  117. 5:09

    longer chasing this um this this raw

  118. 5:12

    throughput in a steady state. We often

  119. 5:14

    need to optimize uh for example

  120. 5:16

    interactive latency and they're very um

  121. 5:19

    highly volatile and client-driven

  122. 5:21

    context because user and you know client

  123. 5:24

    define the prompt structure. So this

  124. 5:26

    introduced several critical challenges.

  125. 5:28

    First of all, KV cache management

  126. 5:30

    becomes super volatile because the

  127. 5:32

    context is client determined as I said.

  128. 5:34

    So oftentimes we face this like you know

  129. 5:37

    frequent evictions and rewrites and

  130. 5:40

    secondly we also need to tune um the

  131. 5:43

    engine like VM with upper layer uh

  132. 5:45

    scheduling and routing.

  133. 5:47

    It needs that coordination such as

  134. 5:49

    prefix routing especially when latency

  135. 5:52

    becomes a primary uh scheduling matrix

  136. 5:54

    rather than like a secondary or

  137. 5:56

    afterthought. And thirdly, we also need

  138. 5:58

    to rethink our metrics. For example, we

  139. 6:00

    need to measure cats throughput

  140. 6:02

    separately. Why? Because on the right,

  141. 6:04

    it's really clear that economic stakes

  142. 6:07

    is very high. So, this is the uh

  143. 6:09

    anthropic API pricing. You also heard

  144. 6:11

    from earlier sessions. There's 10x cost

  145. 6:13

    difference between cash and non-cash

  146. 6:15

    tokens. So, 10x difference on your um

  147. 6:18

    token balance sheet is is pretty serious

  148. 6:20

    impact on your business.

  149. 6:23

    So next let's let's look at how the KV

  150. 6:25

    cache is um both utilized and managed in

  151. 6:28

    LMD. So LMD router has this really

  152. 6:31

    flexible um endpoint picker plugins we

  153. 6:34

    call the EP that can route the request

  154. 6:36

    to the optimal pods and that meet the KV

  155. 6:38

    cache locality and also the load

  156. 6:41

    criteria. So the EP continue probe each

  157. 6:44

    pods like VM pod matrix to score each

  158. 6:47

    pod on like the running for example

  159. 6:49

    running and waiting request and then the

  160. 6:51

    KV cache utilization also prefix uh

  161. 6:54

    cache availability and so we can

  162. 6:56

    schedule requests to the optimal pod

  163. 6:58

    with the lowest load and also highest

  164. 7:00

    possibility to um to of a cache hit. So

  165. 7:03

    um going down from to the KV cache

  166. 7:06

    management layer actually you also heard

  167. 7:07

    from earlier session right before this.

  168. 7:10

    So for agentic sessions when you have u

  169. 7:12

    hot warm and cold cache our current

  170. 7:15

    effort focus on for example um more

  171. 7:17

    offloading tiers like NVME SSD and also

  172. 7:20

    uh file system XF along with KV ccentric

  173. 7:23

    store um like uh moon cake and also

  174. 7:26

    implementing smarter and session a wire

  175. 7:28

    eviction policies such as priority and

  176. 7:30

    also session pinning to uh ensure this

  177. 7:33

    uh really important you know the the

  178. 7:35

    context persists exactly when and where

  179. 7:38

    it's needed.

  180. 7:41

    So, I'm gonna play this um video really

  181. 7:44

    quick. Uh it's a it's a short demo.

  182. 7:46

    >> Stand here so you can look at it.

  183. 7:48

    >> Okay.

  184. 7:52

    So,

  185. 7:54

    okay. So, this is a example of a KV

  186. 7:56

    cache bar routing. As you see, when we

  187. 7:58

    send the very first request and it

  188. 8:00

    populate the KV cache, it takes roughly

  189. 8:02

    3 seconds. And when we actually look at

  190. 8:06

    where it's you know the KV cache uh is

  191. 8:08

    going there's no KV cache hit because

  192. 8:10

    it's the very first turn. And then when

  193. 8:12

    we have the second turn the request

  194. 8:13

    actually reuse a KV cache because as you

  195. 8:16

    see the system prompt is the same and

  196. 8:18

    this time takes about one seconds. And

  197. 8:20

    then when you actually look at the uh

  198. 8:21

    pod address exactly the same because we

  199. 8:24

    define the KV cache. Now going to the

  200. 8:26

    third turn a new request with different

  201. 8:28

    system prompt. Now it takes about three

  202. 8:30

    uh seconds and as you see you know right

  203. 8:34

    now and we don't find any KV cache here

  204. 8:36

    because you can tell it's different pod

  205. 8:38

    address and then if you just change the

  206. 8:41

    user prompt and keep the same system

  207. 8:43

    prompt and the next turn you you reuse

  208. 8:46

    the KB cache and in this in this time it

  209. 8:48

    takes roughly about uh one second. Yeah.

  210. 8:51

    So it's a pretty intuitive demo and um

  211. 8:53

    I'll turn it to Ashish to talk about the

  212. 8:56

    next side but before that what does

  213. 8:57

    problem does it solve? So often times

  214. 8:59

    the prefix routing KB cache routing

  215. 9:02

    helps you solve the TTFD problem and of

  216. 9:04

    course you'll improve your lat uh your

  217. 9:06

    your throughput but oftentimes for

  218. 9:08

    agentic workload is not just a TTFT your

  219. 9:10

    throughput is about your inter token

  220. 9:12

    latency how do we solve that so preview

  221. 9:15

    decode disagregation is a really uh

  222. 9:17

    powerful technique but there are times

  223. 9:19

    there work at times it doesn't work so

  224. 9:21

    I'll turn it to Ashish to give you a

  225. 9:23

    preview of um of the PD uh disregation

  226. 9:28

    So before we dive into PD, let's just uh

  227. 9:31

    look at what LLMD is. So LLMD is a high

  228. 9:34

    performance Kubernetes native and

  229. 9:36

    actually now works on non-cubernetes

  230. 9:38

    environments as well. Distributed LM LLM

  231. 9:41

    inference framework hosted under the

  232. 9:43

    CNCF umbrella. LLMD provides a unified

  233. 9:47

    intelligent control plane designed

  234. 9:49

    specifically for agentic era of

  235. 9:50

    inference workloads. Well, Euchin

  236. 9:52

    already talked about the router and the

  237. 9:54

    EP at the top of the slide. Um, the

  238. 9:58

    other aspects are workload APIs such as

  239. 10:00

    leader worker set and disagregated set

  240. 10:02

    that orchestrates complex multi-

  241. 10:05

    multi-node model execution and then

  242. 10:08

    autoscalers that monitors capacity

  243. 10:10

    bounds and real-time traffic mixes to

  244. 10:13

    independently scale up and scale down uh

  245. 10:16

    your pods depending on the system load.

  246. 10:19

    So now look now let's look at uh prefill

  247. 10:21

    decode disagregation in detail. Um uh

  248. 10:25

    okay so why does PD exist in the first

  249. 10:29

    place? So one of the most powerful

  250. 10:30

    patterns implemented by LLMD is prefill

  251. 10:33

    decode disagregation and you must have

  252. 10:35

    heard from some of the previous talks as

  253. 10:37

    well. So what happens is in in a nonPD

  254. 10:40

    situation in aggregated serving one pod

  255. 10:43

    is responsible for optimizing both your

  256. 10:46

    time to first token and your inter token

  257. 10:48

    latencies. Uh but in PD prefill and

  258. 10:51

    decode become independently scalable

  259. 10:53

    inference pods. But to understand why we

  260. 10:55

    actually need this we have to look at

  261. 10:58

    the physics of LLM execution.

  262. 11:00

    colloccating uh both prefill and decode

  263. 11:03

    tasks on the same GPU creates something

  264. 11:05

    called as phase interference. Prefill

  265. 11:08

    phase is the phase that creates the KV

  266. 11:10

    caches for your initial prompt. It wants

  267. 11:13

    high compute. It's highly bursty uh

  268. 11:17

    utilizes GPUs at uh high flops and and

  269. 11:21

    thrives on large batch parallelism to

  270. 11:24

    process the prompts and builds the

  271. 11:25

    initial KV cache. The decode phase on

  272. 11:28

    the other hand is generating one token

  273. 11:30

    at a time and it's more me memory

  274. 11:32

    bandwidth hungry. It's highly latency

  275. 11:34

    sensitive and requires high heavy cache

  276. 11:37

    residency. So in a in a in a traditional

  277. 11:40

    aggregated pod if you if there's a

  278. 11:43

    sudden influx of a long prefilled palm,

  279. 11:46

    it will completely stall the ongoing

  280. 11:48

    decode token generation process causing

  281. 11:50

    massive problems and jitter in user

  282. 11:53

    streaming latency.

  283. 11:56

    So, so how does PD actually work in

  284. 11:58

    practice in LMD? So, LNMD uses um uh you

  285. 12:03

    know like okay, we'll start with step

  286. 12:05

    one. A incoming request hits the gateway

  287. 12:07

    router which dynamically evaluates

  288. 12:10

    cluster states using something known as

  289. 12:12

    the endpoint picker you talked about and

  290. 12:15

    schedules the request to use PD

  291. 12:17

    disagregation selecting the optimal

  292. 12:19

    prefill and decode workers. The router

  293. 12:22

    then coordinates the transaction

  294. 12:23

    directly with the designated pre-fill

  295. 12:25

    worker. The pre-fill worker processes

  296. 12:27

    the prompt, construct the initial KV

  297. 12:29

    cache of the prompt and outputs the

  298. 12:32

    standard KV transfer metadata. Um, and

  299. 12:35

    the target decode worker actually pulls

  300. 12:37

    the computed KV caches um, uh, across

  301. 12:41

    the network fab fabric utilizing uh, the

  302. 12:44

    KV transfer metadata that the uh, uh,

  303. 12:47

    prefill pod had generated. Um okay so

  304. 12:52

    with that yes that's kind of how uh PD

  305. 12:55

    is implemented in practice in LMD and

  306. 12:57

    next I would like to show you some uh

  307. 12:59

    experimental results on where PD

  308. 13:01

    actually shines. So in this graph you

  309. 13:03

    can see that um

  310. 13:06

    uh in in in the standard aggregated

  311. 13:08

    deployment which is the top red line uh

  312. 13:11

    the P99 ITL uh hovers roughly around 900

  313. 13:15

    milliseconds and you can you can see

  314. 13:17

    some fluctuations um up and down and but

  315. 13:22

    the the bottom blue line is the P99 uh

  316. 13:26

    inter token latency on a PD deployment

  317. 13:28

    and you can see that it's drastically

  318. 13:30

    almost nine times better at 100

  319. 13:32

    millconds and it's also much smoother uh

  320. 13:35

    than the aggregated serving

  321. 13:40

    and uh this is some of our own internal

  322. 13:43

    results at Red Hat. So for a GPOSS 12B

  323. 13:46

    model uh 16 H100s

  324. 13:49

    uh the aggregated config is four

  325. 13:52

    replicas tensor parallelism 4 and the

  326. 13:54

    disagregated is two prefilled 2D code

  327. 13:56

    all with tensor parallelism 4. It's a

  328. 13:58

    highly multi-turn workload with a 10,000

  329. 14:01

    token prefix and 128 tokens for every

  330. 14:05

    turn every turn. So, so this is a great

  331. 14:08

    chart like you can see at the bottom

  332. 14:09

    most line is a standard aggregated

  333. 14:11

    config that's uh is doing the default

  334. 14:14

    Kubernetes scheduling and uh and it's

  335. 14:17

    aggregated. So that's kind of our

  336. 14:18

    baseline and then the middle blue line

  337. 14:21

    is still aggregated but with the LLMD uh

  338. 14:25

    KV cache aware routing and you can

  339. 14:27

    almost see the gains just just based on

  340. 14:29

    the routing and the red line is actually

  341. 14:31

    the PD uh the pre-fill decode config

  342. 14:35

    with two pre-fill and two decode workers

  343. 14:37

    and you can actually see that like it's

  344. 14:39

    very similar to the aggregated config at

  345. 14:41

    the lower concurrency regimes and uh

  346. 14:44

    even and and very similar at the higher

  347. 14:45

    concurrency regimes but it's actually

  348. 14:47

    the middle part of the concurrency

  349. 14:49

    regime that PD actually shines

  350. 14:53

    and and these are some of the the

  351. 14:56

    classic parita curves that we see when

  352. 14:58

    you actually do PD and uh aggregated

  353. 15:01

    side by side. So these results are again

  354. 15:03

    from the GPTOSS 12B model 64 H100s

  355. 15:07

    aggregated is eight replicas TP8 and

  356. 15:10

    this a is uh three prefilled 5D code

  357. 15:13

    again TP8 and a pre-filled heavy

  358. 15:15

    workload with like 5,000 average input

  359. 15:18

    sequence length and 500 output sequence

  360. 15:20

    length and you can actually see the blue

  361. 15:22

    line is the the PD curve and the red

  362. 15:24

    line is the aggregated curve and the PD

  363. 15:26

    curve kind of dominates um uh the

  364. 15:30

    aggregate curve across the entire

  365. 15:32

    interactivity spectrum.

  366. 15:37

    Okay, but I don't want to leave you guys

  367. 15:39

    that PD is the answer to everything and

  368. 15:40

    it's a magic bullet. But um it's uh it's

  369. 15:43

    essentially a separation phase

  370. 15:45

    separation trade-off and not a magic

  371. 15:46

    bullet. So we created this uh matrix to

  372. 15:49

    help you decide when PD might be uh good

  373. 15:53

    for you. So if you're managing long

  374. 15:55

    context uh with high ISL OSL ratios and

  375. 16:00

    you if you have a large model that

  376. 16:01

    you're serving that can that you can

  377. 16:04

    apply rich model parallelism techniques

  378. 16:07

    um you're facing that middle concurrency

  379. 16:09

    regime uh that I I showed you in the

  380. 16:11

    previous graphs and and the very

  381. 16:14

    important part is that if you want uh

  382. 16:16

    strict ITL streaming requirements like

  383. 16:18

    you want the you want the token

  384. 16:19

    generation to be uh much more smooth um

  385. 16:22

    then you want to consider PD but we also

  386. 16:25

    saw that it requires transfer of KV

  387. 16:27

    caches from your pre-filled workers to

  388. 16:29

    your decode workers. So you must pro

  389. 16:31

    process an advanced uh high-sp speeded

  390. 16:33

    network fabric like uh RDMMA or rocky to

  391. 16:37

    support that KV cache transfer. And if

  392. 16:39

    you do not have such requirements, short

  393. 16:41

    moderate context, any model size, low

  394. 16:46

    concurrency regimes or uh if you have

  395. 16:48

    strict TTF requirements because you can

  396. 16:50

    actually tune them on an aggregate

  397. 16:52

    serving um and you the biggest point is

  398. 16:55

    like if you don't have the network

  399. 16:56

    fabric to support those KV cache

  400. 16:57

    transfers. So you might actually just

  401. 16:59

    want to stick with aggregated.

  402. 17:02

    So here is my key takeaway from all of

  403. 17:04

    this. So architecting this complex

  404. 17:06

    platform requires balancing a lot of u

  405. 17:09

    knobs and a highly multi-dimensional

  406. 17:11

    design space all of which is supported

  407. 17:13

    in LLMD. As you saw the scheduler must

  408. 17:16

    support or constantly evaluate SLO

  409. 17:18

    targets uh QEPs KV cache locality

  410. 17:21

    metrics PD ratios and network topologies

  411. 17:24

    to be able to route the request to the

  412. 17:26

    optimal FOD. While in while the PD

  413. 17:29

    design space you you need dynamic PD

  414. 17:32

    rate matching to adapt to PD ratios

  415. 17:35

    because you know you can start with a

  416. 17:36

    static PD ratio but it needs to evolve

  417. 17:38

    with the autoscaler as the traffic

  418. 17:40

    changes um and you need uh yeah

  419. 17:44

    autoscaling to scale PD pools

  420. 17:46

    independently

  421. 17:47

    um and constantly tweaking model

  422. 17:50

    parallelism techniques like tensor

  423. 17:51

    parallelism data parallelism uh to meet

  424. 17:54

    your SLOs's.

  425. 17:56

    So um I think with these uh I will hand

  426. 17:59

    it over to Euchen to anchor some of the

  427. 18:01

    concepts that we showed with the real

  428. 18:04

    world case study of serving the GLM 5.2

  429. 18:07

    model uh which is uh still ongoing as we

  430. 18:10

    speak.

  431. 18:11

    >> Yeah, still ongoing. You probably have

  432. 18:13

    seen tons of uh impressive numbers of

  433. 18:16

    GLM 5.2 on B200 when we talk to our

  434. 18:19

    customers and they usually don't have

  435. 18:21

    you know the luxury of B200. They have a

  436. 18:23

    lot of H200. So we have to figure out

  437. 18:25

    how to like put all the knobs together

  438. 18:27

    and make GM 5.2 work really well for

  439. 18:30

    cluster of of H200. So uh we let's

  440. 18:33

    anchor all the concept together. Um we

  441. 18:35

    went through for example the uh KV cache

  442. 18:38

    routing PD disagregation. We kind of

  443. 18:40

    call them a wildl path in LMD and also

  444. 18:43

    we combine with different parallelism

  445. 18:45

    strategies to so we can uh independently

  446. 18:48

    uh scale prefuel paths because for

  447. 18:50

    agentic workload is super uh long you

  448. 18:52

    know like heavy prefill. So uh in this

  449. 18:55

    case we designed the prefuel pool using

  450. 18:57

    up to three workers optimized for uh

  451. 18:59

    high throughput uh with deep and then

  452. 19:02

    for decoup we use uh one dedicated

  453. 19:04

    worker and um that's optimized for for

  454. 19:06

    low latency. So we use Nixo for

  455. 19:08

    efficient KV transfer between the pools

  456. 19:11

    and also with the each worker we have

  457. 19:13

    the leader worker set group uh with TP1

  458. 19:15

    DP8 and also uh EP8 uh expert

  459. 19:18

    parallelism 8. So the architecture is

  460. 19:21

    just highly modular because you can uh

  461. 19:23

    actually scale the throughput by simply

  462. 19:24

    adding uh preview workers without

  463. 19:26

    reconfiguring and um the decoup. So uh

  464. 19:30

    this highlights how AMD effectly

  465. 19:32

    effectively managed the complexity of

  466. 19:34

    combining like PB and DB and EPI scale.

  467. 19:38

    And also we found some interesting fun

  468. 19:39

    fact actually a couple days ago. Um B B

  469. 19:42

    B B B B B B B B B B B B B B B B B B B B

  470. 19:42

    B B B B B B B B B B B B B B B B B B Bf6

  471. 19:43

    uh BF16 KV cache actually is faster than

  472. 19:46

    using like FPA uh KV cache for longer

  473. 19:49

    preview. Um this is also like we

  474. 19:51

    continue to explore and found like more

  475. 19:53

    interesting patterns, but more

  476. 19:55

    importantly uh we want to kind of just

  477. 19:57

    show the result really quick. So um for

  478. 19:59

    this uh data set agentic workload data

  479. 20:01

    set the ISO OSL ratio is pretty high 45

  480. 20:04

    to1 ratio preview is uh is really the

  481. 20:07

    constraint you can tell um with 2P even

  482. 20:09

    1D we have um 4x passer TDFT and also 60

  483. 20:14

    uh% more requests and this is continuous

  484. 20:17

    like work in progress so the next step

  485. 20:19

    is we need to also put the upper layer

  486. 20:22

    lower TTFT and also adding more more

  487. 20:24

    preview replicas so um I know we're

  488. 20:27

    running out of time really quick. Uh we

  489. 20:30

    um the fundamental shift for agentic

  490. 20:31

    workload we're continuing to uh have

  491. 20:34

    this um uh agentic north uh northstar uh

  492. 20:37

    with session graph orchestration program

  493. 20:39

    award scheduling uh state reuse life

  494. 20:42

    cycle and also the uh agentic benchmark

  495. 20:45

    um we're working on. So you can find

  496. 20:47

    them uh in AMD upstream AMD and also you

  497. 20:51

    know feel free to join the SIG group and

  498. 20:54

    uh and contribute and um this is the

  499. 20:57

    very last slide. So distri distributed

  500. 20:59

    inference is not challenge uh every

  501. 21:00

    single comp a single company can solve

  502. 21:02

    along. We're proud to be uh building

  503. 21:05

    this uh future in the open alongside our

  504. 21:07

    incredible ecosystem collaborators uh

  505. 21:10

    core wave Google IBM Nvidia growing list

  506. 21:13

    of launch partners and industry

  507. 21:15

    adopters. So if you're passionate about

  508. 21:17

    the future of opensource inference, we

  509. 21:19

    invite you to join us. We do have a

  510. 21:21

    booth downstairs. Feel free to stop by,

  511. 21:23

    ask us any questions. And uh thank you

  512. 21:25

    so much for your time.

  513. 21:27

    [applause]

  514. 21:44

    >> [music]