sebae banner ad-300x250
sebae intro coupon 30 off
sebae banner 728x900
sebae banner 300x250

When you should NOT use prefill/decode disaggregation

0 views
0%

When you should NOT use prefill/decode disaggregation

Everyone’s talking about PD disaggregation. Greg Pereira, Senior Machine Learning Engineer at Red Hat, explains when it actually costs you: 10 uncached input tokens and 1,000 output tokens means opening a KV transfer between two GPUs is more work than just running it on one. That’s why llm-d ships a prefix-based PD decider.

From vLLM Office Hours #55: Mooncake + vLLM/llm-d Deep Dive. Full session: https://youtu.be/RFyeBEy1AP8

Join us live every other Thursday: https://red.ht/office-hours

#vLLM #llm-d #KVCache #AIInference #RedHat #MLOps

Date: September 18, 2026