0 views

Everyone’s talking about PD disaggregation. Greg Pereira, Senior Machine Learning Engineer at Red Hat, explains when it actually costs you: 10 uncached input tokens and 1,000 output tokens means opening a KV transfer between two GPUs is more work than just running it on one. That’s why llm-d ships a prefix-based PD decider.
From vLLM Office Hours #55: Mooncake + vLLM/llm-d Deep Dive. Full session: https://youtu.be/RFyeBEy1AP8
Join us live every other Thursday: https://red.ht/office-hours
#vLLM #llm-d #KVCache #AIInference #RedHat #MLOps
Date: September 18, 2026











