0 views

Greg Pereira, Senior Machine Learning Engineer at Red Hat, explains why one routing decision isn’t enough: prefill cares deeply about landing on a pod that already holds its prefix, so it only has to compute the delta. Decode doesn’t care at all, since it can pull KV from any prefiller. That’s why llm-d runs separate scheduling profiles for each phase.
From vLLM Office Hours #55: Mooncake + vLLM/llm-d Deep Dive. Full session: https://youtu.be/RFyeBEy1AP8
Join us live every other Thursday: https://red.ht/office-hours
#vLLM #llm-d #KVCache #AIInference #RedHat #MLOps
Date: September 17, 2026











