0 views

Greg Pereira, Senior Machine Learning Engineer at Red Hat, discusses the scale-out problem: add a second vLLM replica behind a traditional load balancer and you drop to roughly 50% KV cache reuse. It only gets worse from there.
From vLLM Office Hours #55: Mooncake + vLLM/llm-d Deep Dive. Full session: https://youtu.be/RFyeBEy1AP8
Join us live every other Thursday: https://red.ht/office-hours
#vLLM #llm-d #KVCache #AIInference #RedHat #MLOps
Date: September 21, 2026









![[vLLM Office Hours #38] vLLM 2025 Retrospective & 2026 Roadmap – December 18, 2025](https://videos.sebae.net/wp-content/uploads/2025/12/hqdefault-509.jpg)

