![[vLLM Office Hours #58] - Intro to llm-d Semantic Classifier - October 1, 2026](https://i2.ytimg.com/vi/icrWIlo4QKc/hqdefault.jpg)
Welcome to vLLM office hours! These bi-weekly sessions are your chance to stay current with the vLLM ecosystem, ask questions, and hear directly from contributors and power users.
This week’s special topic: Intro to llm-d Semantic Classifier.
vLLM project update from core maintainer Michael Goin, covering a new hardware backend from Tenstorrent, HiSparse for sparse-MLA decode, native tiered KV cache offloading, and a month of optimization work on Kimi K3. Plus two releases, v0.29 and v0.30, which turned Model Runner V2 on by default, generalized adaptive verification to all draft-model speculators, landed watermarking support, and sped up startup and scale-out.
Then our special topic with Christopher Nuland, Chief AI Architect at Red Hat: llm-d-sc, a lightweight classification microservice built for speed. Where KV-cache routing decides which pod already has your prefix warm, semantic classification decides which model tier a request belongs to in the first place. Christopher covers why it was split out as its own service, how it integrates with llm-d and the Praxis proxy, how it compares to vLLM Semantic Router and Switchyard, what it deliberately does not try to solve, and a look at domain-based routing from an in-flight NASA proof of concept.
Slides: https://docs.google.com/presentation/d/1ZnH1l1tiTTiOBmZ62gqBGn9-SBSBKfN8nhIuB22aYdw
Want to join the discussion live on Google Meet? Get a calendar invite by filling out this form: https://red.ht/office-hours
Timestamps:
00:00 Intro
02:09 vLLM project update
12:29 vLLM v0.29 and v0.30 releases
16:35 llm-d Semantic Classifier deep dive
42:50 Q&A











