sebae banner ad-300x250
sebae intro coupon 30 off
sebae banner 728x900
sebae banner 300x250

Enterprise gen AI inference demo with vLLM on Red Hat AI stack

0 views
0%

Enterprise gen AI inference demo with vLLM on Red Hat AI stack

How can you scale enterprise LLM inference to production? In this demo from Red Hat Summit, Grace Ableidinger, AI Developer Advocate at Red Hat, breaks down how to scale AI workloads from a single RAG chatbot to a multi-model enterprise architecture. Learn how AI Gateway manages control plane routing, how vLLM handles paged attention and continuous batching, and how llm-d optimizes routing across replicas. Discover optimization techniques including model quantization, speculative decoding, and prefill-decode disaggregation to reduce latency and maximize GPU efficiency.

00:00 Introduction to enterprise gen AI inference
00:15 AI Gateway, vLLM, and llm-d architecture overview
00:49 Scaling a single GPU RAG chatbot with vLLM
01:16 Model quantization with vLLM and Hugging Face
01:50 Speculative decoding for faster token generation
02:20 Horizontal scaling and smart routing with llm-d
03:34 Prefill-decode disaggregation for agentic workflows
04:06 Exploring the Red Hat AI model catalog

Find more gen AI resources from Red Hat:
✨ Explore Red Hat AI → https://www.redhat.com/en/products/ai
📖 Read the Red Hat AI blog → https://www.redhat.com/en/blog/channel/artificial-intelligence
🤝 Browse models on Red Hat AI Hugging Face → https://huggingface.co/RedHat

#RedHatAI #OpenShiftAI #vLLM #LLMD #GenAI #AIInference #LLM #MachineLearning

Date: August 11, 2026