sebae banner ad-300x250
sebae intro coupon 30 off
sebae banner 728x900
sebae banner 300x250

Self host Gemma 4: Deploy LLMs on Cloud Run GPUs

0 views
0%

Self host Gemma 4: Deploy LLMs on Cloud Run GPUs

GCP credit β†’ https://goo.gle/handson-ep7-lab1
Lab β†’ https://goo.gle/guardians

In this episode, we deploy Google’s Gemma 4 model to Cloud Run two completely different ways, each with real trade-offs you need to understand before choosing one for production.

πŸ”¨ Ollama β€” model baked into the container. Instant cold starts. Rebuild to update.
⚑ vLLM β€” model mounted from Cloud Storage via FUSE. Slower first boot, but swap models without redeploying.

Both use Cloud Run GPUs, scale to zero, and ship through automated CI/CD with Cloud Build.

We build both. You decide which fits. πŸ‘‡
πŸ“¦ CI/CD with Cloud Build
πŸ–₯️ GPU accelerated serverless inference
πŸ”„ Baked in vs. decoupled model architecture
πŸš€ Scale to zero
βš–οΈ Cold start speed vs. production agility

Chapters:
0:00 – Intro
6:08 – Getting started with Agentverse lab
7:57 – Laying the foundations of the citadel
16:07 – Forging the power core: Self hosted LLMs
28:02 – Forging the citadel’s central core: Deploy vLLM
43:59 – Summary

More resources:
Cloud Run GPU documentation β†’ https://goo.gle/4sEbTvG
Ollama documentation β†’ https://goo.gle/3Qdi64w
vLLM documentation β†’ https://goo.gle/4cvvxE9
Cloud Storage FUSE β†’ https://goo.gle/4cQAb0V

Watch more Hands on AI β†’ https://www.youtube.com/watch?v=qCBreTfjFHQ&list=PLIivdWyY5sqKnJOvP89yF8t9mWuzMTcbM
πŸ”” Subscribe to Google Cloud Tech β†’ https://goo.gle/GoogleCloudTech

#Gemma4 #CloudRun

Speakers: Ayo Adedeji, Annie Wang
Products Mentioned: Agent Development Kit, Gemini API, Cloud Run

Date: April 18, 2026