Gemma4 Self-Hosting Guide

Find the best Gemma configuration for your hardware. Search by GPU, CPU, or RAM to see what works, how fast, and what quality to expect.

CPU: AMD Ryzen 9 5900X 12-Core Processor
RAM: 121GB
GPU: GPU (via Ollama host)
Best: gemma4:31b (87% at N/A)
functiongemma-270m:latest Params268.10M QuantQ4_K_M Backendollama Score0% SpeedN/A
CPU: AMD Ryzen 9 5900X 12-Core Processor
RAM: 121GB
GPU: NVIDIA GeForce RTX 3090
Best: gemma4:e4b (14% at N/A)
gemma-4-12B-it-Q4_K_M Params12B QuantQ4_K_M Backendllama-cpp (b9496) Score10% Speed77 tok/s
gemma-4-12B-it-Q4_K_M Params12B QuantQ4_K_M Backendllama-cpp (b9496) Score4% Speed77 tok/s
CPU: AMD Ryzen 9 5900X 12-Core Processor
RAM: 121GB
GPU: NVIDIA GeForce RTX 3090 (~8.4 GB VRAM used)
Best: gemma-4-12B-it-Q4_K_M (18% at 71 tok/s)
CPU: AMD Ryzen 9 5900X 12-Core Processor
RAM: 121GB
GPU: CPU only
Best: gemma-4-26B-A4B-it-Q4_K_M (73% at N/A)
CPU: Intel(R) Core(TM) i7-4790K CPU @ 4.00GHz (8 cores)
RAM: 31.9 GB
GPU: CPU only
Best: gemma3:1b (94% at N/A)

Backend Comparison

BackendBest ForGPU SupportNotes
OllamaMost users, GPU setupsCUDA, Metal, ROCmEasiest setup, automatic model management
llama.cppFlexible quantizationCUDA, Metal, VulkanMore quant options, manual model files
gemma.cppCPU-first setupsCPU only (for now)Google-native, Gemma 2/3 only currently

Hardware Tiers

  • High-end GPU (24+ GB VRAM): Run Gemma 4 31B Dense or 26B MoE at full precision. RTX 3090/4090, A100, etc.
  • Mid-range GPU (8-16 GB VRAM): Gemma 4 26B MoE with quantization, or Gemma 4 E4B unquantized.
  • Apple Silicon (32+ GB unified): Gemma 4 26B MoE via Ollama Metal. 48+ GB can try 31B Dense.
  • CPU only (16+ GB RAM): Gemma 4 E4B or Gemma 3 4B via Ollama. Viable for interactive use at 140+ tok/s.
  • CPU only (8-16 GB RAM): Gemma 3 4B or Gemma 2 via gemma.cpp. Smaller but functional.