VRAM calculator: how much GPU memory does your model need?
Pick a model, a quantization and a context length. The calculator adds up the weights, the KV cache and the runtime overhead, then shows which GPUs can run it and the cheapest place to rent one.
GPUs that can run it
- 1× L424 GB VRAM · Vast.ai · 24/7 month: $241Best value$0.33/hr →
- 1× RTX A600048 GB VRAM · RunPod · 24/7 month: $241$0.33/hr →
- 1× RTX 409024 GB VRAM · Novita AI · 24/7 month: $245$0.34/hr →
- 1× RTX 4000 Ada20 GB VRAM · Akamai Cloud (Linode) · 24/7 month: $380$0.52/hr →
- 1× RTX 509032 GB VRAM · Vast.ai · 24/7 month: $438$0.60/hr →
- 1× L40S48 GB VRAM · RunPod · 24/7 month: $577$0.79/hr →
Without a GPU
A CPU-only VPS needs about 8 GB of RAM for this. It works, but expect slow answers above ~8B parameters.
Cheapest VPS plans with enough RAM
- Contabo · Cloud VPS 44 vCPU · 8 GB RAM$6.26/mo →
- IONOS · VPS L+4 vCPU · 8 GB RAM$8.00/mo →
- OVHcloud · VPS-24 vCPU · 8 GB RAM$8.50/mo →
Weights = parameters × bits per weight. KV cache from the model’s own layer, head and window sizes. We add 1 GiB for the runtime and keep 5% of each card free.
What uses the memory
Three things: the model weights (parameters × bits per weight), the KV cache that grows with every token of context and every parallel user, and a fixed overhead for the runtime.
The architecture numbers come from each model’s official configuration file, so sliding-window models like Gemma 3 and gpt-oss are counted correctly instead of being over-estimated.
Picking a quantization
4-bit (Q4_K_M) is the popular default: about a quarter of the memory of FP16 with a small quality loss. 8-bit is close to the original. Use FP16 only for fine-tuning or when memory is not a concern.
Questions people ask
How much VRAM for an 8B model?
About 6 GB at 4-bit with a short context, or around 17 GB at FP16. A 12–16 GB card handles it comfortably.
Why does long context use so much memory?
The KV cache stores keys and values for every token in the window, in every layer. Doubling the context doubles that part.
Does it work for Ollama?
Yes — Ollama uses llama.cpp with GGUF files, which is exactly what the quantization options describe.
Sources
Prices come from each provider’s own pricing page on the date shown. Providers change prices often — the checkout has the final word.