Free tool · runs in your browser · nothing sent anywhere
LLM VRAM Calculator
Will this model fit on my GPU? Pick a model and your graphics card or Mac. You'll see whether it fits, which quant (GGUF Q4 to FP16) to run, and how much context you can keep.
Rent on RunPod → · Rent on Vast.ai →
How much VRAM does each model need?
Total VRAM at an 8K context, including the KV cache and a working-memory margin. Use the calculator above for other context lengths.
| Model | Q4 | Q5 | Q8 | FP16 |
|---|---|---|---|---|
| Llama 3.2 3B | 3 GB | 4 GB | 5 GB | 8 GB |
| Qwen3 4B | 4 GB | 4 GB | 6 GB | 10 GB |
| Gemma 3 4B | 4 GB | 4 GB | 6 GB | 10 GB |
| Mistral 7B | 5 GB | 6 GB | 9 GB | 17 GB |
| Llama 3.1 8B | 6 GB | 7 GB | 10 GB | 19 GB |
| Qwen3 8B | 6 GB | 7 GB | 10 GB | 19 GB |
| Gemma 3 12B | 9 GB | 11 GB | 16 GB | 29 GB |
| Qwen3 14B | 9 GB | 11 GB | 17 GB | 32 GB |
| Phi-4 14B | 10 GB | 12 GB | 18 GB | 34 GB |
| Mistral Small 24B | 15 GB | 19 GB | 29 GB | 54 GB |
| Gemma 3 27B | 19 GB | 23 GB | 34 GB | 63 GB |
| Qwen3 32B | 21 GB | 25 GB | 39 GB | 72 GB |
| DeepSeek-R1 Distill 32B | 21 GB | 25 GB | 39 GB | 72 GB |
| Llama 3.3 70B | 44 GB | 54 GB | 83 GB | 156 GB |
| Qwen2.5 72B | 45 GB | 55 GB | 85 GB | 161 GB |
Estimates, not guarantees. Real usage varies by a few percent between llama.cpp, Ollama, vLLM and others.
What can I run on 8GB, 12GB, 16GB, 24GB or 48GB?
| VRAM | Biggest models that fit comfortably (Q4, 8K context) |
|---|---|
| 8GB | Llama 3.1 8B, Qwen3 8B, Mistral 7B |
| 12GB | Phi-4 14B, Qwen3 14B, Gemma 3 12B |
| 16GB | Phi-4 14B, Qwen3 14B, Gemma 3 12B |
| 24GB | Qwen3 32B, DeepSeek-R1 Distill 32B, Gemma 3 27B |
| 32GB | Qwen3 32B, DeepSeek-R1 Distill 32B, Gemma 3 27B |
| 48GB | Llama 3.3 70B, Qwen3 32B, DeepSeek-R1 Distill 32B |
| 80GB | Qwen2.5 72B, Llama 3.3 70B, Qwen3 32B |
On a Mac, use roughly two thirds of your unified memory (three quarters on 48GB and up) as your "VRAM", which is what the calculator does when you pick a Mac.
How this works
Three things use your VRAM: the model weights, the KV cache (which grows with context length), and a bit of overhead for the framework. We estimate weights from parameter count and quantization, size the KV cache from the model's real architecture (layers, KV heads, head size) and your chosen context, and add a working-memory margin. We call it "comfortable" only when there's about 8% headroom left, because running at 100% is how you get out-of-memory crashes mid-answer.
Everything runs in your browser. Nothing you pick is sent anywhere.
Fits? Now run it.
If your model fits, the next step is a frontend to actually use it. Our tested picks are in the self-hosted ChatGPT alternatives guide. If it doesn't fit, you've got three options: a smaller quant, a smaller model, or renting a bigger card by the hour.
Skip the setup
Once you know it fits, our SelfHost AI Stack kit gets the whole thing running in one command: chat UI, local models, and private web search, with a setup guide for people who've never touched Docker.
Get the kit — £29 →FAQ
How much VRAM do I need for a 7B or 8B model?
About 6 GB at Q4 with an 8K context, so any 8 GB card runs it. At Q8 you need around 10 GB, and unquantized FP16 needs roughly 19 GB.
How much VRAM do I need to run a 70B model?
Around 44 GB at Q4 with an 8K context. That means two 24 GB cards, a 48 GB workstation card, a Mac with 64 GB or more, or renting a GPU by the hour.
Is 24GB of VRAM enough for local LLMs?
Yes for most people. 24 GB runs 24–32B models at Q4 with room for context, which is where local models start to feel properly capable. 70B models need about double that.
What do Q4, Q8 and GGUF mean?
They describe quantization: storing each weight in fewer bits. Q4 uses about 4 bits per weight and keeps most of the quality at roughly a quarter of the FP16 size. GGUF is the file format llama.cpp, Ollama and LM Studio use for these quantized models.
Does context length use VRAM?
Yes. The KV cache grows with every token of context. On a 32B model, going from 8K to 32K context adds several GB, which is often what tips a model from fitting to not fitting.
Can a Mac run local LLMs?
Yes. Apple Silicon shares memory between CPU and GPU, and by default macOS lets the GPU use roughly two thirds to three quarters of it. A 64 GB Mac can run 70B models at Q4.