Free tool · runs in your browser · nothing sent anywhere

LLM VRAM Calculator

Will this model fit on my GPU? Pick a model and your graphics card or Mac. You'll see whether it fits, which quant (GGUF Q4 to FP16) to run, and how much context you can keep.

8,192 tok
Pick a model and GPU
Model weights: —
KV cache (context): —
Overhead: —
Total needed: — of — available
Cheaper than guessing wrong. If it doesn't fit, or you want to test before buying a card, rent the exact GPU by the hour first. A couple of dollars beats an $800 mistake.
Rent on RunPod → · Rent on Vast.ai →

How much VRAM does each model need?

Total VRAM at an 8K context, including the KV cache and a working-memory margin. Use the calculator above for other context lengths.

ModelQ4Q5Q8FP16
Llama 3.2 3B3 GB4 GB5 GB8 GB
Qwen3 4B4 GB4 GB6 GB10 GB
Gemma 3 4B4 GB4 GB6 GB10 GB
Mistral 7B5 GB6 GB9 GB17 GB
Llama 3.1 8B6 GB7 GB10 GB19 GB
Qwen3 8B6 GB7 GB10 GB19 GB
Gemma 3 12B9 GB11 GB16 GB29 GB
Qwen3 14B9 GB11 GB17 GB32 GB
Phi-4 14B10 GB12 GB18 GB34 GB
Mistral Small 24B15 GB19 GB29 GB54 GB
Gemma 3 27B19 GB23 GB34 GB63 GB
Qwen3 32B21 GB25 GB39 GB72 GB
DeepSeek-R1 Distill 32B21 GB25 GB39 GB72 GB
Llama 3.3 70B44 GB54 GB83 GB156 GB
Qwen2.5 72B45 GB55 GB85 GB161 GB

Estimates, not guarantees. Real usage varies by a few percent between llama.cpp, Ollama, vLLM and others.

What can I run on 8GB, 12GB, 16GB, 24GB or 48GB?

VRAMBiggest models that fit comfortably (Q4, 8K context)
8GBLlama 3.1 8B, Qwen3 8B, Mistral 7B
12GBPhi-4 14B, Qwen3 14B, Gemma 3 12B
16GBPhi-4 14B, Qwen3 14B, Gemma 3 12B
24GBQwen3 32B, DeepSeek-R1 Distill 32B, Gemma 3 27B
32GBQwen3 32B, DeepSeek-R1 Distill 32B, Gemma 3 27B
48GBLlama 3.3 70B, Qwen3 32B, DeepSeek-R1 Distill 32B
80GBQwen2.5 72B, Llama 3.3 70B, Qwen3 32B

On a Mac, use roughly two thirds of your unified memory (three quarters on 48GB and up) as your "VRAM", which is what the calculator does when you pick a Mac.

How this works

Three things use your VRAM: the model weights, the KV cache (which grows with context length), and a bit of overhead for the framework. We estimate weights from parameter count and quantization, size the KV cache from the model's real architecture (layers, KV heads, head size) and your chosen context, and add a working-memory margin. We call it "comfortable" only when there's about 8% headroom left, because running at 100% is how you get out-of-memory crashes mid-answer.

Everything runs in your browser. Nothing you pick is sent anywhere.

Fits? Now run it.

If your model fits, the next step is a frontend to actually use it. Our tested picks are in the self-hosted ChatGPT alternatives guide. If it doesn't fit, you've got three options: a smaller quant, a smaller model, or renting a bigger card by the hour.

Skip the setup

Once you know it fits, our SelfHost AI Stack kit gets the whole thing running in one command: chat UI, local models, and private web search, with a setup guide for people who've never touched Docker.

Get the kit — £29 →

FAQ

How much VRAM do I need for a 7B or 8B model?

About 6 GB at Q4 with an 8K context, so any 8 GB card runs it. At Q8 you need around 10 GB, and unquantized FP16 needs roughly 19 GB.

How much VRAM do I need to run a 70B model?

Around 44 GB at Q4 with an 8K context. That means two 24 GB cards, a 48 GB workstation card, a Mac with 64 GB or more, or renting a GPU by the hour.

Is 24GB of VRAM enough for local LLMs?

Yes for most people. 24 GB runs 24–32B models at Q4 with room for context, which is where local models start to feel properly capable. 70B models need about double that.

What do Q4, Q8 and GGUF mean?

They describe quantization: storing each weight in fewer bits. Q4 uses about 4 bits per weight and keeps most of the quality at roughly a quarter of the FP16 size. GGUF is the file format llama.cpp, Ollama and LM Studio use for these quantized models.

Does context length use VRAM?

Yes. The KV cache grows with every token of context. On a 32B model, going from 8K to 32K context adds several GB, which is often what tips a model from fitting to not fitting.

Can a Mac run local LLMs?

Yes. Apple Silicon shares memory between CPU and GPU, and by default macOS lets the GPU use roughly two thirds to three quarters of it. A 64 GB Mac can run 70B models at Q4.