LLM VRAM calculator
Estimate GPU memory for local inference from model size, precision, context length, and KV cache. Use presets, then tune for your stack.
Dense 8B style. Good default for single consumer GPUs.
Estimate
- Weights: ~4.0 GB
- KV cache (fp16-ish): ~4.3 GB
- Total with overhead: ~9.1 GB
Rough GPU fit: 8 GB card if under ~7 GB, 12 GB if under ~11 GB, 24 GB if under ~22 GB, 48 GB+ for larger totals. Leave headroom for the OS and framework.
Related tools
GPU / LLM cost · Fine-tune cost · AI video cost · Local and GPU hub · API cost calculator · Cheapest API models
Related guides
GPU LLM cost calculator: compare API spend to self-host GPUs
Sources and references
Official documentation used for definitions, counting methods, or rate cards. Always confirm critical budgets on the provider page.
FAQ
- Is this Exact VRAM from a profiler?
- No. It is a planning estimate from parameters, precision, context, and a simple KV model. Always verify with your serving stack (vLLM, llama.cpp, TensorRT-LLM) on real hardware.
- Why do MoE models look huge?
- Weight VRAM follows total parameters stored. Active FLOPs per token can be much lower. Edit the preset if your checkpoint or expert offload differs.
- What about CUDA graphs and fragmentation?
- The overhead slider covers activations and allocator slack. Raise it if you hit OOM with long context or concurrent sessions.