LLM VRAM Calculator
Enter the parameter count and quantization method to estimate the VRAM needed for inference.
Loading...
About this tool
- The result is a rough estimate based on the model's weight size, and may differ from actual VRAM usage.
- Bits / weight can be picked from quantization presets or adjusted manually to match a known GGUF file size.
- For inference with a long context, KV cache usage grows, so set the runtime overhead higher.
- All calculation happens in the browser; input is never sent to a server.