Skip to main content

LLM VRAM Calculator

Enter the parameter count and quantization method to estimate the VRAM needed for inference.

Loading...

About this tool

  • The result is a rough estimate based on the model's weight size, and may differ from actual VRAM usage.
  • Bits / weight can be picked from quantization presets or adjusted manually to match a known GGUF file size.
  • For inference with a long context, KV cache usage grows, so set the runtime overhead higher.
  • All calculation happens in the browser; input is never sent to a server.