Ollama Local LLM VRAM & Speed Calculator
Estimate exact GPU memory requirements, KV cache allocation, quantization overhead, and expected tokens/second throughput across NVIDIA, AMD, and Apple Silicon hardware.
Model & Quantization Spec
GPU & Memory Subsystem
VRAM Allocation Breakdown Calculated Specs
ollama run llama3:70b-instruct-q4_K_M
Where these numbers come from
The memory total is weights plus KV cache plus buffers. Weights are parameter count multiplied by bits per weight, which is what the quantisation setting changes. The speed estimate divides a card's memory bandwidth by the bytes read per generated token, because single-stream generation is bound by memory bandwidth rather than by compute. Both figures are estimates from published specifications, so treat them as a sizing guide rather than a measurement of your machine.
Each guide below works through one part of that arithmetic:
- How much VRAM you need to run local LLMs with Ollama explains every term in the memory total, including why context length is a second memory bill.
- Ollama tokens per second covers the bandwidth division behind the speed estimate, and how to measure your real rate.
- Best GPU for Ollama compares capacity against bandwidth by tier, from 8 GB up to 48 GB across two cards.
- Best Ollama models for 8GB VRAM works the same sums against the tightest common budget.
- Ollama on Apple Silicon adjusts the sizing for unified memory, where the GPU gets only part of the pool.
- Ollama without a GPU applies the same arithmetic to system RAM for CPU-only inference.
- Ollama not using GPU is the next stop if the model fits on paper but still runs on the CPU.