#quantization
-
Best Ollama Models for 8GB VRAM: What Fits
Best Ollama models for 8GB VRAM: which parameter counts and quantisations fit once the desktop takes its share, and what to run at each context length.
-
Ollama Tokens per Second: What Sets Your Speed
Ollama tokens per second explained: how memory bandwidth, quantisation, and context length set generation rate, plus how to measure your own numbers.
-
How Much VRAM Do You Need to Run Local LLMs with Ollama?
VRAM sizing for local LLMs: how quantisation, parameter count, and context length set your GPU memory bill, and which model fits 8, 12, 16, or 24 GB.