#performance
-
Ollama Tokens per Second: What Sets Your Speed
Ollama tokens per second explained: how memory bandwidth, quantisation, and context length set generation rate, plus how to measure your own numbers.
-
Ollama Without a GPU: CPU-Only Speed and RAM
Ollama without a GPU: what CPU-only inference costs in tokens per second, how much system RAM each model class needs, and when it is still worth running.
-
Ollama Not Using GPU: How to Diagnose and Fix It
Ollama not using GPU: read the processor split, confirm the device is visible and supported, and fix the capacity or context setting that caused it.