#gpu
-
How to Run Ollama in Docker: CPU, NVIDIA and AMD GPU Setup
How to run Ollama in Docker: the one-line CPU container, NVIDIA and AMD GPU passthrough, a reboot-safe Compose file, and keeping port 11434 private.
-
Best GPU for Ollama: 8GB to 48GB Cards Compared
Best GPU for Ollama by the two specs that matter: VRAM capacity decides what loads, memory bandwidth decides how fast it answers. Compared by tier.
-
Best Ollama Models for 8GB VRAM: What Fits
Best Ollama models for 8GB VRAM: which parameter counts and quantisations fit once the desktop takes its share, and what to run at each context length.
-
Ollama Tokens per Second: What Sets Your Speed
Ollama tokens per second explained: how memory bandwidth, quantisation, and context length set generation rate, plus how to measure your own numbers.
-
Ollama Not Using GPU: How to Diagnose and Fix It
Ollama not using GPU: read the processor split, confirm the device is visible and supported, and fix the capacity or context setting that caused it.
-
How Much VRAM Do You Need to Run Local LLMs with Ollama?
VRAM sizing for local LLMs: how quantisation, parameter count, and context length set your GPU memory bill, and which model fits 8, 12, 16, or 24 GB.