Common Problems

Diagnostics and typical issues when running LLMs locally.

🇷🇺 Russian version: troubleshooting.ru.md


← Local models


Problem Cause Solution
Out of memory Model too large for RAM Use smaller model or Q4 quantization
Slow generation CPU mode, no GPU Check GPU usage: ollama ps
Model not found Typo in name ollama list to see available
API connection refused Ollama not running ollama serve or check tray icon
GPU not used Missing drivers Install CUDA Toolkit / NVIDIA drivers
Context too short Default 2048 OLLAMA_CONTEXT_LENGTH=32768
Model downloads slowly Large file (4-40 GB) Check internet, use a download manager
Chinese text in response Wrong model for language Use Qwen 3.5 for Russian

Diagnostics

# Check Ollama status
curl http://localhost:11434/api/version

# Check running models
ollama ps

# View server logs
journalctl -u ollama -f  # Linux
tail -f ~/.ollama/logs/server.log  # Mac

Reinstall

# Mac
brew reinstall ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# Windows
# Re-run OllamaSetup.exe

← Back


In section: getting-started · running-models · models · catalog · quantization · memory-and-context · tools · advanced-setup · troubleshooting · apple-silicon
Related sections: Zero Level · AI Agents · Use Cases
Navigation: ← Local Models · ↑ Back to main · 🇷🇺 Русский