Common Problems
Diagnostics and typical issues when running LLMs locally.
🇷🇺 Russian version: troubleshooting.ru.md
| Problem | Cause | Solution |
|---|---|---|
| Out of memory | Model too large for RAM | Use smaller model or Q4 quantization |
| Slow generation | CPU mode, no GPU | Check GPU usage: ollama ps |
| Model not found | Typo in name | ollama list to see available |
| API connection refused | Ollama not running | ollama serve or check tray icon |
| GPU not used | Missing drivers | Install CUDA Toolkit / NVIDIA drivers |
| Context too short | Default 2048 | OLLAMA_CONTEXT_LENGTH=32768 |
| Model downloads slowly | Large file (4-40 GB) | Check internet, use a download manager |
| Chinese text in response | Wrong model for language | Use Qwen 3.5 for Russian |
Diagnostics
# Check Ollama status
curl http://localhost:11434/api/version
# Check running models
ollama ps
# View server logs
journalctl -u ollama -f # Linux
tail -f ~/.ollama/logs/server.log # Mac
Reinstall
# Mac
brew reinstall ollama
# Linux
curl -fsSL https://ollama.com/install.sh | sh
# Windows
# Re-run OllamaSetup.exe
In section: getting-started · running-models · models · catalog · quantization · memory-and-context · tools · advanced-setup · troubleshooting · apple-silicon
Related sections: Zero Level · AI Agents · Use Cases
Navigation: ← Local Models · ↑ Back to main · 🇷🇺 Русский