🐧 Installing Ollama on Linux
Complete guide for Linux: CPU, CUDA, Docker, systemd, troubleshooting.
🇷🇺 Russian version: setup-linux.ru.md
← Install on Windows · Learning path →
Contents
- Simple install (CPU)
- Install with GPU (NVIDIA CUDA)
- Install via Docker
- Server configuration (systemd)
- Run your first model
- Coding setup
- Common issues
- Whats next
1. Simple install (CPU)
If you don’t have a discrete NVIDIA GPU, Ollama will run on the CPU. Slower than a GPU setup, but perfectly usable for small models (up to 7B).
Install (one command)
curl -fsSL https://ollama.com/install.sh | sh
After installation, Ollama runs as a system service:
# Check status
systemctl status ollama
# View logs
journalctl -u ollama -f
Verification
ollama --version
curl http://localhost:11434/api/version
2. Install with GPU (NVIDIA CUDA)
Step 1. Make sure NVIDIA drivers are installed
nvidia-smi
This should show your graphics card, driver version, and VRAM size.
Step 2. Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
The installer will automatically detect CUDA and enable GPU acceleration.
Step 3. Verify GPU is used
ollama ps
Should show 100% GPU when a model is running. If it shows 0% GPU, something is wrong.
If GPU is not detected
# NVIDIA Container Toolkit (required for Docker)
sudo apt install nvidia-container-toolkit # Debian/Ubuntu
sudo dnf install nvidia-container-toolkit # Fedora
# Or check environment variables
export OLLAMA_CUDA=1
ollama serve
3. Install via Docker
Recommended for servers and fine-tuning:
# With GPU (NVIDIA)
docker run -d --gpus all -p 11434:11434 \
-v ollama:/root/.ollama \
--name ollama ollama/ollama
# Without GPU (CPU only)
docker run -d -p 11434:11434 \
-v ollama:/root/.ollama \
--name ollama ollama/ollama
Environment variables for Docker:
docker run -d --gpus all -p 11434:11434 \
-e OLLAMA_CONTEXT_LENGTH=32768 \
-e OLLAMA_FLASH_ATTENTION=1 \
-e OLLAMA_KV_CACHE_TYPE=q8_0 \
-v ollama:/root/.ollama \
--name ollama ollama/ollama
4. Server configuration (systemd)
If Ollama is installed natively (not via Docker):
sudo systemctl edit ollama.service
Add:
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_CONTEXT_LENGTH=32768"
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"
Restart:
sudo systemctl daemon-reload
sudo systemctl restart ollama
Note:
OLLAMA_HOST=0.0.0.0opens Ollama to network access. If this is your personal machine, keep127.0.0.1(local access only).
5. Run your first model
ollama run qwen3.5:4b
After download (~3.4 GB) you will see the >>> prompt.
Try:
Write a bash script that finds all files larger than 100 MB
Explain the difference between soft and hard links in Linux
6. Coding setup
VS Code + Continue
Install VS Code and the Continue extension. Configuration:
{
"models": [{
"title": "Local Qwen",
"provider": "ollama",
"model": "qwen2.5-coder:7b",
"apiBase": "http://localhost:11434"
}]
}
Aider
pip install aider-chat
export OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/qwen2.5-coder:7b
More: ../use-cases/coding.md
7. Common issues
Permission denied
# Add your user to the docker group (if using Docker)
sudo usermod -aG docker $USER
# Or run Ollama as the current user
ollama serve
GPU not used in Docker
# Install nvidia-container-toolkit
sudo apt install nvidia-container-toolkit
sudo systemctl restart docker
# Use the --gpus all flag
docker run -d --gpus all ...
Model crashes with out of memory
- Reduce context:
OLLAMA_CONTEXT_LENGTH=4096 - Use stronger quantization:
qwen3.5:4b:q3_k_m - Close other RAM-heavy programs
“libcuda.so not found”
sudo apt install nvidia-cuda-toolkit # Ubuntu/Debian
sudo dnf install cuda # Fedora
Stop the server
sudo systemctl stop ollama # if installed as a service
docker stop ollama # if running in Docker
8. Whats next
| If you want | Go to |
|---|---|
| Choose a model for your task | ../local-models/models.md |
| Set up AI coding | ../use-cases/coding.md |
| Compare tools | ../local-models/tools.md |
| Learn about quantization | ../local-models/quantization.md |
| Back to learning path | learning-path.md |
| Back to navigation | README.md |
In section: what-is-ai · how-models-work · cloud-vs-local · hardware-guide · glossary · faq · learning-path · setup-windows · setup-linux
Related sections: Local Models · AI Agents · Use Cases
Navigation: ← Zero Level · ↑ Back to main · 🇷🇺 Русский