🐧 Installing Ollama on Linux

Complete guide for Linux: CPU, CUDA, Docker, systemd, troubleshooting.

🇷🇺 Russian version: setup-linux.ru.md


← Install on Windows · Learning path →


Contents

  1. Simple install (CPU)
  2. Install with GPU (NVIDIA CUDA)
  3. Install via Docker
  4. Server configuration (systemd)
  5. Run your first model
  6. Coding setup
  7. Common issues
  8. Whats next

1. Simple install (CPU)

If you don’t have a discrete NVIDIA GPU, Ollama will run on the CPU. Slower than a GPU setup, but perfectly usable for small models (up to 7B).

Install (one command)

curl -fsSL https://ollama.com/install.sh | sh

After installation, Ollama runs as a system service:

# Check status
systemctl status ollama

# View logs
journalctl -u ollama -f

Verification

ollama --version
curl http://localhost:11434/api/version

2. Install with GPU (NVIDIA CUDA)

Step 1. Make sure NVIDIA drivers are installed

nvidia-smi

This should show your graphics card, driver version, and VRAM size.

Step 2. Install Ollama

curl -fsSL https://ollama.com/install.sh | sh

The installer will automatically detect CUDA and enable GPU acceleration.

Step 3. Verify GPU is used

ollama ps

Should show 100% GPU when a model is running. If it shows 0% GPU, something is wrong.

If GPU is not detected

# NVIDIA Container Toolkit (required for Docker)
sudo apt install nvidia-container-toolkit   # Debian/Ubuntu
sudo dnf install nvidia-container-toolkit   # Fedora

# Or check environment variables
export OLLAMA_CUDA=1
ollama serve

3. Install via Docker

Recommended for servers and fine-tuning:

# With GPU (NVIDIA)
docker run -d --gpus all -p 11434:11434 \
  -v ollama:/root/.ollama \
  --name ollama ollama/ollama

# Without GPU (CPU only)
docker run -d -p 11434:11434 \
  -v ollama:/root/.ollama \
  --name ollama ollama/ollama

Environment variables for Docker:

docker run -d --gpus all -p 11434:11434 \
  -e OLLAMA_CONTEXT_LENGTH=32768 \
  -e OLLAMA_FLASH_ATTENTION=1 \
  -e OLLAMA_KV_CACHE_TYPE=q8_0 \
  -v ollama:/root/.ollama \
  --name ollama ollama/ollama

4. Server configuration (systemd)

If Ollama is installed natively (not via Docker):

sudo systemctl edit ollama.service

Add:

[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
Environment="OLLAMA_CONTEXT_LENGTH=32768"
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"

Restart:

sudo systemctl daemon-reload
sudo systemctl restart ollama

Note: OLLAMA_HOST=0.0.0.0 opens Ollama to network access. If this is your personal machine, keep 127.0.0.1 (local access only).


5. Run your first model

ollama run qwen3.5:4b

After download (~3.4 GB) you will see the >>> prompt.

Try:

Write a bash script that finds all files larger than 100 MB
Explain the difference between soft and hard links in Linux

6. Coding setup

VS Code + Continue

Install VS Code and the Continue extension. Configuration:

{
  "models": [{
    "title": "Local Qwen",
    "provider": "ollama",
    "model": "qwen2.5-coder:7b",
    "apiBase": "http://localhost:11434"
  }]
}

Aider

pip install aider-chat
export OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/qwen2.5-coder:7b

More: ../use-cases/coding.md


7. Common issues

Permission denied

# Add your user to the docker group (if using Docker)
sudo usermod -aG docker $USER

# Or run Ollama as the current user
ollama serve

GPU not used in Docker

# Install nvidia-container-toolkit
sudo apt install nvidia-container-toolkit
sudo systemctl restart docker

# Use the --gpus all flag
docker run -d --gpus all ...

Model crashes with out of memory

“libcuda.so not found”

sudo apt install nvidia-cuda-toolkit   # Ubuntu/Debian
sudo dnf install cuda                  # Fedora

Stop the server

sudo systemctl stop ollama     # if installed as a service
docker stop ollama             # if running in Docker

8. Whats next

If you want Go to
Choose a model for your task ../local-models/models.md
Set up AI coding ../use-cases/coding.md
Compare tools ../local-models/tools.md
Learn about quantization ../local-models/quantization.md
Back to learning path learning-path.md
Back to navigation README.md

In section: what-is-ai · how-models-work · cloud-vs-local · hardware-guide · glossary · faq · learning-path · setup-windows · setup-linux
Related sections: Local Models · AI Agents · Use Cases
Navigation: ← Zero Level · ↑ Back to main · 🇷🇺 Русский