🪟 Installing Ollama on Windows
Complete guide for Windows: native install, WSL2, GPU acceleration, troubleshooting.
🇷🇺 Russian version: setup-windows.ru.md
← Hardware guide · Install on Linux →
Contents
- Two installation methods
- Method 1: Native install (recommended)
- Method 2: Install via WSL2
- Verification
- Run your first model
- Coding setup
- Common issues
- Whats next
1. Two installation methods
| Method | Difficulty | GPU | For |
|---|---|---|---|
| Native | Easy (one command) | NVIDIA CUDA | Most users |
| WSL2 | Medium | NVIDIA CUDA | Developers, Continue.dev |
If you have an NVIDIA GPU, both methods give you GPU acceleration (models run fast).
If you have AMD or integrated graphics, Ollama runs on CPU (slower, but works).
2. Method 1: Native install (recommended)
Step 1. Install
- Download: ollama.com/download/windows
- Run
OllamaSetup.exe - Wait (Ollama icon appears in system tray)
Step 2. Verify
Open PowerShell or Command Prompt:
ollama --version
Should show: ollama version 0.x.x
Step 3. Confirm the server is running
Ollama starts automatically at login. Icon in system tray (bottom right) — an alpaca dog.
Check:
curl http://localhost:11434/api/version
Should return: {"version":"0.x.x"}
3. Method 2: Install via WSL2
This method offers more capabilities (Docker, Linux tools, Continue.dev).
Step 1. Install WSL2
Open PowerShell as Administrator:
wsl --install
Reboot your computer. After reboot, a Linux window will open — create a user.
Step 2. Install Ollama inside WSL2
curl -fsSL https://ollama.com/install.sh | sh
Step 3. Use from Windows
Ollama inside WSL2 is available at localhost:11434 from Windows.
VS Code, Continue.dev and other programs connect to it as local.
4. Verification
Make sure everything works:
# Check version
ollama --version
# Check installed models
ollama list
# Search available models
ollama search qwen
5. Run your first model
ollama run qwen3.5:4b
After downloading (~3.4 GB) you’ll see the >>> prompt — the model is ready to chat.
What to try:
Write a Python function that checks if a number is prime
Explain the difference between list and tuple in Python
Translate to Russian: "I need to install an AI model locally"
Commands during chat:
/bye— exit/clear— clear historyCtrl+C— stop generation
6. Coding setup
VS Code + Continue
After installing Ollama and running a model:
- Install VS Code
- Install Continue extension (from the marketplace)
- In Continue settings, add the model:
{ "models": [{ "title": "Local Qwen", "provider": "ollama", "model": "qwen2.5-coder:7b", "apiBase": "http://localhost:11434" }] }
More: ../use-cases/coding.md
Aider
pip install aider-chat
set OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/qwen2.5-coder:7b
7. Common issues
Ollama does not see GPU (NVIDIA)
- Install CUDA Toolkit
- Update NVIDIA drivers
- Check:
ollama psshould show100% GPU
WSL2 is slow
- Ensure WSL2, not WSL1:
wsl -l -v - Disable Large Send Offload in network adapter settings
Model does not fit in RAM
- Use Q4 quantization
- Choose a smaller model:
phi4-miniinstead ofqwen3.5:7b
“connection refused” error
- Make sure Ollama is running (tray icon)
- Check:
curl http://localhost:11434/api/version
8. Whats next
| If you want | Go to |
|---|---|
| Choose models | ../local-models/running-models.md |
| Set up AI coding | ../use-cases/coding.md |
| Open WebUI | docker run -d -p 3000:8080 ghcr.io/open-webui/open-webui:main |
| Back to learning path | learning-path.md |
| Back to navigation | README.md |
In section: what-is-ai · how-models-work · cloud-vs-local · hardware-guide · glossary · faq · learning-path · setup-windows · setup-linux
Related sections: Local Models · AI Agents · Use Cases
Navigation: ← Zero Level · ↑ Back to main · 🇷🇺 Русский