🪟 Installing Ollama on Windows

Complete guide for Windows: native install, WSL2, GPU acceleration, troubleshooting.

🇷🇺 Russian version: setup-windows.ru.md


← Hardware guide · Install on Linux →


Contents

  1. Two installation methods
  2. Method 1: Native install (recommended)
  3. Method 2: Install via WSL2
  4. Verification
  5. Run your first model
  6. Coding setup
  7. Common issues
  8. Whats next

1. Two installation methods

Method Difficulty GPU For
Native Easy (one command) NVIDIA CUDA Most users
WSL2 Medium NVIDIA CUDA Developers, Continue.dev

If you have an NVIDIA GPU, both methods give you GPU acceleration (models run fast).
If you have AMD or integrated graphics, Ollama runs on CPU (slower, but works).


Step 1. Install

  1. Download: ollama.com/download/windows
  2. Run OllamaSetup.exe
  3. Wait (Ollama icon appears in system tray)

Step 2. Verify

Open PowerShell or Command Prompt:

ollama --version

Should show: ollama version 0.x.x

Step 3. Confirm the server is running

Ollama starts automatically at login. Icon in system tray (bottom right) — an alpaca dog.

Check:

curl http://localhost:11434/api/version

Should return: {"version":"0.x.x"}


3. Method 2: Install via WSL2

This method offers more capabilities (Docker, Linux tools, Continue.dev).

Step 1. Install WSL2

Open PowerShell as Administrator:

wsl --install

Reboot your computer. After reboot, a Linux window will open — create a user.

Step 2. Install Ollama inside WSL2

curl -fsSL https://ollama.com/install.sh | sh

Step 3. Use from Windows

Ollama inside WSL2 is available at localhost:11434 from Windows. VS Code, Continue.dev and other programs connect to it as local.


4. Verification

Make sure everything works:

# Check version
ollama --version

# Check installed models
ollama list

# Search available models
ollama search qwen

5. Run your first model

ollama run qwen3.5:4b

After downloading (~3.4 GB) you’ll see the >>> prompt — the model is ready to chat.

What to try:

Write a Python function that checks if a number is prime
Explain the difference between list and tuple in Python
Translate to Russian: "I need to install an AI model locally"

Commands during chat:


6. Coding setup

VS Code + Continue

After installing Ollama and running a model:

  1. Install VS Code
  2. Install Continue extension (from the marketplace)
  3. In Continue settings, add the model:
    {
      "models": [{
        "title": "Local Qwen",
        "provider": "ollama",
        "model": "qwen2.5-coder:7b",
        "apiBase": "http://localhost:11434"
      }]
    }
    

More: ../use-cases/coding.md

Aider

pip install aider-chat
set OLLAMA_API_BASE=http://127.0.0.1:11434
aider --model ollama_chat/qwen2.5-coder:7b

7. Common issues

Ollama does not see GPU (NVIDIA)

WSL2 is slow

Model does not fit in RAM

“connection refused” error


8. Whats next

If you want Go to
Choose models ../local-models/running-models.md
Set up AI coding ../use-cases/coding.md
Open WebUI docker run -d -p 3000:8080 ghcr.io/open-webui/open-webui:main
Back to learning path learning-path.md
Back to navigation README.md

In section: what-is-ai · how-models-work · cloud-vs-local · hardware-guide · glossary · faq · learning-path · setup-windows · setup-linux
Related sections: Local Models · AI Agents · Use Cases
Navigation: ← Zero Level · ↑ Back to main · 🇷🇺 Русский