How to Find and Run a Model

Practical guide: from first command to advanced HuggingFace workflows.

🇷🇺 Russian version: running-models.ru.md


← Local models · Model selection →


Covers: Ollama, LM Studio, HuggingFace integration.

# Search for models
ollama search qwen

# Pull and run
ollama run qwen3.5:4b

# Custom model from Modelfile
ollama create mymodel -f ./Modelfile
ollama run mymodel

LM Studio

  1. Download from lmstudio.ai
  2. Search and download models from the GUI
  3. Load a model and chat
  4. Start local API server for dev

HuggingFace + Ollama

# Import any GGUF model from HuggingFace
ollama pull hf.co/bartowski/qwen3.5-4b-GGUF

# Manual import
ollama create mymodel --file Modelfile

API usage

import requests

r = requests.post("http://localhost:11434/api/generate", json={
    "model": "qwen3.5:4b",
    "prompt": "What is AI?",
    "stream": False
})
print(r.json()["response"])

What’s next

Go to Description
models.md Choosing the right model
advanced-setup.md Modelfile, API tuning
tools.md Compare all inference engines
Back README.md

← Back to navigation


In section: getting-started · running-models · models · catalog · quantization · memory-and-context · tools · advanced-setup · troubleshooting · apple-silicon
Related sections: Zero Level · AI Agents · Use Cases
Navigation: ← Local Models · ↑ Back to main · 🇷🇺 Русский