Tools Comparison

Complete comparison of local inference tools: Ollama, LM Studio, MLX, llama.cpp, vLLM, GPT4All, and more.

πŸ‡·πŸ‡Ί Russian version: tools.ru.md


← Local models Β· Setup β†’


Contents

  1. Overview
  2. Scenarios: which tool to choose
  3. Ollama
  4. LM Studio
  5. llama.cpp
  6. MLX (Apple Silicon)
  7. vLLM
  8. GPT4All
  9. Jan
  10. Open WebUI
  11. Aider
  12. Continue.dev
  13. Enchanted
  14. LocalAI
  15. MindWork AI Studio
  16. Comparison table
  17. Benchmarks on Apple Silicon
  18. Scaling and Production
  19. Decision Tree
  20. Installation guides
  21. Terminology glossary
  22. What’s next

1. Overview

There are many tools for running local LLMs. They differ in:


2. Scenarios: which tool to choose

Scenario Tool Why
Download and test models (discovery) LM Studio Built-in HuggingFace search by category, visual browser
Need an API for your app (integration) Ollama 3 commands, OpenAI API, any language
Maximum speed on Mac (performance) MLX (mlx-lm) or LM Studio +20–40% tok/s, –50% RAM
Need to fine-tune a model (training) MLX (mlx-lm) LoRA/QLoRA, native Apple, Python API
Private RAG on documents (RAG) GPT4All (single user) or Open WebUI (team) LocalDocs / 9 vector databases
AI assistant in VS Code (coding-IDE) Continue.dev + Ollama Tab autocomplete, inline editing, @codebase
Autonomous coding agent in CLI (coding-agent) Aider + Ollama Architect/editor split, repo map
Production server (deployment) vLLM (Linux) or LM Studio (Mac) Continuous batching, PagedAttention
Complete beginner (beginner) LM Studio GUI, no terminal commands needed
Maximum privacy (privacy) Ollama + GPT4All Everything local, zero external connections
Long context 100K+ (context) llama.cpp (raw) Speculative decoding, KV cache control
Batch processing thousands of prompts (batch) vLLM (Linux) or llama.cpp (Mac) Continuous batching

3. Ollama

The most user-friendly option. One command to download and run any model.

brew install ollama
ollama run qwen3.5:4b
Parameter Value
Interface CLI + HTTP API
Open Source MIT
Stars 148K+
Apple Silicon Metal (native) + MLX (v0.19+, 32GB+ Mac)
GPU support Metal (Mac), CUDA, ROCm, Vulkan
GitHub ollama/ollama

How it works under the hood:

GPU Support Matrix:

Platform Backend Status
Apple Silicon Metal Native
Apple Silicon MLX (v0.19+, 32GB+ RAM only, limited models)
NVIDIA Linux CUDA
NVIDIA Windows CUDA
AMD Linux ROCm
AMD Windows ROCm Experimental
Intel GPU Vulkan Limited
CPU only llama.cpp Always fallback

Key features:

When to choose Ollama:

Best for: General use, agents, beginners, production

API Endpoints

Ollama provides an HTTP API on port 11434. All requests are POST, responses in JSON.

Endpoint Method Description
/api/chat POST Chat completion (streaming)
/api/generate POST Text generation (completion)
/api/embed POST Embeddings (new API, replaces /api/embeddings)
/api/embeddings POST Embeddings (deprecated)
/api/tags GET List local models
/api/ps GET List loaded models
/api/show POST Model info
/api/create POST Create from Modelfile
/api/pull POST Download model
/api/push POST Upload to registry
/api/copy POST Copy model
/api/delete DELETE Delete model
/api/version GET Ollama version
/api/blobs/:digest HEAD/POST Check/create blob
/api/experimental/web_search POST Web search (experimental)
/api/experimental/web_fetch POST Fetch pages (experimental)

Basic examples:

# Generate (no streaming)
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5:4b",
  "prompt": "What is the meaning of life?",
  "stream": false
}'

# Chat
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:4b",
  "messages": [{"role": "user", "content": "Hello!"}],
  "stream": false
}'

# JSON mode
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:4b",
  "messages": [{"role": "user", "content": "Extract the name from: My name is John"}],
  "format": "json",
  "stream": false
}'

# List installed models
curl http://localhost:11434/api/tags

# Loaded models (memory, CPU)
curl http://localhost:11434/api/ps

# Unload model from memory
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3.5:4b",
  "keep_alive": 0
}'

# Embeddings
curl http://localhost:11434/api/embed -d '{
  "model": "all-minilm",
  "input": ["text for vectorization"]
}'

Structured output (JSON Schema):

curl http://localhost:11434/api/chat -d '{
  "model": "qwen3.5:4b",
  "messages": [{"role": "user", "content": "Extract data: John, 25, New York"}],
  "format": {
    "type": "object",
    "properties": {
      "name": {"type": "string"},
      "age": {"type": "integer"},
      "city": {"type": "string"}
    },
    "required": ["name", "age", "city"]
  },
  "stream": false
}'

OpenAI-compatible endpoints

Any OpenAI library can connect to Ollama by changing base_url:

Endpoint Description
/v1/chat/completions Chat (OpenAI format)
/v1/completions Completion (OpenAI format)
/v1/embeddings Embeddings (OpenAI format)
/v1/models List models
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama"  # any value
)

response = client.chat.completions.create(
    model="qwen3.5:4b",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
# curl version
curl http://localhost:11434/v1/chat/completions -d '{
  "model": "qwen3.5:4b",
  "messages": [{"role": "user", "content": "Hello!"}]
}'

Modelfile Parameters

Modelfile β€” model configuration in Ollama (analogous to Dockerfile for LLMs):

FROM qwen3.5:4b
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM "You are a professional Python developer."

All PARAMETER instructions:

Parameter Default Description
num_ctx 2048* Context window size
num_predict -1 (∞) Max tokens to generate
temperature 0.8 Response creativity
top_k 40 Top-K sampling
top_p 0.9 Nucleus sampling
min_p 0.0 Min-P sampling
seed 0 Seed (reproducibility)
stop β€” Stop sequences
repeat_penalty 1.1 Repeat penalty
repeat_last_n 64 Penalty window (0=off)
num_batch auto Batch size
num_thread auto CPU threads
numa false NUMA (Linux, multi-socket)
use_mmap true Memory-mapped I/O

*num_ctx in Modelfile = 2048, but via API auto-selected based on VRAM. On M1 16GB β†’ 4096 (see README).

Other Modelfile instructions:

Instruction Description
FROM Base model (required)
TEMPLATE Prompt template (Go template)
SYSTEM System message
ADAPTER LoRA adapter
LICENSE License
MESSAGE Few-shot examples
DRAFT Draft model (speculative decoding)
REQUIRES Minimum Ollama version

Environment Variables

Key OLLAMA_* variables for configuration:

Variable Default Description
OLLAMA_HOST 127.0.0.1:11434 Server IP and port
OLLAMA_MODELS ~/.ollama/models Models path
OLLAMA_KEEP_ALIVE 5m Model lifetime in memory
OLLAMA_NUM_PARALLEL 1 Parallel requests
OLLAMA_MAX_LOADED_MODELS auto Max models on GPU
OLLAMA_LOAD_TIMEOUT 5m Load timeout
OLLAMA_KV_CACHE_TYPE f16 KV cache quantization (q8_0, q4_0)
OLLAMA_FLASH_ATTENTION false Flash attention (memory saving)
OLLAMA_GPU_OVERHEAD 0 VRAM reserve (bytes)
OLLAMA_MAX_QUEUE 512 Max request queue
OLLAMA_DEBUG false Debug mode
OLLAMA_NOHISTORY false Disable history
OLLAMA_NO_CLOUD false No cloud functions

4. LM Studio

GUI application. Download, load, and chat with models without terminal.

Parameter Value
Interface GUI + Headless daemon
Open Source Closed source
Apple Silicon MLX (auto-detection) + llama.cpp
OS macOS, Windows, Linux
Website lmstudio.ai

Key features:

LM Studio vs Ollama on M4 Pro 24GB (Qwen3-Coder-30B MoE):

Metric LM Studio (MLX) Ollama (llama.cpp) Difference
Throughput 102 tok/s 70 tok/s +46%
TTFT 291 ms 175 ms Ollama faster
GPU Power 12.4 W 15.4 W –20%
Efficiency 8.2 tok/s/W 4.5 tok/s/W +82%
Process Memory 21.4 GB 41.6 GB –49%

(Source: asiai.dev)

When to choose LM Studio:

Best for: Beginners who prefer GUI, quick testing, visual model comparison


5. llama.cpp

The engine behind Ollama. Lower level, more control.

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
make -j

# Run a model
./main -m model.gguf -p "Hello" -n 128
Parameter Value
Interface CLI + Server
Open Source MIT
Stars 70K+
Apple Silicon Metal (native)
OS All platforms
GitHub ggml-org/llama.cpp

Unique features:

Speculative decoding (M3 Ultra 192GB, Llama 3.1 70B Q4):

Mode tok/s Speedup
Direct 9.4 1Γ—
Speculative (70B+70B) 11.3 1.2Γ—
Speculative (70B+8B) 15.1 1.6Γ—

When to choose llama.cpp:

Best for: Power users, fine-tuning, benchmarking, embedding in custom apps


6. MLX (Apple Silicon)

Apple’s ML framework. Optimized for M series chips.

pip install mlx-lm
mlx_lm.generate --model Qwen/Qwen3.5-4B --prompt "Hello"
Parameter Value
Interface Python API + CLI
Open Source Apache 2.0
Stars 18K+
Apple Silicon Native (built by Apple for Apple Silicon)
OS macOS only
GitHub ml-explore/mlx

What is MLX: MLX is Apple’s machine learning framework for Apple Silicon. It includes:

Unlike llama.cpp:

Key features:

Fine-tuning with MLX:

from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen3-8B-4bit")

# LoRA fine-tuning
from mlx_lm import lora
lora.train_lora(
    model=model,
    tokenizer=tokenizer,
    train_set="data.jsonl",
    num_lora_layers=16,
    lora_rank=8,
)

Performance (M5 Max 128GB, Llama 3.1 70B Q4):

Engine tok/s Memory
MLX 18 ~39 GB
llama.cpp Metal 14 ~41 GB
Ollama (CPU) 12 ~41 GB

When to choose MLX:

Best for: Mac users wanting maximum speed, fine-tuning on Apple Silicon


7. vLLM

Production serving engine. For high-throughput, multi-user scenarios.

pip install vllm
vllm serve Qwen/Qwen3.5-4B
Parameter Value
Interface HTTP API (OpenAI-compatible)
Open Source Apache 2.0
Stars 45K+
Apple Silicon Not supported (NVIDIA/AMD only)
GitHub vllm-project/vllm

Key technologies:

When to use vLLM:

Best for: Production serving, multiple users, high throughput


8. GPT4All

Privacy-focused. Runs entirely on CPU, no GPU needed.

pip install gpt4all
Parameter Value
Interface GUI + Python API + CLI
Open Source MIT
Stars 72K+
Apple Silicon Metal
GitHub nomic-ai/gpt4all

LocalDocs (RAG):

GPT4All vs Open WebUI for RAG:

Criteria GPT4All Open WebUI
Setup One .dmg Docker / pip
Documents LocalDocs 9 vector databases
Multi-user
API Local OpenAI-compatible
Resources Lightweight Heavier (Python)
Formats PDF, TXT, MD, DOCX Any via document loaders

Key features:

When to choose GPT4All:

Best for: CPU-only machines, privacy-critical applications


9. Jan

Desktop AI client. Beautiful GUI for running models locally with optional cloud model support.

Parameter Value
Interface GUI (Electron)
Open Source AGPL (core), GUI closed
Stars 35K+
Apple Silicon Metal
GitHub janhq/jan

Key features:

When to choose Jan:


10. Open WebUI

Web interface for Ollama. Feature-rich web UI with RAG, multi-user, and extensive backend support.

Parameter Value
Interface Web (Python + Svelte)
Open Source MIT
Stars 70K+
Install docker run
GitHub open-webui/open-webui

Key features:

When to choose Open WebUI:


11. Aider

CLI coding agent. AI pair programming in your terminal β€” connects to local models via Ollama.

Parameter Value
Interface CLI
Open Source Apache 2.0
Stars 25K+
Local models Yes (via Ollama)
GitHub paul-gauthier/aider

Architecture:

Quality with local models:

When to choose Aider:


12. Continue.dev

IDE plugin. AI assistant for VS Code and JetBrains with local model support.

Parameter Value
Interface IDE plugin (VS Code, JetBrains)
Open Source Apache 2.0
Stars 25K+
Local models Yes (via Ollama)
GitHub continuedev/continue

How it works:

Per-role model configuration:

{
  "models": [
    {
      "title": "Quick chat",
      "provider": "ollama",
      "model": "qwen3:4b"
    },
    {
      "title": "Autocomplete",
      "provider": "ollama",
      "model": "qwen2.5-coder:1.5b"
    },
    {
      "title": "Complex refactoring",
      "provider": "ollama",
      "model": "qwen2.5-coder:14b"
    }
  ]
}

When to choose Continue.dev:


13. Enchanted

Native macOS/iOS client. Minimalistic SwiftUI frontend for Ollama.

Parameter Value
Interface macOS + iOS GUI (SwiftUI)
Open Source MIT
Apple Silicon Native, connects to Ollama
GitHub gluonfield/enchanted

Key features:

When to choose Enchanted:


14. LocalAI

OpenAI API drop-in replacement. Full OpenAI API including TTS, STT, and image generation.

Parameter Value
Interface HTTP API (OpenAI-compatible)
Open Source MIT
Stars 30K+
Install Docker, binaries
GitHub mudler/LocalAI

Backend support (60+):

When to choose LocalAI:


15. MindWork AI Studio

Universal multi-provider GUI. One interface for local and cloud models.

Parameter Value
Interface GUI (macOS, Windows, Linux)
Open Source FSL-1.1-MIT (β†’ MIT in 2 years)
Stars 529
Apple Silicon Native
GitHub MindWorkAI/AI-Studio

Provider support:

Type Providers
Local Ollama, LM Studio, llama.cpp, vLLM
Cloud OpenAI, Anthropic, Google Gemini, Mistral, Perplexity, xAI (Grok), DeepSeek, OpenRouter, Groq
HuggingFace Cerebras, Nebius, Together AI, Fireworks and more

Key features:

When to choose MindWork AI Studio:


16. Comparison table

Feature Ollama LM Studio llama.cpp MLX vLLM GPT4All Jan Open WebUI Aider Continue.dev Enchanted LocalAI MindWork AI Studio
Type Engine Engine+GUI Engine Engine Engine Engine+GUI Desktop UI Web UI CLI Agent IDE Plugin Mobile UI Engine+API Multi-GUI
Setup 1 command Download Build from source pip install pip/Docker pip install Download docker run pip install IDE install Download Docker Download
Platform Mac/Lin/Win Mac/Lin/Win Mac/Lin/Win Mac only Linux Mac/Lin/Win Mac/Lin/Win Web Mac/Lin/Win IDE macOS/iOS Mac/Lin/Win Mac/Lin/Win
GPU support CUDA+Metal CUDA+Metal CUDA+Metal (M only) (CUDA) (CPU only) Metal Via backend Via backend Via backend Via backend Via backend
Open Source MIT MIT Apache 2.0 Apache 2.0 MIT AGPL (core) MIT Apache 2.0 Apache 2.0 MIT MIT FSL-1.1
Quantizations K-quants K-quants K-quants + I-quants 2–6 bit AWQ/GPTQ K-quants Via backend N/A N/A N/A N/A All formats N/A
Tool calling
OpenAI API
Docker (Mac only) Community
Parallel req.
Multi-model Manual Manual
RAG support Manual Manual Manual (LocalDocs) (9 VDBs) (@codebase) (Qdrant)
Speed (M1 7B) 22–25 t/s 22–28 t/s 20–25 t/s 28–35 t/s N/A 10–15 t/s Via backend N/A N/A N/A N/A Via backend N/A

17. Benchmarks on Apple Silicon

17.1 MLX vs llama.cpp vs Ollama (M5 Max 128GB, Llama 3.1 70B Q4)

Backend tok/s Memory Difference
MLX 18 ~39 GB
llama.cpp Metal 14 ~41 GB –22%
Ollama (CPU) 12 ~41 GB –33%

(Source: CraftRigs)

17.2 MLX vs llama.cpp (M4 Max 36GB, Llama 3 8B Q4)

Backend Prefill (tok/s) Generation (tok/s) Memory
llama.cpp Metal 1420 71.3 5.8 GB
MLX 4-bit 1180 65.8 6.1 GB

(Source: Contra Collective)

17.3 MLX vs llama.cpp (M3 Ultra 192GB, Llama 3.1 70B Q4)

Backend Prefill (tok/s) Generation (tok/s) Memory
llama.cpp Metal 380 9.4 41 GB
MLX 4-bit 470 11.1 39 GB

(Source: Contra Collective)

17.4 Systematic comparison (arXiv, M2 Ultra 192GB, Qwen-2.5)

Framework Max tok/s TTFT Throughput
MLX ~230 Medium Maximum
MLC-LLM ~200 Low Best TTFT
llama.cpp ~150 Fast Lightweight
Ollama ~130 Slow Simplicity
PyTorch MPS ~7–9 β€” Not for production

(Source: arXiv:2511.05502)

17.5 Speculative decoding acceleration

Tool Model Without SD (tok/s) With SD (tok/s) Speedup
mlx-lm Llama 3.3 70B 11.2 23.5 2.1Γ—
llama.cpp Llama 3.3 70B + 8B draft 9.4 15.1 1.6Γ—

17.6 Context window impact on speed

Gemma 4 26B MoE on M3 Max 128GB:

Context Prefill (tok/s) Generation (tok/s)
1K 937 41.5
16K 1015 30.8
64K 754 15.5
128K 534 5.6

(Source: PubliVault)

17.7 M4 Pro 24 GB, Qwen3-Coder-30B MoE

Metric LM Studio (MLX) Ollama (llama.cpp) Difference
Throughput 102 tok/s 70 tok/s +46%
TTFT 291 ms 175 ms Ollama faster
GPU Power 12.4 W 15.4 W –20%
Memory 21.4 GB 41.6 GB –49%

17.8 MLX backend in Ollama

Since March 2026, Ollama can use MLX backend on Mac with 32 GB+ RAM:


18. Scaling and Production

18.1 Concurrent users

Tool Max concurrent Depends on Mechanism
vLLM 100+ GPU memory Continuous batching
Ollama 1–4 RAM, OLLAMA_NUM_PARALLEL Sequential (pre-v0.5), parallel (v0.5+)
LM Studio 1–2 RAM Server mode
llama.cpp 1–8 RAM, batch size Server mode
MLX 1–2 RAM No server mode
GPT4All 1 β€” Local only
LocalAI 4–8 RAM, backend Server mode

18.2 Docker support

Tool Docker Official image
Ollama ollama/ollama
LM Studio β€”
MLX (Mac only) mlx-community
llama.cpp ghcr.io/ggml-org/llama.cpp
vLLM vllm/vllm-openai
GPT4All β€”
Jan β€”
Open WebUI ghcr.io/open-webui/open-webui
LocalAI localai/localai
Aider Community

18.3 GPU memory management

Tool Unload Offload KV cache control Multi-GPU
Ollama On idle OLLAMA_GPU_LAYERS OLLAMA_KV_CACHE_TYPE
MLX (manual) N/A (always GPU)
llama.cpp -ngl --cache-type-k, --cache-type-v
vLLM --gpu-memory-utilization --kv-cache-dtype

19. Decision Tree

Want to run LLMs locally?
β”‚
β”œβ”€ Complete beginner, no terminal
β”‚  └─ LM Studio ─────────────────────────────── GUI, auto-MLX, model browser
β”‚
β”œβ”€ Need one command and API
β”‚  └─ Ollama ─────────────────────────────────── brew install + ollama run
β”‚     β”‚
β”‚     β”œβ”€ Want a web UI β†’ Open WebUI ─────────── Docker, RAG, multi-user
β”‚     β”œβ”€ Want IDE assistant β†’ Continue.dev ──── VS Code, tab autocomplete
β”‚     └─ Want CLI agent β†’ Aider ─────────────── architect/editor, repo map
β”‚
β”œβ”€ Mac, need maximum speed
β”‚  └─ MLX (via LM Studio or mlx-lm) ─────────── +20–40%, –50% RAM
β”‚
β”œβ”€ Need production server on Linux
β”‚  └─ vLLM ───────────────────────────────────── Continuous batching, multi-GPU
β”‚
β”œβ”€ Need RAG on documents
β”‚  β”œβ”€ Single user β†’ GPT4All
β”‚  └─ Team β†’ Open WebUI
β”‚
β”œβ”€ Need training / fine-tuning
β”‚  └─ MLX (mlx-lm) ──────────────────────────── LoRA/QLoRA on Apple Silicon
β”‚
β”œβ”€ Need full control and flexibility
β”‚  └─ llama.cpp (raw) ───────────────────────── GGUF, speculative decoding
β”‚
β”œβ”€ Need full OpenAI API (TTS, STT, images)
β”‚  └─ LocalAI ───────────────────────────────── 60+ backends, Docker
β”‚
β”œβ”€ Need a desktop client for local + cloud
β”‚  └─ Jan ───────────────────────────────────── Model Hub, remote support
β”‚
└─ Need one GUI for all providers
   └─ MindWork AI Studio ────────────────────── Local + cloud, plugins, RAG

20. Installation guides

Ollama

# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

# Run a model
ollama run qwen3:8b

# API
curl http://localhost:11434/api/generate -d '{
  "model": "qwen3:8b",
  "prompt": "Hello!",
  "stream": false
}'

# Modelfile
cat > Modelfile << 'EOF'
FROM qwen3:8b
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM "You are a professional Python developer."
EOF
ollama create my-coder -f Modelfile

LM Studio

# 1. Download .dmg from lmstudio.ai
# 2. Open β†’ Search β†’ find a model
# 3. Download β†’ Load β†’ Chat
# Server mode:
# Settings β†’ Server β†’ Enable β†’ Port 1234

MLX

# Install
pip install mlx-lm

# Inference
python -m mlx_lm.generate \
  --model mlx-community/Qwen3-8B-4bit \
  --prompt "Hello, how are you?" \
  --max-tokens 256

# Chat
python -m mlx_lm.chat \
  --model mlx-community/Qwen3-8B-4bit

# HTTP server
python -m mlx_lm.server \
  --model mlx-community/Qwen3-8B-4bit

# Fine-tuning
python -m mlx_lm.lora \
  --model mlx-community/Qwen3-8B-4bit \
  --data data.jsonl \
  --num-layers 16 \
  --lora-rank 8

llama.cpp

# Build
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && make -j

# Download GGUF
wget https://huggingface.co/Qwen/Qwen3-8B-GGUF/resolve/main/qwen3-8b-q4_k_m.gguf

# Run
./main -m qwen3-8b-q4_k_m.gguf \
  -p "Hello!" \
  -n 256 \
  -ngl 99  # all layers on GPU

# Server
./server -m qwen3-8b-q4_k_m.gguf \
  --port 8080 \
  -ngl 99

# Embeddings
./embedding -m qwen3-8b-q4_k_m.gguf \
  -p "Some text for embedding"

vLLM

# Install (Linux only)
pip install vllm

# Server
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen3-8B \
  --dtype auto \
  --gpu-memory-utilization 0.9

# API
curl http://localhost:8000/v1/chat/completions -d '{
  "model": "Qwen/Qwen3-8B",
  "messages": [{"role": "user", "content": "Hello!"}]
}'

Aider + Ollama

# Install
python -m pip install aider-chat

# Run with local model
export OLLAMA_API_BASE=http://localhost:11434
aider --model ollama/qwen2.5-coder:7b

# Architect mode (2 models)
aider --model ollama/qwen2.5-coder:7b \
  --editor-model ollama/qwen3:4b

# Modes
aider --chat-mode ask      # questions about code
aider --chat-mode code     # writing code
aider --chat-mode architect # architect + editor

Open WebUI

# Docker
docker run -d -p 3000:8080 \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  ghcr.io/open-webui/open-webui:main

# Connect to Ollama
# Settings β†’ Connections β†’ Ollama API URL: http://host.docker.internal:11434

GPT4All

# macOS
brew install --cask gpt4all

# Or download from https://gpt4all.io/

# Python API
pip install gpt4all

Jan

# Download from https://jan.ai/
# Or via Homebrew:
brew install --cask jan

# Open β†’ Model Hub β†’ Download β†’ Start chatting

LocalAI

# Docker
docker run -p 8080:8080 \
  -v $PWD/models:/models \
  localai/localai:latest

# LLM
curl http://localhost:8080/v1/chat/completions -d '{
  "model": "gpt-4",
  "messages": [{"role": "user", "content": "Hello!"}]
}'

# TTS
curl http://localhost:8080/v1/audio/speech -d '{
  "model": "tts-1",
  "input": "Hello, world!",
  "voice": "en_US-amy-medium"
}'

Enchanted

# Download from App Store (macOS and iOS)
# Or build from source:
git clone https://github.com/gluonfield/enchanted.git
# Open in Xcode β†’ Build β†’ Run

# Make sure Ollama is running on your Mac
# Configure the server URL in settings
# Start chatting from any device on your network

MindWork AI Studio

# Download from https://mindwork.ai/
# Available for macOS, Windows, Linux

# Or build from source:
git clone https://github.com/MindWorkAI/AI-Studio.git
# Follow build instructions in the repository

21. Terminology glossary

Term Meaning
TTFT (Time To First Token) Latency before the first token of a response
Continuous batching Dynamically adding/removing requests from a batch during processing
PagedAttention Efficient KV cache management that eliminates memory fragmentation
Speculative decoding Small β€œdraft” model generates tokens β†’ large model verifies them
KV cache Key/Value attention cache β€” the main RAM consumer for long contexts
GGUF Model format for llama.cpp with built-in quantization metadata
Modelfile Ollama configuration for creating custom models (analogous to Dockerfile)

22. What’s next

Go to Description
advanced-setup.md Modelfile, API tuning
benchmarks/apple-silicon.md Speed on Mac
quantization.md Compression guide
Back README.md

← Back to navigation


In section: getting-started Β· running-models Β· models Β· catalog Β· quantization Β· memory-and-context Β· tools Β· advanced-setup Β· troubleshooting Β· apple-silicon
Related sections: Zero Level Β· AI Agents Β· Use Cases
Navigation: ← Local Models Β· ↑ Back to main Β· πŸ‡·πŸ‡Ί Русский