Tools Comparison
Complete comparison of local inference tools: Ollama, LM Studio, MLX, llama.cpp, vLLM, GPT4All, and more.
π·πΊ Russian version: tools.ru.md
Contents
- Overview
- Scenarios: which tool to choose
- Ollama
- LM Studio
- llama.cpp
- MLX (Apple Silicon)
- vLLM
- GPT4All
- Jan
- Open WebUI
- Aider
- Continue.dev
- Enchanted
- LocalAI
- MindWork AI Studio
- Comparison table
- Benchmarks on Apple Silicon
- Scaling and Production
- Decision Tree
- Installation guides
- Terminology glossary
- Whatβs next
1. Overview
There are many tools for running local LLMs. They differ in:
- Supported formats (GGUF, MLX, AWQ)
- Hardware support (CPU, CUDA, Metal)
- Quantization options (Q4, Q5, Q8, etc.)
- API compatibility (OpenAI, custom)
- Ease of setup
2. Scenarios: which tool to choose
| Scenario | Tool | Why |
|---|---|---|
| Download and test models (discovery) | LM Studio | Built-in HuggingFace search by category, visual browser |
| Need an API for your app (integration) | Ollama | 3 commands, OpenAI API, any language |
| Maximum speed on Mac (performance) | MLX (mlx-lm) or LM Studio | +20β40% tok/s, β50% RAM |
| Need to fine-tune a model (training) | MLX (mlx-lm) | LoRA/QLoRA, native Apple, Python API |
| Private RAG on documents (RAG) | GPT4All (single user) or Open WebUI (team) | LocalDocs / 9 vector databases |
| AI assistant in VS Code (coding-IDE) | Continue.dev + Ollama | Tab autocomplete, inline editing, @codebase |
| Autonomous coding agent in CLI (coding-agent) | Aider + Ollama | Architect/editor split, repo map |
| Production server (deployment) | vLLM (Linux) or LM Studio (Mac) | Continuous batching, PagedAttention |
| Complete beginner (beginner) | LM Studio | GUI, no terminal commands needed |
| Maximum privacy (privacy) | Ollama + GPT4All | Everything local, zero external connections |
| Long context 100K+ (context) | llama.cpp (raw) | Speculative decoding, KV cache control |
| Batch processing thousands of prompts (batch) | vLLM (Linux) or llama.cpp (Mac) | Continuous batching |
3. Ollama
The most user-friendly option. One command to download and run any model.
brew install ollama
ollama run qwen3.5:4b
| Parameter | Value |
|---|---|
| Interface | CLI + HTTP API |
| Open Source | |
| Stars | 148K+ |
| Apple Silicon | Metal (native) + MLX (v0.19+, 32GB+ Mac) |
| GPU support | Metal (Mac), CUDA, ROCm, Vulkan |
| GitHub | ollama/ollama |
How it works under the hood:
- Wrapper around llama.cpp with automatic model downloading
- Models stored in
~/.ollama/models/blobs/in GGUF format - On run: checks locally β downloads if missing β loads into RAM β starts llama.cpp backend
- HTTP API on port 11434, compatible with OpenAI
GPU Support Matrix:
| Platform | Backend | Status |
|---|---|---|
| Apple Silicon | Metal | |
| Apple Silicon | MLX | |
| NVIDIA Linux | CUDA | |
| NVIDIA Windows | CUDA | |
| AMD Linux | ROCm | |
| AMD Windows | ROCm | |
| Intel GPU | Vulkan | |
| CPU only | llama.cpp |
Key features:
- OpenAI-compatible API (
/v1/chat/completions) - Tool calling (function calling) for agents
Modelfileβ custom models (temperature, system prompt, template)- Automatic GPU detection (CUDA + Metal)
- Docker support
- Parallel requests via
OLLAMA_NUM_PARALLEL - Multimodal (vision models)
OLLAMA_KV_CACHE_TYPEβ KV cache quantization controlOLLAMA_MAX_LOADED_MODELSβ multiple models in memory
When to choose Ollama:
Need a simple API for integration (Continue.dev, Aider, Open WebUI)
Want one command
ollama run <model>with no configNeed cross-platform (Mac + Linux + Windows)
Ollama 0.19+ β automatic MLX backend for compatible models
Need maximum speed (MLX is 20β40% faster)
Need fine-grained inference control (llama.cpp gives more)
Best for: General use, agents, beginners, production
API Endpoints
Ollama provides an HTTP API on port 11434. All requests are POST, responses in JSON.
| Endpoint | Method | Description |
|---|---|---|
/api/chat |
POST | Chat completion (streaming) |
/api/generate |
POST | Text generation (completion) |
/api/embed |
POST | Embeddings (new API, replaces /api/embeddings) |
/api/embeddings |
POST | Embeddings (deprecated) |
/api/tags |
GET | List local models |
/api/ps |
GET | List loaded models |
/api/show |
POST | Model info |
/api/create |
POST | Create from Modelfile |
/api/pull |
POST | Download model |
/api/push |
POST | Upload to registry |
/api/copy |
POST | Copy model |
/api/delete |
DELETE | Delete model |
/api/version |
GET | Ollama version |
/api/blobs/:digest |
HEAD/POST | Check/create blob |
/api/experimental/web_search |
POST | Web search (experimental) |
/api/experimental/web_fetch |
POST | Fetch pages (experimental) |
Basic examples:
# Generate (no streaming)
curl http://localhost:11434/api/generate -d '{
"model": "qwen3.5:4b",
"prompt": "What is the meaning of life?",
"stream": false
}'
# Chat
curl http://localhost:11434/api/chat -d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": false
}'
# JSON mode
curl http://localhost:11434/api/chat -d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Extract the name from: My name is John"}],
"format": "json",
"stream": false
}'
# List installed models
curl http://localhost:11434/api/tags
# Loaded models (memory, CPU)
curl http://localhost:11434/api/ps
# Unload model from memory
curl http://localhost:11434/api/generate -d '{
"model": "qwen3.5:4b",
"keep_alive": 0
}'
# Embeddings
curl http://localhost:11434/api/embed -d '{
"model": "all-minilm",
"input": ["text for vectorization"]
}'
Structured output (JSON Schema):
curl http://localhost:11434/api/chat -d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Extract data: John, 25, New York"}],
"format": {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer"},
"city": {"type": "string"}
},
"required": ["name", "age", "city"]
},
"stream": false
}'
OpenAI-compatible endpoints
Any OpenAI library can connect to Ollama by changing base_url:
| Endpoint | Description |
|---|---|
/v1/chat/completions |
Chat (OpenAI format) |
/v1/completions |
Completion (OpenAI format) |
/v1/embeddings |
Embeddings (OpenAI format) |
/v1/models |
List models |
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:11434/v1",
api_key="ollama" # any value
)
response = client.chat.completions.create(
model="qwen3.5:4b",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
# curl version
curl http://localhost:11434/v1/chat/completions -d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Modelfile Parameters
Modelfile β model configuration in Ollama (analogous to Dockerfile for LLMs):
FROM qwen3.5:4b
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM "You are a professional Python developer."
All PARAMETER instructions:
| Parameter | Default | Description |
|---|---|---|
num_ctx |
2048* | Context window size |
num_predict |
-1 (β) | Max tokens to generate |
temperature |
0.8 | Response creativity |
top_k |
40 | Top-K sampling |
top_p |
0.9 | Nucleus sampling |
min_p |
0.0 | Min-P sampling |
seed |
0 | Seed (reproducibility) |
stop |
β | Stop sequences |
repeat_penalty |
1.1 | Repeat penalty |
repeat_last_n |
64 | Penalty window (0=off) |
num_batch |
auto | Batch size |
num_thread |
auto | CPU threads |
numa |
false | NUMA (Linux, multi-socket) |
use_mmap |
true | Memory-mapped I/O |
*
num_ctxin Modelfile = 2048, but via API auto-selected based on VRAM. On M1 16GB β 4096 (see README).
Other Modelfile instructions:
| Instruction | Description |
|---|---|
FROM |
Base model (required) |
TEMPLATE |
Prompt template (Go template) |
SYSTEM |
System message |
ADAPTER |
LoRA adapter |
LICENSE |
License |
MESSAGE |
Few-shot examples |
DRAFT |
Draft model (speculative decoding) |
REQUIRES |
Minimum Ollama version |
Environment Variables
Key OLLAMA_* variables for configuration:
| Variable | Default | Description |
|---|---|---|
OLLAMA_HOST |
127.0.0.1:11434 |
Server IP and port |
OLLAMA_MODELS |
~/.ollama/models |
Models path |
OLLAMA_KEEP_ALIVE |
5m |
Model lifetime in memory |
OLLAMA_NUM_PARALLEL |
1 |
Parallel requests |
OLLAMA_MAX_LOADED_MODELS |
auto | Max models on GPU |
OLLAMA_LOAD_TIMEOUT |
5m |
Load timeout |
OLLAMA_KV_CACHE_TYPE |
f16 |
KV cache quantization (q8_0, q4_0) |
OLLAMA_FLASH_ATTENTION |
false | Flash attention (memory saving) |
OLLAMA_GPU_OVERHEAD |
0 |
VRAM reserve (bytes) |
OLLAMA_MAX_QUEUE |
512 |
Max request queue |
OLLAMA_DEBUG |
false | Debug mode |
OLLAMA_NOHISTORY |
false | Disable history |
OLLAMA_NO_CLOUD |
false | No cloud functions |
4. LM Studio
GUI application. Download, load, and chat with models without terminal.
| Parameter | Value |
|---|---|
| Interface | GUI + Headless daemon |
| Open Source | |
| Apple Silicon | MLX (auto-detection) + llama.cpp |
| OS | macOS, Windows, Linux |
| Website | lmstudio.ai |
Key features:
- Built-in model search on HuggingFace (by category, size)
- One-click download and load
- Automatic backend selection: MLX for Mac, CUDA for NVIDIA
- Drag-and-drop GGUF files
- Built-in chat with system prompt, template, parameters
- Local API server (OpenAI compatible)
- Headless daemon since v0.3 (runs in background)
- Model comparison mode
LM Studio vs Ollama on M4 Pro 24GB (Qwen3-Coder-30B MoE):
| Metric | LM Studio (MLX) | Ollama (llama.cpp) | Difference |
|---|---|---|---|
| Throughput | 102 tok/s | 70 tok/s | +46% |
| TTFT | 291 ms | 175 ms | Ollama faster |
| GPU Power | 12.4 W | 15.4 W | β20% |
| Efficiency | 8.2 tok/s/W | 4.5 tok/s/W | +82% |
| Process Memory | 21.4 GB | 41.6 GB | β49% |
(Source: asiai.dev)
When to choose LM Studio:
Just starting out β donβt want to use CLI
Want to quickly test different models
Need maximum speed on Mac (MLX)
Need automation / CI (no CLI)
Need full open source
Best for: Beginners who prefer GUI, quick testing, visual model comparison
5. llama.cpp
The engine behind Ollama. Lower level, more control.
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
make -j
# Run a model
./main -m model.gguf -p "Hello" -n 128
| Parameter | Value |
|---|---|
| Interface | CLI + Server |
| Open Source | |
| Stars | 70K+ |
| Apple Silicon | Metal (native) |
| OS | All platforms |
| GitHub | ggml-org/llama.cpp |
Unique features:
- GGUF format: K-quants (Q4_K_M, Q5_K_M) and I-quants (IQ2_XXS, IQ3_XXS)
- All GGUF quantizations (Q2_K through Q8_0)
- Speculative decoding (1.5β2Γ faster)
- KV cache quantization (q8_0, q4_0)
- Built-in embedding endpoint
- Flash attention
- Batch processing
- LoRA adapters on the fly
- Extensive parameter control
Speculative decoding (M3 Ultra 192GB, Llama 3.1 70B Q4):
| Mode | tok/s | Speedup |
|---|---|---|
| Direct | 9.4 | 1Γ |
| Speculative (70B+70B) | 11.3 | 1.2Γ |
| Speculative (70B+8B) | 15.1 | 1.6Γ |
When to choose llama.cpp:
Need cross-platform (GGUF everywhere)
Want full control over inference
Use non-standard quantizations
Want simplicity of one command
Best for: Power users, fine-tuning, benchmarking, embedding in custom apps
6. MLX (Apple Silicon)
Appleβs ML framework. Optimized for M series chips.
pip install mlx-lm
mlx_lm.generate --model Qwen/Qwen3.5-4B --prompt "Hello"
| Parameter | Value |
|---|---|
| Interface | Python API + CLI |
| Open Source | |
| Stars | 18K+ |
| Apple Silicon | Native (built by Apple for Apple Silicon) |
| OS | macOS only |
| GitHub | ml-explore/mlx |
What is MLX: MLX is Appleβs machine learning framework for Apple Silicon. It includes:
- mlx β core (arrays, autograd, optimizers)
- mlx-lm β LLM layer (inference + fine-tuning)
- mlx-examples β examples (LoRA, QLoRA, adapters)
Unlike llama.cpp:
- Not a wrapper β native framework for Apple Silicon
- Supports training (fine-tuning, LoRA, QLoRA), not just inference
- Up to 2Γ faster than llama.cpp on Mac (native optimization)
- Uses uniform quantization (4-bit, 8-bit), not K-quants
Key features:
- Up to 46% faster than llama.cpp on Apple Silicon
- 49% less memory usage
- Native Metal acceleration
- 2/3/4/6 bit quantization
- Good for fine-tuning (LoRA)
Fine-tuning with MLX:
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen3-8B-4bit")
# LoRA fine-tuning
from mlx_lm import lora
lora.train_lora(
model=model,
tokenizer=tokenizer,
train_set="data.jsonl",
num_lora_layers=16,
lora_rank=8,
)
Performance (M5 Max 128GB, Llama 3.1 70B Q4):
| Engine | tok/s | Memory |
|---|---|---|
| MLX | 18 | ~39 GB |
| llama.cpp Metal | 14 | ~41 GB |
| Ollama (CPU) | 12 | ~41 GB |
When to choose MLX:
Maximum speed on Mac
Need fine-tuning / LoRA / QLoRA
Python ecosystem
Mac only (not cross-platform)
Need a GUI
Best for: Mac users wanting maximum speed, fine-tuning on Apple Silicon
7. vLLM
Production serving engine. For high-throughput, multi-user scenarios.
pip install vllm
vllm serve Qwen/Qwen3.5-4B
| Parameter | Value |
|---|---|
| Interface | HTTP API (OpenAI-compatible) |
| Open Source | |
| Stars | 45K+ |
| Apple Silicon | |
| GitHub | vllm-project/vllm |
Key technologies:
- PagedAttention β efficient KV cache management (eliminates fragmentation)
- Continuous batching β dynamic add/remove requests from batch
- Prefix caching β cache common prompt prefixes
- Tensor parallelism β distribute model across multiple GPUs
- Speculative decoding β built-in
- AWQ/GPTQ quantization
- Kubernetes ready
When to use vLLM:
- Production server with high load (10+ concurrent users)
- Batch processing thousands of prompts
- Multi-GPU cluster
Not for Mac (NVIDIA/AMD Linux only)
Best for: Production serving, multiple users, high throughput
8. GPT4All
Privacy-focused. Runs entirely on CPU, no GPU needed.
pip install gpt4all
| Parameter | Value |
|---|---|
| Interface | GUI + Python API + CLI |
| Open Source | |
| Stars | 72K+ |
| Apple Silicon | Metal |
| GitHub | nomic-ai/gpt4all |
LocalDocs (RAG):
- Built-in document indexing (PDF, TXT, MD, DOCX)
- Vector database (custom Nomic implementation)
- Local file search β no data sent to the cloud
- Folder attachments (up to 10 collections)
GPT4All vs Open WebUI for RAG:
| Criteria | GPT4All | Open WebUI |
|---|---|---|
| Setup | One .dmg |
Docker / pip |
| Documents | LocalDocs | 9 vector databases |
| Multi-user | ||
| API | Local | OpenAI-compatible |
| Resources | Lightweight | Heavier (Python) |
| Formats | PDF, TXT, MD, DOCX | Any via document loaders |
Key features:
- No GPU required
- Built-in RAG
- Local plugin ecosystem
- Very easy setup
When to choose GPT4All:
Only need RAG on local documents
Minimal setup (one click)
Need an API server / integrations
Multi-user access
Best for: CPU-only machines, privacy-critical applications
9. Jan
Desktop AI client. Beautiful GUI for running models locally with optional cloud model support.
| Parameter | Value |
|---|---|
| Interface | GUI (Electron) |
| Open Source | |
| Stars | 35K+ |
| Apple Silicon | Metal |
| GitHub | janhq/jan |
Key features:
- Model Hub β search, browse, and download models
- Vision model support (via llama.cpp)
- Remote model support (OpenAI, Anthropic, local servers)
- Thread management (conversation history)
- Local server mode
When to choose Jan:
Need a polished desktop client
Want to combine local and cloud models in one app
Need lightweight (Electron uses significant RAM)
10. Open WebUI
Web interface for Ollama. Feature-rich web UI with RAG, multi-user, and extensive backend support.
| Parameter | Value |
|---|---|
| Interface | Web (Python + Svelte) |
| Open Source | |
| Stars | 70K+ |
| Install | docker run |
| GitHub | open-webui/open-webui |
Key features:
- Connect to Ollama, OpenAI, Anthropic, Google, AWS Bedrock
- 9 vector databases: Chroma, Milvus, Qdrant, Weaviate, PGVector, Elastic, Meilisearch, Pinecone, Supabase
- Agentic RAG β agent decides when and how to search documents
- Multi-user mode (teams)
- Web search, image generation, audio input
- Themes and customization
When to choose Open WebUI:
Need a web interface for a team
Need advanced RAG with agentic search
Need support for 9+ vector databases
Only need single-user (overkill)
11. Aider
CLI coding agent. AI pair programming in your terminal β connects to local models via Ollama.
| Parameter | Value |
|---|---|
| Interface | CLI |
| Open Source | |
| Stars | 25K+ |
| Local models | Yes (via Ollama) |
| GitHub | paul-gauthier/aider |
Architecture:
- Architect/Editor split β one plans, the other writes code
- Repo map β automatic repository map for context
- Map-refine β refines map as changes are made
- Lint & test integration β automatic code verification
Quality with local models:
- Qwen 2.5 Coder 7B β ~50% of GPT-4o quality on Aider tasks
- DeepSeek R1 7B β good for refactoring
- Qwen3-Coder-30B β ~70% of GPT-4o
When to choose Aider:
Need an autonomous coding agent in CLI
Need refactoring via prompts
Need a GUI / visual interface
12. Continue.dev
IDE plugin. AI assistant for VS Code and JetBrains with local model support.
| Parameter | Value |
|---|---|
| Interface | IDE plugin (VS Code, JetBrains) |
| Open Source | |
| Stars | 25K+ |
| Local models | Yes (via Ollama) |
| GitHub | continuedev/continue |
How it works:
- Connects to any OpenAI-compatible API (Ollama, LM Studio, vLLM)
- Chat β conversation with model inside IDE
- Inline editing β select code β prompt β model modifies
- Tab autocomplete β completions as you type (via separate model)
- @codebase β full repository context via embeddings
Per-role model configuration:
{
"models": [
{
"title": "Quick chat",
"provider": "ollama",
"model": "qwen3:4b"
},
{
"title": "Autocomplete",
"provider": "ollama",
"model": "qwen2.5-coder:1.5b"
},
{
"title": "Complex refactoring",
"provider": "ollama",
"model": "qwen2.5-coder:14b"
}
]
}
When to choose Continue.dev:
Work in VS Code / JetBrains
Want an AI assistant without Copilot subscription
Need tab autocomplete with local models
Donβt use an IDE
13. Enchanted
Native macOS/iOS client. Minimalistic SwiftUI frontend for Ollama.
| Parameter | Value |
|---|---|
| Interface | macOS + iOS GUI (SwiftUI) |
| Open Source | |
| Apple Silicon | Native, connects to Ollama |
| GitHub | gluonfield/enchanted |
Key features:
- Native macOS and iOS apps (SwiftUI)
- Connects to any Ollama server via URI (
http://localhost:11434) - Access from iPhone to Mac over local network
- Minimalistic Apple-style interface
- Markdown rendering, image support
When to choose Enchanted:
Need an iPhone/iPad client in addition to Mac
Like minimalistic design
Donβt need iOS β Ollama CLI or Open WebUI is simpler
14. LocalAI
OpenAI API drop-in replacement. Full OpenAI API including TTS, STT, and image generation.
| Parameter | Value |
|---|---|
| Interface | HTTP API (OpenAI-compatible) |
| Open Source | |
| Stars | 30K+ |
| Install | Docker, binaries |
| GitHub | mudler/LocalAI |
Backend support (60+):
- llama.cpp, transformers, diffusers, whisper, piper-tts, stable-audio
- Image generation, text-to-speech, speech-to-text
- Embeddings, reranking
- Model gallery (YAML configs)
When to choose LocalAI:
Need full OpenAI API (including TTS, STT, images)
Already have Docker infrastructure
Only need LLM (overkill)
15. MindWork AI Studio
Universal multi-provider GUI. One interface for local and cloud models.
| Parameter | Value |
|---|---|
| Interface | GUI (macOS, Windows, Linux) |
| Open Source | |
| Stars | 529 |
| Apple Silicon | Native |
| GitHub | MindWorkAI/AI-Studio |
Provider support:
| Type | Providers |
|---|---|
| Local | Ollama, LM Studio, llama.cpp, vLLM |
| Cloud | OpenAI, Anthropic, Google Gemini, Mistral, Perplexity, xAI (Grok), DeepSeek, OpenRouter, Groq |
| HuggingFace | Cerebras, Nebius, Together AI, Fireworks and more |
Key features:
- Unified interface for local and cloud models β switch without restart
- Assistants for common tasks (translation, document analysis, slide generation)
- Image generation (DALL-E, Stable Diffusion)
- RAG (Qdrant, external data via ERI)
- Lua plugins, i18n, enterprise configurations
- Free for commercial use
When to choose MindWork AI Studio:
Need one GUI for all providers (local + cloud)
Want to quickly switch between models without configuration
Need assistants for business tasks
Only need local inference β Ollama or LM Studio is simpler
Need CLI/API for automation β not a replacement for Ollama
16. Comparison table
| Feature | Ollama | LM Studio | llama.cpp | MLX | vLLM | GPT4All | Jan | Open WebUI | Aider | Continue.dev | Enchanted | LocalAI | MindWork AI Studio |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Type | Engine | Engine+GUI | Engine | Engine | Engine | Engine+GUI | Desktop UI | Web UI | CLI Agent | IDE Plugin | Mobile UI | Engine+API | Multi-GUI |
| Setup | 1 command | Download | Build from source | pip install | pip/Docker | pip install | Download | docker run |
pip install | IDE install | Download | Docker | Download |
| Platform | Mac/Lin/Win | Mac/Lin/Win | Mac/Lin/Win | Mac only | Linux | Mac/Lin/Win | Mac/Lin/Win | Web | Mac/Lin/Win | IDE | macOS/iOS | Mac/Lin/Win | Mac/Lin/Win |
| GPU support | Via backend | Via backend | Via backend | Via backend | Via backend | ||||||||
| Open Source | |||||||||||||
| Quantizations | K-quants | K-quants | K-quants + I-quants | 2β6 bit | AWQ/GPTQ | K-quants | Via backend | N/A | N/A | N/A | N/A | All formats | N/A |
| Tool calling | |||||||||||||
| OpenAI API | |||||||||||||
| Docker | |||||||||||||
| Parallel req. | |||||||||||||
| Multi-model | Manual | Manual | |||||||||||
| RAG support | Manual | Manual | Manual | ||||||||||
| Speed (M1 7B) | 22β25 t/s | 22β28 t/s | 20β25 t/s | 28β35 t/s | N/A | 10β15 t/s | Via backend | N/A | N/A | N/A | N/A | Via backend | N/A |
17. Benchmarks on Apple Silicon
17.1 MLX vs llama.cpp vs Ollama (M5 Max 128GB, Llama 3.1 70B Q4)
| Backend | tok/s | Memory | Difference |
|---|---|---|---|
| MLX | 18 | ~39 GB | |
| llama.cpp Metal | 14 | ~41 GB | β22% |
| Ollama (CPU) | 12 | ~41 GB | β33% |
(Source: CraftRigs)
17.2 MLX vs llama.cpp (M4 Max 36GB, Llama 3 8B Q4)
| Backend | Prefill (tok/s) | Generation (tok/s) | Memory |
|---|---|---|---|
| llama.cpp Metal | 1420 | 71.3 | 5.8 GB |
| MLX 4-bit | 1180 | 65.8 | 6.1 GB |
(Source: Contra Collective)
17.3 MLX vs llama.cpp (M3 Ultra 192GB, Llama 3.1 70B Q4)
| Backend | Prefill (tok/s) | Generation (tok/s) | Memory |
|---|---|---|---|
| llama.cpp Metal | 380 | 9.4 | 41 GB |
| MLX 4-bit | 470 | 11.1 | 39 GB |
(Source: Contra Collective)
17.4 Systematic comparison (arXiv, M2 Ultra 192GB, Qwen-2.5)
| Framework | Max tok/s | TTFT | Throughput |
|---|---|---|---|
| MLX | ~230 | Medium | Maximum |
| MLC-LLM | ~200 | Low | Best TTFT |
| llama.cpp | ~150 | Fast | Lightweight |
| Ollama | ~130 | Slow | Simplicity |
| PyTorch MPS | ~7β9 | β | Not for production |
(Source: arXiv:2511.05502)
17.5 Speculative decoding acceleration
| Tool | Model | Without SD (tok/s) | With SD (tok/s) | Speedup |
|---|---|---|---|---|
| mlx-lm | Llama 3.3 70B | 11.2 | 23.5 | 2.1Γ |
| llama.cpp | Llama 3.3 70B + 8B draft | 9.4 | 15.1 | 1.6Γ |
17.6 Context window impact on speed
Gemma 4 26B MoE on M3 Max 128GB:
| Context | Prefill (tok/s) | Generation (tok/s) |
|---|---|---|
| 1K | 937 | 41.5 |
| 16K | 1015 | 30.8 |
| 64K | 754 | 15.5 |
| 128K | 534 | 5.6 |
(Source: PubliVault)
17.7 M4 Pro 24 GB, Qwen3-Coder-30B MoE
| Metric | LM Studio (MLX) | Ollama (llama.cpp) | Difference |
|---|---|---|---|
| Throughput | 102 tok/s | 70 tok/s | +46% |
| TTFT | 291 ms | 175 ms | Ollama faster |
| GPU Power | 12.4 W | 15.4 W | β20% |
| Memory | 21.4 GB | 41.6 GB | β49% |
17.8 MLX backend in Ollama
Since March 2026, Ollama can use MLX backend on Mac with 32 GB+ RAM:
- Qwen 3.5-35B-A3B: 58 β 112 tok/s on M5 Max (+93%)
- Currently works with Qwen models, Llama/Mistral coming
18. Scaling and Production
18.1 Concurrent users
| Tool | Max concurrent | Depends on | Mechanism |
|---|---|---|---|
| vLLM | 100+ | GPU memory | Continuous batching |
| Ollama | 1β4 | RAM, OLLAMA_NUM_PARALLEL |
Sequential (pre-v0.5), parallel (v0.5+) |
| LM Studio | 1β2 | RAM | Server mode |
| llama.cpp | 1β8 | RAM, batch size | Server mode |
| MLX | 1β2 | RAM | No server mode |
| GPT4All | 1 | β | Local only |
| LocalAI | 4β8 | RAM, backend | Server mode |
18.2 Docker support
| Tool | Docker | Official image |
|---|---|---|
| Ollama | ollama/ollama |
|
| LM Studio | β | |
| MLX | mlx-community |
|
| llama.cpp | ghcr.io/ggml-org/llama.cpp |
|
| vLLM | vllm/vllm-openai |
|
| GPT4All | β | |
| Jan | β | |
| Open WebUI | ghcr.io/open-webui/open-webui |
|
| LocalAI | localai/localai |
|
| Aider | Community |
18.3 GPU memory management
| Tool | Unload | Offload | KV cache control | Multi-GPU |
|---|---|---|---|---|
| Ollama | OLLAMA_GPU_LAYERS |
OLLAMA_KV_CACHE_TYPE |
||
| MLX | N/A (always GPU) | |||
| llama.cpp | -ngl |
--cache-type-k, --cache-type-v |
||
| vLLM | --gpu-memory-utilization |
--kv-cache-dtype |
19. Decision Tree
Want to run LLMs locally?
β
ββ Complete beginner, no terminal
β ββ LM Studio βββββββββββββββββββββββββββββββ GUI, auto-MLX, model browser
β
ββ Need one command and API
β ββ Ollama βββββββββββββββββββββββββββββββββββ brew install + ollama run
β β
β ββ Want a web UI β Open WebUI βββββββββββ Docker, RAG, multi-user
β ββ Want IDE assistant β Continue.dev ββββ VS Code, tab autocomplete
β ββ Want CLI agent β Aider βββββββββββββββ architect/editor, repo map
β
ββ Mac, need maximum speed
β ββ MLX (via LM Studio or mlx-lm) βββββββββββ +20β40%, β50% RAM
β
ββ Need production server on Linux
β ββ vLLM βββββββββββββββββββββββββββββββββββββ Continuous batching, multi-GPU
β
ββ Need RAG on documents
β ββ Single user β GPT4All
β ββ Team β Open WebUI
β
ββ Need training / fine-tuning
β ββ MLX (mlx-lm) ββββββββββββββββββββββββββββ LoRA/QLoRA on Apple Silicon
β
ββ Need full control and flexibility
β ββ llama.cpp (raw) βββββββββββββββββββββββββ GGUF, speculative decoding
β
ββ Need full OpenAI API (TTS, STT, images)
β ββ LocalAI βββββββββββββββββββββββββββββββββ 60+ backends, Docker
β
ββ Need a desktop client for local + cloud
β ββ Jan βββββββββββββββββββββββββββββββββββββ Model Hub, remote support
β
ββ Need one GUI for all providers
ββ MindWork AI Studio ββββββββββββββββββββββ Local + cloud, plugins, RAG
20. Installation guides
Ollama
# macOS
brew install ollama
# Linux
curl -fsSL https://ollama.com/install.sh | sh
# Run a model
ollama run qwen3:8b
# API
curl http://localhost:11434/api/generate -d '{
"model": "qwen3:8b",
"prompt": "Hello!",
"stream": false
}'
# Modelfile
cat > Modelfile << 'EOF'
FROM qwen3:8b
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM "You are a professional Python developer."
EOF
ollama create my-coder -f Modelfile
LM Studio
# 1. Download .dmg from lmstudio.ai
# 2. Open β Search β find a model
# 3. Download β Load β Chat
# Server mode:
# Settings β Server β Enable β Port 1234
MLX
# Install
pip install mlx-lm
# Inference
python -m mlx_lm.generate \
--model mlx-community/Qwen3-8B-4bit \
--prompt "Hello, how are you?" \
--max-tokens 256
# Chat
python -m mlx_lm.chat \
--model mlx-community/Qwen3-8B-4bit
# HTTP server
python -m mlx_lm.server \
--model mlx-community/Qwen3-8B-4bit
# Fine-tuning
python -m mlx_lm.lora \
--model mlx-community/Qwen3-8B-4bit \
--data data.jsonl \
--num-layers 16 \
--lora-rank 8
llama.cpp
# Build
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp && make -j
# Download GGUF
wget https://huggingface.co/Qwen/Qwen3-8B-GGUF/resolve/main/qwen3-8b-q4_k_m.gguf
# Run
./main -m qwen3-8b-q4_k_m.gguf \
-p "Hello!" \
-n 256 \
-ngl 99 # all layers on GPU
# Server
./server -m qwen3-8b-q4_k_m.gguf \
--port 8080 \
-ngl 99
# Embeddings
./embedding -m qwen3-8b-q4_k_m.gguf \
-p "Some text for embedding"
vLLM
# Install (Linux only)
pip install vllm
# Server
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen3-8B \
--dtype auto \
--gpu-memory-utilization 0.9
# API
curl http://localhost:8000/v1/chat/completions -d '{
"model": "Qwen/Qwen3-8B",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Aider + Ollama
# Install
python -m pip install aider-chat
# Run with local model
export OLLAMA_API_BASE=http://localhost:11434
aider --model ollama/qwen2.5-coder:7b
# Architect mode (2 models)
aider --model ollama/qwen2.5-coder:7b \
--editor-model ollama/qwen3:4b
# Modes
aider --chat-mode ask # questions about code
aider --chat-mode code # writing code
aider --chat-mode architect # architect + editor
Open WebUI
# Docker
docker run -d -p 3000:8080 \
-v open-webui:/app/backend/data \
--name open-webui \
--restart always \
ghcr.io/open-webui/open-webui:main
# Connect to Ollama
# Settings β Connections β Ollama API URL: http://host.docker.internal:11434
GPT4All
# macOS
brew install --cask gpt4all
# Or download from https://gpt4all.io/
# Python API
pip install gpt4all
Jan
# Download from https://jan.ai/
# Or via Homebrew:
brew install --cask jan
# Open β Model Hub β Download β Start chatting
LocalAI
# Docker
docker run -p 8080:8080 \
-v $PWD/models:/models \
localai/localai:latest
# LLM
curl http://localhost:8080/v1/chat/completions -d '{
"model": "gpt-4",
"messages": [{"role": "user", "content": "Hello!"}]
}'
# TTS
curl http://localhost:8080/v1/audio/speech -d '{
"model": "tts-1",
"input": "Hello, world!",
"voice": "en_US-amy-medium"
}'
Enchanted
# Download from App Store (macOS and iOS)
# Or build from source:
git clone https://github.com/gluonfield/enchanted.git
# Open in Xcode β Build β Run
# Make sure Ollama is running on your Mac
# Configure the server URL in settings
# Start chatting from any device on your network
MindWork AI Studio
# Download from https://mindwork.ai/
# Available for macOS, Windows, Linux
# Or build from source:
git clone https://github.com/MindWorkAI/AI-Studio.git
# Follow build instructions in the repository
21. Terminology glossary
| Term | Meaning |
|---|---|
| TTFT (Time To First Token) | Latency before the first token of a response |
| Continuous batching | Dynamically adding/removing requests from a batch during processing |
| PagedAttention | Efficient KV cache management that eliminates memory fragmentation |
| Speculative decoding | Small βdraftβ model generates tokens β large model verifies them |
| KV cache | Key/Value attention cache β the main RAM consumer for long contexts |
| GGUF | Model format for llama.cpp with built-in quantization metadata |
| Modelfile | Ollama configuration for creating custom models (analogous to Dockerfile) |
22. Whatβs next
| Go to | Description |
|---|---|
| advanced-setup.md | Modelfile, API tuning |
| benchmarks/apple-silicon.md | Speed on Mac |
| quantization.md | Compression guide |
| Back | README.md |
In section: getting-started Β· running-models Β· models Β· catalog Β· quantization Β· memory-and-context Β· tools Β· advanced-setup Β· troubleshooting Β· apple-silicon
Related sections: Zero Level Β· AI Agents Β· Use Cases
Navigation: β Local Models Β· β Back to main Β· π·πΊ Π ΡΡΡΠΊΠΈΠΉ