What is AI, ML, and LLM?
Explained in plain language, without jargon or formulas.
[π·πΊ Russian version: what-is-ai.ru.md]
β Level Zero Β· How Neural Networks Work β
Contents
- AI, ML, LLM β Whatβs the Difference?
- What is a Language Model (LLM)?
- Where Does the Model Get Its Answers?
- Parameters: What Do 7B, 14B, 70B Mean?
- Open-Source vs Proprietary Models
- Whatβs Next
1. AI, ML, LLM β Whatβs the Difference?
These three acronyms are often confused. Letβs break them down.
ββββββββββββββββββββββββββββββββββββββββ
β AI (Artificial Intelligence) β
β Artificial Intelligence β general β
β concept: a machine that "thinks" β
β β
β ββββββββββββββββββββββββββββββββ β
β β ML (Machine Learning) β β
β β Machine Learning β AI that β β
β β learns from data β β
β β β β
β β ββββββββββββββββββββββββ β β
β β β LLM (Large Language β β β
β β β Model) β β β
β β β Large Language β β β
β β β Model β what you β β β
β β β chat with β β β
β β ββββββββββββββββββββββββ β β
β ββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββ
AI β the broadest concept. Includes chess programs, face recognition on your phone, and voice assistants.
ML β a way to create AI: instead of manually programming rules, you βfeedβ the program examples and let it find patterns itself.
LLM β a specific type of ML that works with text. ChatGPT, Claude, Qwen, DeepSeek β these are all LLMs.
For this handbook: we talk almost exclusively about LLMs β language models you can run on your own computer.
2. What is a Language Model (LLM)?
Imagine a βsmart autocompleteβ that works not with words, but with entire texts.
You start a phrase β the model continues. You ask a question β the model βcompletesβ the answer. Everything it does is predict the next word (technically, token) based on previous ones.
Analogy: An LLM is a person who has read almost the entire internet and can now finish your thought. Not because they βunderstand,β but because theyβve seen similar texts millions of times.
Important: the model doesnβt βthinkβ or βunderstandβ in the human sense. Itβs an extremely complex probability calculator: sequence of words in β most probable continuation out. But due to complexity (billions of parameters), the result looks like the model actually understands.
3. Where Does the Model Get Its Answers?
The model does not have internet access when answering. All its βknowledgeβ is what it memorized during training.
The process looks like this:
Stage 1: Training
The model is βfedβ massive amounts of text β books, articles, websites, code. Trillions of words. It learns to predict the next word again and again until it gets good enough at it.
Analogy: imagine youβve never seen chess. Youβre shown a million games, and you start guessing what move is usually made in a given position. You donβt know the rules β youβve just βplayed outβ the statistics.
Training a large model takes months and costs millions of dollars (electricity, server rental). This is done by big companies: Meta, Google, Alibaba, DeepSeek.
Stage 2: Inference
This is what you do when typing a question in chat. The model is already trained β it just applies its knowledge. This is fast and free (only electricity).
Training β model became smart. Inference β it answers questions.
Stage 3: Fine-tuning
If you need the model to understand your narrow domain β you can take a ready model and βfine-tuneβ it on your data.
Analogy: the model graduated regular school. You send it to professional development courses in your specialty. This is much faster and cheaper than teaching from scratch.
4. Parameters: What Do 7B, 14B, 70B Mean?
When looking at a model, you see: Qwen 3.5 7B, Llama 3.1 70B. The number with B β number of parameters (B = billion).
Parameters are the modelβs βneuronsβ. More parameters = potentially smarter model, but also more resources needed.
| Parameters | Analogy | Where to Run | Examples |
|---|---|---|---|
| 1β3B | π mouse brain | Any laptop, 8 GB RAM | Phi-3-mini, Qwen 2.5 1.5B |
| 7β9B | π sweet spot | MacBook / PC, 16 GB RAM | Qwen 3.5 7B, Llama 3.1 8B |
| 14β30B | π needs hardware | 32 GB RAM or GPU | Qwen 3.5 14B, DeepSeek-R1 14B |
| 70B+ | Server, 64 GB+ RAM | Llama 3 70B, DeepSeek-R1 671B |
Golden rule: more parameters = smarter model, but proportionally more memory needed. On MacBook Air 16 GB β your ceiling is 7β9B. On MacBook Pro 48 GB β you can run 30B.
Important nuance: 70B model is not 10x smarter than 7B. Itβs 20β30% smarter but needs 10x more resources. For most everyday tasks, 7β9B models are more than enough.
5. Open-Source vs Proprietary Models
| Β | Proprietary (Closed) | Open-Source (Open) |
|---|---|---|
| Examples | GPT-4, Claude, Gemini | Llama 3, Qwen, Mistral, DeepSeek |
| Who Creates | OpenAI, Anthropic, Google | Meta, Alibaba, Mistral, community |
| Where It Runs | Only on company servers | Can download and run locally |
| Cost | $10β200/mo or per token | Free |
| Can Customize | No | Yes (fine-tune, modify) |
| Privacy | Your data goes to server | Everything stays with you |
For local running we only care about open-source models. Their weights are open β you can download the model file and run it anywhere.
Leading open-source models as of July 2026:
- Qwen 3.5 (Alibaba) β best for Russian language
- Llama 3.1 / 4 (Meta) β benchmark, good English
- DeepSeek-R1 / V3 β excellent reasoning
- Mistral β efficient models
- Gemma 4 (Google) β gaining popularity
- Phi-4 (Microsoft) β small and smart
6. Whatβs Next
| If You Want To | Go To |
|---|---|
| Understand how neural networks work (by analogy) | how-models-work.md |
| Choose: cloud or local | cloud-vs-local.md |
| Check what hardware you need | hardware-guide.md |
| Skip theory and install now | ../local-models/getting-started.md |
| Back to navigation | README.md |
In section: what-is-ai Β· how-models-work Β· cloud-vs-local Β· hardware-guide Β· glossary Β· faq Β· learning-path Β· setup-windows Β· setup-linux
Related sections: Local Models Β· AI Agents Β· Use Cases
Navigation: β Zero Level Β· β Back to main Β· π·πΊ Π ΡΡΡΠΊΠΈΠΉ