Agent Architecture
Core components, agent loop, state management, and orchestration β what makes an agent work.
π·πΊ Russian version: architecture.ru.md
Contents
1. Agent components
Any AI agent, regardless of complexity, consists of four mandatory components:
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β AI AGENT β
β β
β ββββββββββββββββ ββββββββββββββββ ββββββββββββββββββ β
β β LLM β β Tools β β Orchestrator β β
β β (brain) β β (hands) β β (control) β β
β ββββββββ¬ββββββββ ββββββββ¬ββββββββ βββββββββ¬ββββββββββ β
β β β β β
β βΌ βΌ βΌ β
β Makes decisions Calls functions, Manages loop, β
β API, reads files, state, memory β
β searches the web β
β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β
β β Memory (state) β β
β β Step history Β· Context Β· Results β β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
1.1 LLM β decision making model
| Role | What it does | Example |
|---|---|---|
| Reasoner | Analyzes situation and plans | βUser asked about weather. Need to call weather APIβ |
| Decider | Chooses next step | βI have API result β I can answerβ |
| Generator | Formulates final response | βIts currently +22C and clear in Tokyoβ |
For local agents, the same models are used for chatting, but there are nuances:
Requirements:
- Tool calling β model must support function calling (Qwen 3.5+, Llama 3.1+)
- Structured output β model must return JSON by schema
- Long context β agent loop quickly fills context, need 32K+ tokens
1.2 Tools
Tools are functions the model can call. They extend the modelβs capabilities beyond its knowledge.
| Type | Example | Purpose |
|---|---|---|
| Search | Web search, vector search | Find current information |
| Read | File reading, PDF, websites | Get data for analysis |
| Write | File creation, email sending | Act in the system |
| Compute | Calculator, Python code | Precise calculations |
| API | External service calls | System integration |
| Communication | Telegram, Slack, email | Talk to humans |
Each tool is described by JSON Schema β the model βreadsβ the description and decides when to call it.
tool = {
"type": "function",
"function": {
"name": "search_web",
"description": "Search the web",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"}
},
"required": ["query"]
}
}
}
1.3 Orchestrator (control loop)
The orchestrator is the code that manages the sequence: send request to model β get decision β execute tool β return result to model β repeat.
This is my job as Sisyphus. The orchestrator decides:
- When to hand control to a tool
- When to return response to user
- When to stop (step/time/token limit)
- What to do on error (retry, fallback, inform user)
1.4 Memory (State)
The agent must remember what it has done. Details in memory.md:
| Memory type | Stores | Example |
|---|---|---|
| Short-term | Current dialogue with model | All messages in loop |
| Long-term | Information between sessions | Vector DB with projects |
| Working | Current task state | Which step is executing |
2. Agent loop
The fundamental cycle:
- User sends request β system appends to history
- LLM receives full history (system prompt + all previous steps)
- LLM decides: answer text OR call a tool
- If tool: system executes function β result added as observation β back to step 2
- If answer: text returned to user
Basic cycle
ββββββββββββ
β START β
ββββββ¬ββββββ
β
βΌ
ββββββββββββ βββββββββββββββ
β LLM ββββββΆβ Need a β
β decides β β tool? β
ββββββββββββ ββββββββ¬βββββββ
β² β ββββββββββββ
β ββββββββ€ YES ββββΆ Call β
β β β β tool β
β β β ββββββ¬ββββββ
β β β β
β β β βΌ
β β β ββββββββββββ
β β β β Result β
β β β ββββββ¬ββββββ
β β β β
β β ββββββββββββββββ
β β
β ββββββ΄βββββ
β β NO β
β βΌ β
β ββββββββββ β
ββββ€ANSWER ββββββ
β user β
ββββββββββ
Implementation in Python (no frameworks)
import requests, json
def agent_loop(model: str, user_input: str, tools: list, max_steps: int = 10):
"""Simple agent loop."""
messages = [
{"role": "system", "content": "You are a helpful assistant. Use tools when needed."},
{"role": "user", "content": user_input}
]
for step in range(max_steps):
response = requests.post("http://localhost:11434/api/chat", json={
"model": model,
"messages": messages,
"tools": tools,
"stream": False
})
data = response.json()
msg = data["message"]
messages.append(msg)
if msg.get("tool_calls"):
for tc in msg["tool_calls"]:
func_name = tc["function"]["name"]
args = tc["function"]["arguments"]
print(f" Step {step+1}: calling {func_name}({args})")
result = execute_tool(func_name, args)
messages.append({
"role": "tool",
"name": func_name,
"content": json.dumps(result)
})
else:
return msg["content"]
return "Step limit reached"
This is literally the core of any agent. LangGraph, CrewAI, Agno all do the same thing internally.
Why go framework-free?
Writing the loop manually at least once is critically important because:
- You understand what happens βunder the hoodβ of the framework
- You can debug when something goes wrong
- You are not dependent on abstractions that hide important details
- You can add your own logic (pauses, conditions, logging)
3. State and memory
Agent state is all information accumulated during its work.
agent_state = {
"messages": [],
"step": 0,
"max_steps": 10,
"tools_used": [],
"errors": [],
"intermediate_results": {},
"plan": [],
"user_intent": "",
}
How frameworks manage state
LangGraph β StateGraph with typed state:
class AgentState(TypedDict):
messages: Annotated[list, add_messages]
next_agent: str
CrewAI β each agent has internal state (role, goal, memory).
Agno β session_state for data between runs.
Common issues
| Problem | Why it happens | Solution |
|---|---|---|
| Context overflow | Each step adds tokens | Step limit, compression |
| Context loss | Model forgets dialogue start | Summarize history |
| Contradictory decisions | Too much data visible | Clean context, keep focus |
| Data leakage | State stores sensitive info | Scoping, clear after completion |
4. Context window
| Action | Tokens (approx) |
|---|---|
| System prompt | 200β500 |
| User request | 50β200 |
| Model decision (reasoning) | 200β1000 |
| Tool call (JSON) | 100β300 |
| Tool result | 200β5000 |
| One loop step | ~500β5000 tokens |
| 10 steps | ~5000β50000 tokens |
On M1 16 GB with num_ctx: 4096 you hit the limit at 3-5 steps.
Solutions
- Increase
num_ctxβ if it fits in RAM - Limit steps β
max_steps=5for most tasks - Context compression β summarize when too long
- KV cache quantization β
OLLAMA_KV_CACHE_TYPE=q4_0gives 4Γ more space
def compress_if_needed(messages, max_messages=20):
if len(messages) > max_messages:
system = messages[0]
recent = messages[-(max_messages-1):]
return [system] + recent
return messages
5. Orchestration
This is what I do as Sisyphus. One agent is just a loop. Multiple agents working together is orchestration.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β ORCHESTRATOR β
β β
β ββββββββββββ ββββββββββββ ββββββββββββββββββββ β
β β Agent 1 β β Agent 2 β β Agent 3 β β
β β PM β β Analyst β β Developer β β
β ββββββ¬ββββββ ββββββ¬ββββββ βββββββββ¬βββββββββββ β
β β β β β
β ββββββββββββββββ΄βββββββββββββββββ β
β β β
β Solve tasks, β
β exchange results β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββ
What the orchestrator decides:
- Which agent takes the task
- Order of work (sequential / parallel / hierarchical)
- How agents exchange data
- What to do if an agent fails
- When the task is complete
def orchestrate(agents: list, task: str):
result = None
for agent in agents:
context = f"Previous result: {result}" if result else ""
result = agent.run(f"{task}\n{context}")
return result
Real orchestration is more complex β see multi-agent.md and 02-agent-team tutorial.
6. Whats next
| If you want | Go to | |ββββ-|ββ-| | Study specific agent patterns (with code) | patterns.md | | How agents store information | memory.md | | How it works with frameworks | frameworks.md | | Write your first agent | tutorials/01-first-agent.md | | Back to navigation | README.md | β
In section: architecture Β· evaluation Β· frameworks Β· memory Β· multi-agent Β· ollama-for-agents Β· orchestrators Β· patterns Β· prompting Β· ready-made Β· safety Β· skills
Related sections: Zero Level Β· Local Models Β· Use Cases Β· Resources
Navigation: β AI Agents Β· β Back to main Β· π·πΊ Π ΡΡΡΠΊΠΈΠΉ