Agent Architecture

Core components, agent loop, state management, and orchestration β€” what makes an agent work.

πŸ‡·πŸ‡Ί Russian version: architecture.ru.md


← AI agents Β· Patterns β†’


Contents

  1. Agent components
  2. Agent loop
  3. State and memory
  4. Context window
  5. Orchestration
  6. Whats next

1. Agent components

Any AI agent, regardless of complexity, consists of four mandatory components:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                      AI AGENT                             β”‚
β”‚                                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  β”‚     LLM      β”‚  β”‚   Tools      β”‚  β”‚  Orchestrator  β”‚  β”‚
β”‚  β”‚   (brain)    β”‚  β”‚   (hands)    β”‚  β”‚  (control)     β”‚  β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚         β”‚                 β”‚                   β”‚           β”‚
β”‚         β–Ό                 β–Ό                   β–Ό           β”‚
β”‚  Makes decisions    Calls functions,     Manages loop,    β”‚
β”‚                     API, reads files,    state, memory    β”‚
β”‚                     searches the web                      β”‚
β”‚                                                          β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚  β”‚              Memory (state)                       β”‚    β”‚
β”‚  β”‚  Step history Β· Context Β· Results                β”‚    β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1.1 LLM β€” decision making model

Role What it does Example
Reasoner Analyzes situation and plans β€œUser asked about weather. Need to call weather API”
Decider Chooses next step β€œI have API result β€” I can answer”
Generator Formulates final response β€œIts currently +22C and clear in Tokyo”

For local agents, the same models are used for chatting, but there are nuances:

Requirements:

1.2 Tools

Tools are functions the model can call. They extend the model’s capabilities beyond its knowledge.

Type Example Purpose
Search Web search, vector search Find current information
Read File reading, PDF, websites Get data for analysis
Write File creation, email sending Act in the system
Compute Calculator, Python code Precise calculations
API External service calls System integration
Communication Telegram, Slack, email Talk to humans

Each tool is described by JSON Schema β€” the model β€œreads” the description and decides when to call it.

tool = {
    "type": "function",
    "function": {
        "name": "search_web",
        "description": "Search the web",
        "parameters": {
            "type": "object",
            "properties": {
                "query": {"type": "string", "description": "Search query"}
            },
            "required": ["query"]
        }
    }
}

1.3 Orchestrator (control loop)

The orchestrator is the code that manages the sequence: send request to model β†’ get decision β†’ execute tool β†’ return result to model β†’ repeat.

This is my job as Sisyphus. The orchestrator decides:

1.4 Memory (State)

The agent must remember what it has done. Details in memory.md:

Memory type Stores Example
Short-term Current dialogue with model All messages in loop
Long-term Information between sessions Vector DB with projects
Working Current task state Which step is executing

2. Agent loop

The fundamental cycle:

  1. User sends request β†’ system appends to history
  2. LLM receives full history (system prompt + all previous steps)
  3. LLM decides: answer text OR call a tool
  4. If tool: system executes function β†’ result added as observation β†’ back to step 2
  5. If answer: text returned to user

Basic cycle

  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  START   β”‚
  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ LLM      │────▢│ Need a      β”‚
  β”‚ decides  β”‚     β”‚ tool?       β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
       β–²                  β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚           β”Œβ”€β”€β”€β”€β”€β”€β”€  YES ───▢  Call    β”‚
       β”‚           β”‚      β”‚         β”‚ tool     β”‚
       β”‚           β”‚      β”‚         β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
       β”‚           β”‚      β”‚              β”‚
       β”‚           β”‚      β”‚              β–Ό
       β”‚           β”‚      β”‚         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚           β”‚      β”‚         β”‚ Result   β”‚
       β”‚           β”‚      β”‚         β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
       β”‚           β”‚      β”‚              β”‚
       β”‚           β”‚      β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚           β”‚
       β”‚      β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β”
       β”‚      β”‚   NO    β”‚
       β”‚      β–Ό         β”‚
       β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
       └───ANSWER  β”‚β—„β”€β”€β”€β”˜
          β”‚ user   β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Implementation in Python (no frameworks)

import requests, json

def agent_loop(model: str, user_input: str, tools: list, max_steps: int = 10):
    """Simple agent loop."""
    messages = [
        {"role": "system", "content": "You are a helpful assistant. Use tools when needed."},
        {"role": "user", "content": user_input}
    ]

    for step in range(max_steps):
        response = requests.post("http://localhost:11434/api/chat", json={
            "model": model,
            "messages": messages,
            "tools": tools,
            "stream": False
        })
        data = response.json()
        msg = data["message"]
        messages.append(msg)

        if msg.get("tool_calls"):
            for tc in msg["tool_calls"]:
                func_name = tc["function"]["name"]
                args = tc["function"]["arguments"]
                print(f"  Step {step+1}: calling {func_name}({args})")

                result = execute_tool(func_name, args)

                messages.append({
                    "role": "tool",
                    "name": func_name,
                    "content": json.dumps(result)
                })
        else:
            return msg["content"]

    return "Step limit reached"

This is literally the core of any agent. LangGraph, CrewAI, Agno all do the same thing internally.

Why go framework-free?

Writing the loop manually at least once is critically important because:


3. State and memory

Agent state is all information accumulated during its work.

agent_state = {
    "messages": [],
    "step": 0,
    "max_steps": 10,
    "tools_used": [],
    "errors": [],
    "intermediate_results": {},
    "plan": [],
    "user_intent": "",
}

How frameworks manage state

LangGraph β€” StateGraph with typed state:

class AgentState(TypedDict):
    messages: Annotated[list, add_messages]
    next_agent: str

CrewAI β€” each agent has internal state (role, goal, memory).

Agno β€” session_state for data between runs.

Common issues

Problem Why it happens Solution
Context overflow Each step adds tokens Step limit, compression
Context loss Model forgets dialogue start Summarize history
Contradictory decisions Too much data visible Clean context, keep focus
Data leakage State stores sensitive info Scoping, clear after completion

4. Context window

Action Tokens (approx)
System prompt 200–500
User request 50–200
Model decision (reasoning) 200–1000
Tool call (JSON) 100–300
Tool result 200–5000
One loop step ~500–5000 tokens
10 steps ~5000–50000 tokens

On M1 16 GB with num_ctx: 4096 you hit the limit at 3-5 steps.

Solutions

  1. Increase num_ctx β€” if it fits in RAM
  2. Limit steps β€” max_steps=5 for most tasks
  3. Context compression β€” summarize when too long
  4. KV cache quantization β€” OLLAMA_KV_CACHE_TYPE=q4_0 gives 4Γ— more space
def compress_if_needed(messages, max_messages=20):
    if len(messages) > max_messages:
        system = messages[0]
        recent = messages[-(max_messages-1):]
        return [system] + recent
    return messages

5. Orchestration

This is what I do as Sisyphus. One agent is just a loop. Multiple agents working together is orchestration.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                   ORCHESTRATOR                       β”‚
β”‚                                                      β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚ Agent 1  β”‚  β”‚ Agent 2  β”‚  β”‚    Agent 3       β”‚   β”‚
β”‚  β”‚ PM       β”‚  β”‚ Analyst  β”‚  β”‚  Developer       β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚       β”‚              β”‚                β”‚              β”‚
β”‚       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜              β”‚
β”‚                      β”‚                               β”‚
β”‚               Solve tasks,                            β”‚
β”‚               exchange results                        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

What the orchestrator decides:

def orchestrate(agents: list, task: str):
    result = None
    for agent in agents:
        context = f"Previous result: {result}" if result else ""
        result = agent.run(f"{task}\n{context}")
    return result

Real orchestration is more complex β€” see multi-agent.md and 02-agent-team tutorial.


6. Whats next

| If you want | Go to | |β€”β€”β€”β€”-|β€”β€”-| | Study specific agent patterns (with code) | patterns.md | | How agents store information | memory.md | | How it works with frameworks | frameworks.md | | Write your first agent | tutorials/01-first-agent.md | | Back to navigation | README.md | β€”


In section: architecture Β· evaluation Β· frameworks Β· memory Β· multi-agent Β· ollama-for-agents Β· orchestrators Β· patterns Β· prompting Β· ready-made Β· safety Β· skills
Related sections: Zero Level Β· Local Models Β· Use Cases Β· Resources
Navigation: ← AI Agents Β· ↑ Back to main Β· πŸ‡·πŸ‡Ί Русский