Agent Memory: Short-term, Long-term, Vector

How agents store, retrieve, and manage context β€” from simple message history to semantic search with vector databases.

πŸ‡·πŸ‡Ί Russian version: memory.ru.md


← AI agents Β· Prompting β†’


Contents

  1. Why agents need memory
  2. Types of memory
  3. Short-term memory (working)
  4. Long-term memory (persistent)
  5. Memory in popular frameworks
  6. Problems and solutions
  7. Whats next

1. Why agents need memory

Without memory, every agent call starts from scratch. The model doesnt remember what it said earlier, what the user asked, or what decisions were made.

With memory:


2. Types of memory

Type Lifespan Storage Speed
Short-term One session Messages list Instant
Long-term (files) Indefinitely JSON, SQLite Fast
Long-term (vectors) Indefinitely Chroma, Qdrant Medium
Episodic Per task Structured logs Fast
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                  AGENT MEMORY                       β”‚
β”‚                                                     β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚         SHORT-TERM (working)             β”‚       β”‚
β”‚  β”‚  Β· Current dialog with LLM               β”‚       β”‚
β”‚  β”‚  Β· All messages in the cycle             β”‚       β”‚
β”‚  β”‚  Β· Limited by context window             β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β”‚                                                     β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚         LONG-TERM (persistent)           β”‚       β”‚
β”‚  β”‚  Β· Between work sessions                 β”‚       β”‚
β”‚  β”‚  Β· Vector DB / files / DB                β”‚       β”‚
β”‚  β”‚  Β· Not limited by context                β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β”‚                                                     β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”       β”‚
β”‚  β”‚         WORKING (task state)             β”‚       β”‚
β”‚  β”‚  Β· Current plan and progress             β”‚       β”‚
β”‚  β”‚  Β· What's done, what's left              β”‚       β”‚
β”‚  β”‚  Β· Intermediate results                  β”‚       β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

3. Short-term memory (working)

The simplest form of memory: the agent sees the entire message history.

messages = [
    {"role": "system", "content": "You are a PM agent..."},
    {"role": "user", "content": "What tasks need to be done?"},
    {"role": "assistant", "content": "We need to: 1. Write API 2. Build UI"},
    {"role": "user", "content": "Estimate timelines for each item"},
    # each new message is added here
]

The problem: context overflow

Each step of the agent loop adds ~500-5000 tokens. After 10 steps, the context can grow to 50000 tokens, exceeding the model limit.

Context compression

When history gets too long, compress it:

def compress_messages(messages, max_messages=10, model="qwen3.5:4b"):
    """Compress history if its too long."""

    if len(messages) <= max_messages:
        return messages

    # Take system prompt + last N-1 messages
    system = [m for m in messages if m["role"] == "system"][:1]
    recent = messages[-(max_messages-1):]

    # If still too long summarize the middle
    total_tokens = sum(len(m.get("content", "")) for m in messages)
    if total_tokens > 10000:
        middle = messages[1:-(max_messages-1)]
        summary = requests.post("http://localhost:11434/api/chat", json={
            "model": model,
            "messages": [
                {"role": "system", "content": "Summarize the dialogue history briefly"},
                {"role": "user", "content": json.dumps(middle)}
            ],
            "stream": False
        })
        summary_text = summary.json()["message"]["content"]

        return system + [
            {"role": "system", "content": f"Brief history: {summary_text}"}
        ] + recent

    return system + recent

When to compress


4. Long-term memory (persistent)

Short-term memory lives while the agent runs. Long-term memory persists between sessions.

Method Storage When to use
JSON files agent_memory.json Simple projects, single user
SQLite memory.db Multiple agents, structured data
Vector DB Chroma, Qdrant Semantic memory search
Git Commits in repository For coding agents

Example: file-based memory

import os, json

class FileMemory:
    """Simple long-term memory in a JSON file."""

    def __init__(self, filepath="agent_memory.json"):
        self.filepath = filepath
        self.data = self._load()

    def _load(self) -> dict:
        if os.path.exists(self.filepath):
            with open(self.filepath) as f:
                return json.load(f)
        return {"projects": {}, "decisions": [], "facts": []}

    def save(self):
        with open(self.filepath, "w") as f:
            json.dump(self.data, f, ensure_ascii=False, indent=2)

    def remember_fact(self, fact: str):
        """Remember a fact about the project."""
        self.data["facts"].append({
            "fact": fact,
            "timestamp": __import__("datetime").datetime.now().isoformat()
        })
        self.save()

    def get_project_state(self, project: str) -> dict:
        """Get project state."""
        return self.data["projects"].get(project, {})

    def update_project(self, project: str, key: str, value):
        """Update project state."""
        if project not in self.data["projects"]:
            self.data["projects"][project] = {}
        self.data["projects"][project][key] = value
        self.save()


# Usage in an agent
memory = FileMemory()

# PM agent remembers a decision
memory.remember_fact("We decided to use FastAPI for the backend")
memory.update_project("awesome-ai-handbook", "status", "in development")

# Another agent reads
print(memory.get_project_state("awesome-ai-handbook"))

Example: vector memory (ChromaDB)

import chromadb
from chromadb.utils import embedding_functions

class VectorMemory:
    """Long-term memory with semantic search."""

    def __init__(self, collection_name="agent_memory"):
        self.client = chromadb.Client()
        self.collection = self.client.get_or_create_collection(
            name=collection_name,
            embedding_function=embedding_functions.DefaultEmbeddingFunction()
        )

    def add(self, text: str, metadata: dict = None):
        """Add information to memory."""
        self.collection.add(
            documents=[text],
            metadatas=[metadata or {}],
            ids=[str(hash(text))]
        )

    def search(self, query: str, n_results: int = 3) -> list:
        """Find similar records."""
        results = self.collection.query(
            query_texts=[query],
            n_results=n_results
        )
        return results["documents"][0]


# Agent remembers context
memory = VectorMemory()
memory.add(
    "The user asked to build a TODO app with a web interface",
    {"project": "todo-app", "type": "requirement"}
)

# Later another agent searches
relevant = memory.search("What did the user want?")
print(relevant)

LangGraph checkpointing

from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import StateGraph, MessagesState, START
from langchain_ollama import ChatOllama

llm = ChatOllama(model="qwen3.5:4b")

def call_model(state: MessagesState):
    response = llm.invoke(state["messages"])
    return {"messages": [response]}

graph = StateGraph(MessagesState)
graph.add_node("agent", call_model)
graph.add_edge(START, "agent")

memory_saver = MemorySaver()
agent = graph.compile(checkpointer=memory_saver)

config = {"configurable": {"thread_id": "project-123"}}
agent.invoke({"messages": [("user", "Hello!")]}, config)
agent.invoke({"messages": [("user", "What did I just say?")]}, config)

CrewAI built-in memory

from crewai import Crew, Process
from crewai.memory import Memory

crew = Crew(
    agents=[...],
    tasks=[...],
    memory=Memory(
        short_term=True,
        long_term=True,
        entity=True
    )
)

Agno session state

from agno.agent import Agent
from agno.models.ollama import Ollama

agent = Agent(
    model=Ollama(id="qwen3.5:4b"),
    session_state={}
)

agent.session_state["project"] = "awesome-ai-handbook"
agent.run("Remember: we are making an AI handbook")

6. Problems and solutions

Problem Description Solution
Context forgetting Model forgets the beginning of dialogue Summarize every N steps
Context growth Too many messages exceed limit Compression, delete old messages
Conflicting memory Agents write different data Centralized storage, versioning
Data leakage Sensitive data in history Scoping, clear after completion
Embedding dependency Vector search may not find what you expect Always verify relevance of results

7. Whats next

| If you want | Go to | |β€”β€”β€”β€”-|β€”β€”-| | Learn agent prompting | prompting.md | | Orchestrate an agent team | multi-agent.md | | Secure your agents | safety.md | | Build your own team | tutorials/02-agent-team.md | | Back to navigation | README.md | β€”


In section: architecture Β· evaluation Β· frameworks Β· memory Β· multi-agent Β· ollama-for-agents Β· orchestrators Β· patterns Β· prompting Β· ready-made Β· safety Β· skills
Related sections: Zero Level Β· Local Models Β· Use Cases Β· Resources
Navigation: ← AI Agents Β· ↑ Back to main Β· πŸ‡·πŸ‡Ί Русский