AI Agent Memory Explained: Diagrams, Tools & Architecture

AI Agent Memory Explained: Diagrams, Tools & Architecture

AI Agent Memory Explained: Diagrams, Tools & Architecture

All Posts

Learn how AI agent memory works: short-term vs. long-term, episodic vs. semantic, & a reference architecture, plus how Mem0, Letta, and Zep compare.

Ka Ling Wu

Co-Founder & CEO, Upsolve AI

10 min

AI agent memory cover: agent core connected to context window, vector store, knowledge graph, and governed data

AI agent memory is the set of systems that let an agent retain and reuse information beyond a single context window: 

  • who a user is

  • what happened in past sessions

  • and which approaches actually worked

Otherwise, every conversation starts from zero. 

This guide maps the two memory horizons, the three memory types, and a reference agent memory architecture, then compares how Mem0, Letta, and Zep implement them and draws the line where memory should stop and governed data should take over.

Key Takeaways

  • The context window is not memory. It's volatile working memory rebuilt on every call. Long-term memory for AI agents lives in external stores and is retrieved per task.

  • Purpose-built memory beats brute-force context. The Mem0 research paper reports 26% higher accuracy than OpenAI's built-in memory on the LOCOMO benchmark, with 91% lower p95 latency and 90%+ token savings versus stuffing the full history into context.

  • Memory type dictates architecture. Episodic (what happened), semantic (what's true), and procedural (how to act) memories call for different stores and retrieval strategies.

  • Memory is not a source of truth. Learned preferences belong in memory; business data belongs behind a governed semantic layer with permissions, queried live on every request.

Why the Context Window Isn't Memory

It's tempting to treat a million-token context window as memory. It isn't, for three reasons.

First, it's volatile. The context is reassembled on every model call and erased when the session ends. The MemGPT paper framed this precisely, treating the context window as RAM and arguing agents need an OS-style system that pages information in and out of external storage.

Second, recall degrades before the window fills. Stanford's "Lost in the Middle" study showed model performance drops significantly when relevant information sits in the middle of a long context. The LongMemEval benchmark later quantified the downstream effect: commercial chat assistants and long-context LLMs show around a 30% accuracy drop when recalling information across sustained, multi-session interactions.

Third, it's economically inefficient. Re-sending 100K tokens of history to answer a 50-token question multiplies cost and latency on every single turn. This is where a real memory layer earns its keep: store knowledge outside the model, retrieve only what the current task needs.

For more information, check out our full context engineering guide.

Short-Term vs. Long-Term Memory in AI Agents

Short-term (working) memory is everything assembled into the prompt for one call: system instructions, recent turns, tool outputs, and a scratchpad for intermediate reasoning

It's fast and fully visible to the model, but token-limited, billed every call, and gone at session end.

Long-term memory lives in external stores, typically a vector database for similarity search, a knowledge graph for entities and relationships, and a key-value store for stable profile facts. It persists indefinitely and only enters the context when the read path retrieves it.


Short-term memory

Long-term memory

Lifetime

One call / session

Persistent across sessions

Capacity

Token-limited

Effectively unbounded

Cost model

Billed on every call

Storage + retrieval only

Role

Active reasoning

Accumulated knowledge

The two types of memory are connected by a loop. This is a write path that extracts and consolidates knowledge out of working memory, and a read path that retrieves the relevant slice back in. 

Episodic, Semantic, and Procedural: The Three Types of Agent Memory

Episodic memory: what happened

Episodic memory records timestamped events (sessions, decisions, outcomes, corrections). 

Example: "On March 3, the user rejected the report draft for having too much jargon." 

It's the raw material for learning, and it's inherently temporal, which is why event logs and time-aware graphs fit it best. An agent without episodic memory repeats mistakes; worse, it repeats mistakes the user already corrected.

Semantic memory: what's true

Semantic memory holds distilled facts: user preferences, entity attributes, relationships, constraints. 

Example: "The user's company reports revenue in EUR and closes its fiscal year in March."

These are what most people mean by "the agent remembers me," and they're what vector stores and knowledge graphs retrieve well. Critically, semantic memories change. A system that can't update or invalidate a fact will confidently serve stale ones.

Procedural memory: how to act

Procedural memory captures methods: tool-use patterns, reusable workflows, learned skills. This is arguably the highest-leverage one. The research on agent workflow memory (Wang et al.) showed that inducing reusable workflows from past trajectories and feeding them back to the agent improved success rates by 24.6% and 51.1% (relative) on the Mind2Web and WebArena web-navigation benchmarks. In practice, procedural memory often lives as workflow libraries or automatic updates to the agent's own instructions.

AI Agent Memory Architecture: A Reference Design

Here's a reference architecture that generalizes across nearly every serious implementation.

AI agent memory architecture diagram showing short-term working memory in the context window versus long-term memory in external stores, connected by read and write paths

The write path: extract → consolidate → store

After each turn or completed task, the system extracts candidate memories (facts stated, preferences implied, outcomes observed). Then it consolidates: deciding whether each candidate is new (add), refines an existing memory (update), or contradicts one (delete or supersede). This add/update/delete step is the heart of Mem0's published pipeline and the reason extraction-based systems stay compact. Mem0's paper reports keeping memory to roughly 1,800 tokens per conversation, versus around 26,000 tokens for replaying the full history.

The read path: retrieve → rank → inject

At task time, the system retrieves candidates by semantic similarity, entity, or time window, then ranks and filters by relevance, recency, and scope before injecting the survivors into the context. Ranking is where quality is won or lost: retrieval without recency weighting is how an agent congratulates a user on a job they left last year.

Forgetting and scope are features

Two properties separate production systems from demos. 

Forgetting (TTLs, decay scores, and contradiction-driven deletion) keeps retrieval precision from degrading as the store grows. 

Scoping (namespacing memories by user, session, and tenant) is non-negotiable in multi-tenant products, where one customer's retrieved memory leaking into another's context is a security incident, not a quirk.

Agent Memory Systems Compared: Mem0, Letta, and Zep

Three names dominate the current ecosystem, and they represent different philosophies rather than three flavors of the same product.

For where these sit next to retrieval, protocols, and observability, see the full context engineering stack.

System

Core approach

Memory primitive

Open source

Strongest fit

Mem0

Extraction pipeline layered onto any agent

Facts with add/update/delete ops (+ optional graph)

Apache 2.0

Personalization, fast integration

Letta

Memory built into the agent runtime

Self-editing memory blocks (MemGPT lineage)

Apache 2.0

Long-horizon, self-improving agents

Zep

Temporal knowledge graph service

Graph facts with validity time windows

Graphiti engine is open source; Zep is a managed service

Temporal reasoning, evolving enterprise data

Mem0 is a memory layer you bolt onto an existing agent. It watches conversations, extracts salient facts, and consolidates them through explicit add/update/delete operations. Its ECAI 2025 paper reported 26% higher accuracy than OpenAI's memory on LOCOMO with dramatic latency and token savings, and a graph variant (Mem0ᵍ) for relational queries.

Letta, from the Berkeley team behind MemGPT, takes the opposite stance: memory isn't a sidecar, it's the runtime. Agents hold self-editing memory blocks in context and use tools to page information between in-context and external storage. The agent manages its own memory, which suits long-running autonomous agents that should improve over time.

Zep is built around Graphiti, a temporal knowledge graph that stores facts with validity windows. It knows not just that something is true, but when it was true and when it stopped being true. The Zep paper reports up to 18.5% accuracy improvements with 90% lower latency versus baseline implementations on LongMemEval, the benchmark most focused on temporal reasoning. Note that Zep retired its self-hosted Community Edition in 2025; Graphiti itself remains open source.

Where Memory Stops: Governed Data Context

An agent helping a customer with analytics "remembers" that ARR was €4.2M, because the user mentioned it in June. It's now September, the number has changed, and the agent keeps serving the stale figure with total confidence. Memory did its job perfectly. However, the architecture put the wrong thing in it.

Memory stores what the agent has learned about working with you. Governed context serves what is currently true about your data. Preferences, corrections, and workflows belong in memory. Metrics, records, and anything with a compliance or freshness requirement belong behind a semantic layer that defines them once, enforces role-based and row-level permissions on every query, and is hit live, never cached into a memory store where it will rot and escape access controls.

This split matters most in customer-facing, multi-tenant products, where memory systems have no native permission model but your data absolutely does. It's the architecture we build on at Upsolve.ai: our embedded analytics agents persist each user's preferences and context across sessions, while every number they present is resolved at query time through a governed semantic layer with row-level security, so the agent gets more personal over time without ever becoming a second, ungoverned copy of your database.

Frequently Asked Questions

What is AI agent memory? 

AI agent memory is the architecture that lets an agent store, consolidate, and retrieve information across sessions instead of forgetting everything when the context window resets. It spans short-term working memory (inside the prompt) and long-term stores (vector databases, knowledge graphs, key-value stores) connected by write and read paths.

What's the difference between agent memory and RAG? 

RAG retrieves from a mostly static document corpus; agent memory retrieves from a store the agent itself continuously writes to and revises. Memory systems add extraction, consolidation (update/delete), temporal awareness, and per-user scoping that standard RAG pipelines lack.

How do I add long-term memory to an AI agent? 

Start with the loop: extract candidate facts after each turn, consolidate them (add/update/delete) into a vector or graph store scoped per user, then retrieve and rank the top few per task. Frameworks like Mem0, Letta, and Zep implement this loop so you don't build it from scratch — but evaluate against your own conversations before choosing.

Is Mem0, Letta, or Zep better? 

None wins universally. Mem0 is the fastest to integrate for personalization; Letta suits long-running agents that manage their own memory; Zep is strongest when facts change over time and temporal reasoning matters. Vendor benchmarks conflict, so run LongMemEval or LoCoMo on your own data.

What is agent workflow memory? 

Agent workflow memory is procedural memory: reusable action sequences induced from an agent's past successful trajectories and fed back as guidance. Wang et al.'s research showed it lifted web-task success rates by 24.6–51.1% relative to baselines.

Should agents memorize business data? 

No. Facts with freshness or permission requirements (revenue, records, customer data) should be queried live through a governed semantic layer with access controls, not written into memory. Memory should hold learned preferences and workflows, which is what keeps agents both personal and trustworthy.

Sources

  1. Packer et al. (UC Berkeley): MemGPT: Towards LLMs as Operating Systems. The paper that framed context as RAM and memory as paged storage. https://arxiv.org/abs/2310.08560

  2. Liu et al. (Stanford): Lost in the Middle: How Language Models Use Long Contexts. https://arxiv.org/abs/2307.03172

  3. Wu et al.: LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. https://arxiv.org/abs/2410.10813

  4. Chhikara et al. (Mem0, ECAI 2025): Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. https://arxiv.org/abs/2504.19413

  5. Rasmussen et al. (Zep): Zep: A Temporal Knowledge Graph Architecture for Agent Memory. https://arxiv.org/abs/2501.13956

  6. Wang et al. (CMU): Agent Workflow Memory. https://arxiv.org/abs/2409.07429

  7. Maharana et al.: Evaluating Very Long-Term Conversational Memory of LLM Agents (LoCoMo). https://arxiv.org/abs/2402.17753

  8. Mem0: Open-source repository and documentation. https://github.com/mem0ai/mem0

  9. Letta: Documentation for the Letta (formerly MemGPT) agent runtime. https://docs.letta.com

  10. Zep / Graphiti: Graphiti temporal knowledge graph engine. https://github.com/getzep/graphiti

Try Upsolve for Embedded Dashboards & AI Insights

Embed dashboards and AI insights directly into your product, with no heavy engineering required.

Fast setup

Built for SaaS products

30‑day free trial

See Upsolve in Action

Launch customizable dashboards and AI‑powered insights inside your app, fast and with minimal engineering effort. No code.

Follow us

Related Articles

Stop answering the same 10 questions today.

The Platform for Accurate, Reliable, and Trustworthy AI Analytics.

Agent Studio for Data Teams. Encode context. Deploy agents. Deliver clarity.

© 2026 Upsolve AI, Inc.  

|