Agentic Context Engineering: ACE, Anthropic & How to Do It

Agentic Context Engineering: ACE, Anthropic & How to Do It

Agentic Context Engineering: ACE, Anthropic & How to Do It

All Posts

Agentic context engineering explained: the Stanford ACE paper, Anthropic's guidance, and a write/select/compress/isolate framework applied to data agents.

Ka Ling Wu

Co-Founder & CEO, Upsolve AI

10 min

Agentic context engineering diagram: write, select, compress, and isolate operations feeding an AI agent's context window

Agentic context engineering is the process of controlling what information enters an AI agent's context window at each step of its work: which instructions, tools, retrieved data, and history it sees, and in what form. 

In late 2025, a Stanford-led research paper literally titled Agentic Context Engineering, and Anthropic's widely shared essay on effective context engineering for AI agents. 

This article summarizes both, then turns them into a framework you can actually apply, including to the hardest case: agents that query data.

Key Takeaways

  • "Agentic context engineering" means two related things: the general discipline of curating an agent's context, and ACE, a specific Stanford framework where agents improve their own context. Most searchers conflate them; this article covers both.

  • Context is a finite resource, not free storage: Anthropic's engineering team argues that model recall degrades as context grows ("context rot"), so the goal is the smallest set of high-signal tokens, not the biggest context window.

  • Summarization can quietly destroy agents: the ACE paper identifies "brevity bias" and "context collapse" as failure modes of naive context management, and reports +10.6% on agent benchmarks by treating context as an incremental playbook instead.

  • Four operations cover most of the practice: write, select, compress, isolate. This taxonomy, popularized by LangChain, maps cleanly onto everything both papers recommend.

  • Data agents are the stress test: schemas, query results, and multi-tenant boundaries make analytics the domain where context engineering fails loudest, and where it pays off fastest.

What Is Agentic Context Engineering?

Prompt engineering asks "what should I say to the model?" Context engineering asks a broader question: of everything that could go into the model's context window (system prompt, tool definitions, retrieved documents, message history, tool outputs) what actually should, at this specific step?

An agent runs in a loop, and every loop iteration generates new material: tool results, intermediate conclusions, errors, dead ends. As Anthropic puts it, the curation problem repeats every time you decide what to pass to the model. Context engineering for agents is therefore a runtime discipline, not a design-time one.

In October 2025, researchers from Stanford, SambaNova, and UC Berkeley published a paper that named a specific framework ACE: Agentic Context Engineering. In ACE, the agent itself does the engineering: it evolves its own context based on what worked. Both meanings are worth understanding, because the paper's findings change how you should practice the broader discipline.

The ACE Paper: Contexts That Improve Themselves

The core idea of the ACE paper (Zhang et al., published at ICLR 2026) is that you can improve an LLM system without touching its weights. Instead of fine-tuning, you adapt the context: the instructions, strategies, and accumulated evidence the model reads before acting.

That idea wasn't new. What the paper contributes is a diagnosis of why previous attempts at self-improving contexts failed, and an architecture that avoids those failures.

Two Failure Modes: Brevity Bias and Context Collapse

Earlier approaches had a model periodically rewrite its own instructions into a cleaner summary. The ACE authors found two problems with this:

  • Brevity bias: when a model summarizes, it optimizes for concision and drops the domain-specific details that were doing the real work. The summary reads better and performs worse.

  • Context collapse: when a monolithic context gets rewritten end-to-end, each rewrite erodes a little detail. Over many iterations the context degrades toward mush, and performance drops with it.

If you've ever watched a long agent session lose the plot right after an auto-summarization step, you've seen a small-scale version of both.

The Fix: Playbooks, Not Summaries

ACE treats context as an evolving playbook: a structured collection of itemized strategies that gets updated incrementally, never rewritten wholesale. Three roles share the work:

  1. Generator: runs the task and produces trajectories, including failures.

  2. Reflector: examines those trajectories and extracts concrete lessons: what helped, what hurt.

  3. Curator: merges the lessons into the playbook as delta updates, adding, refining, or pruning individual entries.

Because updates are localized, detail accumulates instead of eroding. The reported results are strong: +10.6% over context-adaptation baselines on agent benchmarks and +8.6% on financial reasoning tasks, with roughly 87% lower adaptation latency. On the AppWorld leaderboard, an agent running ACE on open-source DeepSeek-V3.1 matched the top-ranked GPT-4.1-powered production agent on the overall average, and beat it on the harder test-challenge split. Notably, ACE worked without labeled training data: execution feedback (did the code run, did the answer check out) was enough signal.

When managing agent context, prefer structured, incremental edits over wholesale rewrites, and be suspicious of any process that makes context shorter by making it vaguer. ACE is the clearest published example of a self-improving agent loop in action.

Anthropic's Playbook: Effective Context Engineering for AI Agents

Where ACE is about agents curating their own context, Anthropic's essay on effective context engineering for AI agents is about engineers curating it. Its central claim: context is a depleting resource. Transformer attention has to cover n² pairwise relationships between tokens, so as context grows, the model's ability to use any given token accurately declines. Chroma's research on context rot documents the pattern empirically across 18 models: reliability drops as input length grows, even on deliberately simple tasks.

The design principle that follows is the most quotable line in the piece: aim for the smallest set of high-signal tokens that gets the outcome you want. Concretely, the essay recommends:

  • System prompts at the right altitude. Not brittle if-else logic hardcoded into prose, and not vague vibes either. Specific enough to steer, flexible enough to let the model generalize.

  • Minimal, unambiguous tools. If a human engineer can't say which of two tools applies to a situation, the agent won't do better. Bloated tool sets are a top failure mode.

  • Just-in-time retrieval over pre-loading. Give the agent lightweight references (file paths, stored queries, links) and tools to pull data when needed, rather than stuffing everything into context up front. This mirrors how people use file systems and bookmarks instead of memorizing corpuses.

  • Three techniques for long-horizon work: compaction (summarize a near-full context and restart with the summary), structured note-taking (persist notes outside the window and re-read them), and sub-agent architectures, where workers explore with tens of thousands of tokens but return distilled summaries of one or two thousand. That last pattern is the one behind Anthropic's multi-agent research system, which they report beat a single-agent baseline by 90.2% on their internal research eval.

A Practical Framework: Write, Select, Compress, Isolate

The cleanest way to organize all of this is the four-operation taxonomy popularized by LangChain. Every technique in both papers is one of these four moves.

Write: Persist Context Outside the Window

Writing means saving information somewhere durable so it survives beyond the current window: scratchpads, NOTES.md files, memory tools, or ACE-style playbooks. The test of a good write strategy is whether the agent can crash, restart with a fresh context, read its own notes, and continue. ACE's contribution here is how to write: itemized entries with incremental updates, so knowledge accumulates instead of getting paraphrased away.

Select: Pull In Only What This Step Needs

Selection covers retrieval in all its forms: RAG, file reads, memory lookups, tool results. The shift Anthropic describes is from pre-inference retrieval ("embed everything, stuff top-k chunks into the prompt") toward agentic search, where the model navigates with lightweight identifiers and loads data on demand. Metadata does real work here: a file's name, location, and timestamp tell the agent whether to open it at all.

Compress: Shrink What's Already There

Compression retains only the tokens still earning their place. The safest version is tool-result clearing: an agent rarely needs the raw output of a tool it called forty turns ago. Full compaction (summarize and restart) is the heavier lever. Tune it the way Anthropic suggests: maximize recall first so nothing critical is lost, then trim for precision. And remember the ACE caveat: compression is for history, not for curated knowledge.

Isolate: Split Context Across Boundaries

Isolation prevents one task's context from polluting another's. Sub-agents are the main pattern: each worker gets a clean window, burns as many tokens as it needs, and hands back only a distilled result. Isolation is also where security lives, which matters more than most write-ups admit; a boundary that keeps context clean can be the same boundary that keeps data contained.

Operation

Core question

Key techniques

Main risk

Write

What must survive this window?

Scratchpads, memory files, playbooks

Notes nobody re-reads

Select

What does this step need?

Agentic search, RAG, just-in-time loading

Retrieving too much

Compress

What has stopped earning its place?

Tool-result clearing, compaction

Brevity bias, context collapse

Isolate

What must not mix?

Sub-agents, sandboxes, tenancy walls

Losing cross-task coherence

Context Engineering for Data and Analytics Agents

Most context engineering content uses coding agents as the example. Data and analytics agents are harder, and they make the four operations concrete in a way code doesn't. Here's what each looks like when the agent's job is answering questions about a database. This is our world at Upsolve.ai, so these failure modes come from experience.

Select is one of the most significant steps. The naive approach dumps the full schema into the system prompt. A real production warehouse has hundreds of tables and thousands of columns; that's context rot by design, and the agent starts joining on the wrong keys because three tables all have a customer_id. The fix is a semantic layer: curated definitions of metrics, dimensions, and blessed join paths that the agent selects from, instead of raw DDL. It's Anthropic's "right altitude" principle applied to data: not raw schema (too low), not "you can query the warehouse" (too high), but named business concepts with unambiguous meanings.

Compress or drown. A single SELECT * can return more tokens than the entire rest of the context. Analytics agents need the pattern Anthropic describes Claude Code using: write a targeted query, persist the full result outside the window, and pull only aggregates, samples, or shapes into context. The agent reasons over "23,412 rows, revenue concentrated in two segments," not over 23,412 rows.

Write what the data actually means. The highest-value thing a data agent can persist is tribal knowledge: revenue excludes refunds, the events table is unreliable before March, fiscal year starts in February. This is exactly the ACE playbook pattern, and it's where brevity bias hurts most. A summarizer will compress "exclude test accounts (account_id < 1000) from all revenue queries" into "clean the data first," and every future query is silently wrong.

Isolate along tenant lines. In customer-facing analytics, isolation isn't a performance optimization, it's the product. Customer A's question must never be answered with context from customer B's data: not their schema annotations, not their cached results, not their learned playbook entries. The sub-agent boundary and the security boundary should be the same wall.

Where Context Engineering for Agents Goes Wrong

Three mistakes account for most failures we see:

  • Treating a bigger context window as the fix. Million-token windows move the cliff; they don't remove the gradient. Recall still degrades with length, and cost scales with every token you didn't need.

  • Summarizing the wrong things. Compacting redundant history: good. Compacting curated instructions and learned strategies: that's how you get context collapse in slow motion.

  • Engineering context once, at design time. The defining property of agents is that context changes every step. If your "context strategy" is a static prompt template, you've done prompt engineering and renamed it.

Frequently Asked Questions

What is agentic context engineering? 

It's the practice of deciding what information an AI agent sees in its context window at each step of a task: instructions, tools, retrieved data, and history. The term also refers to ACE, a specific Stanford-led framework in which agents incrementally evolve their own context playbooks based on execution feedback.

How is context engineering different from prompt engineering? 

Prompt engineering optimizes the wording of instructions, typically written once. Context engineering manages the entire, constantly changing token state of an agent across a multi-step task, including tool outputs, retrieved data, and memory. Anthropic frames it as the natural progression of prompt engineering as systems move from single completions to agents.

What is context rot? 

Context rot is the measured decline in a model's ability to accurately recall and use information as its context window fills up. Chroma's benchmark research found the degradation across all models tested, which is why "just add more context" fails as a strategy.

Our context rot deep dive covers the evidence and the fixes.

What are brevity bias and context collapse? 

Both are failure modes of letting a model summarize its own context, identified in the ACE paper. Brevity bias is the tendency of summaries to drop specific, high-value detail in favor of concision. Context collapse is the compounding erosion of detail when a context is fully rewritten many times.

Does agentic context engineering replace fine-tuning? 

For many adaptation tasks it's now the cheaper first thing to try. The ACE paper positions context adaptation as an alternative to weight updates and shows it beating other context-adaptation methods with far less latency, working from execution feedback alone with no labeled dataset. It doesn't benchmark head-to-head against fine-tuning, though, so fine-tuning still makes sense when the behavior you need can't be reached through context at all.

How do I apply context engineering to a data or analytics agent? 

Start with select: replace raw schema dumps with a semantic layer of curated metrics and join paths. Then compress: keep full query results outside the context and load only aggregates or samples. Persist data caveats and business definitions as an incrementally updated playbook, and enforce hard per-tenant isolation so one customer's context never leaks into another's.

Sources

  1. Zhang et al., Stanford / SambaNova / UC Berkeley: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (ICLR 2026). https://arxiv.org/abs/2510.04618

  2. Anthropic: Effective Context Engineering for AI Agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents

  3. LangChain: Context Engineering for Agents (write / select / compress / isolate taxonomy). https://www.langchain.com/blog/context-engineering-for-agents

  4. Chroma Research: Context Rot: How Increasing Input Tokens Impacts LLM Performance. https://www.trychroma.com/research/context-rot

  5. Anthropic: How We Built Our Multi-Agent Research System. https://www.anthropic.com/engineering/multi-agent-research-system

Try Upsolve for Embedded Dashboards & AI Insights

Embed dashboards and AI insights directly into your product, with no heavy engineering required.

Fast setup

Built for SaaS products

30‑day free trial

See Upsolve in Action

Launch customizable dashboards and AI‑powered insights inside your app, fast and with minimal engineering effort. No code.

Follow us

Related Articles

Stop answering the same 10 questions today.

The Platform for Accurate, Reliable, and Trustworthy AI Analytics.

Agent Studio for Data Teams. Encode context. Deploy agents. Deliver clarity.

© 2026 Upsolve AI, Inc.  

|