Agentic context engineering explained: the Stanford ACE paper, Anthropic's guidance, and a write/select/compress/isolate framework applied to data agents.

Ka Ling Wu
Co-Founder & CEO, Upsolve AI
10 min
Agentic context engineering is the process of controlling what information enters an AI agent's context window at each step of its work: which instructions, tools, retrieved data, and history it sees, and in what form.
In late 2025, a Stanford-led research paper literally titled Agentic Context Engineering, and Anthropic's widely shared essay on effective context engineering for AI agents.
This article summarizes both, then turns them into a framework you can actually apply, including to the hardest case: agents that query data.
Key Takeaways |
|---|
|
What Is Agentic Context Engineering?
Prompt engineering asks "what should I say to the model?" Context engineering asks a broader question: of everything that could go into the model's context window (system prompt, tool definitions, retrieved documents, message history, tool outputs) what actually should, at this specific step?
An agent runs in a loop, and every loop iteration generates new material: tool results, intermediate conclusions, errors, dead ends. As Anthropic puts it, the curation problem repeats every time you decide what to pass to the model. Context engineering for agents is therefore a runtime discipline, not a design-time one.
In October 2025, researchers from Stanford, SambaNova, and UC Berkeley published a paper that named a specific framework ACE: Agentic Context Engineering. In ACE, the agent itself does the engineering: it evolves its own context based on what worked. Both meanings are worth understanding, because the paper's findings change how you should practice the broader discipline.
The ACE Paper: Contexts That Improve Themselves
The core idea of the ACE paper (Zhang et al., published at ICLR 2026) is that you can improve an LLM system without touching its weights. Instead of fine-tuning, you adapt the context: the instructions, strategies, and accumulated evidence the model reads before acting.
That idea wasn't new. What the paper contributes is a diagnosis of why previous attempts at self-improving contexts failed, and an architecture that avoids those failures.
Two Failure Modes: Brevity Bias and Context Collapse
Earlier approaches had a model periodically rewrite its own instructions into a cleaner summary. The ACE authors found two problems with this:
Brevity bias: when a model summarizes, it optimizes for concision and drops the domain-specific details that were doing the real work. The summary reads better and performs worse.
Context collapse: when a monolithic context gets rewritten end-to-end, each rewrite erodes a little detail. Over many iterations the context degrades toward mush, and performance drops with it.
If you've ever watched a long agent session lose the plot right after an auto-summarization step, you've seen a small-scale version of both.
The Fix: Playbooks, Not Summaries
ACE treats context as an evolving playbook: a structured collection of itemized strategies that gets updated incrementally, never rewritten wholesale. Three roles share the work:
Generator: runs the task and produces trajectories, including failures.
Reflector: examines those trajectories and extracts concrete lessons: what helped, what hurt.
Curator: merges the lessons into the playbook as delta updates, adding, refining, or pruning individual entries.
Because updates are localized, detail accumulates instead of eroding. The reported results are strong: +10.6% over context-adaptation baselines on agent benchmarks and +8.6% on financial reasoning tasks, with roughly 87% lower adaptation latency. On the AppWorld leaderboard, an agent running ACE on open-source DeepSeek-V3.1 matched the top-ranked GPT-4.1-powered production agent on the overall average, and beat it on the harder test-challenge split. Notably, ACE worked without labeled training data: execution feedback (did the code run, did the answer check out) was enough signal.
When managing agent context, prefer structured, incremental edits over wholesale rewrites, and be suspicious of any process that makes context shorter by making it vaguer. ACE is the clearest published example of a self-improving agent loop in action.
Anthropic's Playbook: Effective Context Engineering for AI Agents
Where ACE is about agents curating their own context, Anthropic's essay on effective context engineering for AI agents is about engineers curating it. Its central claim: context is a depleting resource. Transformer attention has to cover n² pairwise relationships between tokens, so as context grows, the model's ability to use any given token accurately declines. Chroma's research on context rot documents the pattern empirically across 18 models: reliability drops as input length grows, even on deliberately simple tasks.
The design principle that follows is the most quotable line in the piece: aim for the smallest set of high-signal tokens that gets the outcome you want. Concretely, the essay recommends:
System prompts at the right altitude. Not brittle if-else logic hardcoded into prose, and not vague vibes either. Specific enough to steer, flexible enough to let the model generalize.
Minimal, unambiguous tools. If a human engineer can't say which of two tools applies to a situation, the agent won't do better. Bloated tool sets are a top failure mode.
Just-in-time retrieval over pre-loading. Give the agent lightweight references (file paths, stored queries, links) and tools to pull data when needed, rather than stuffing everything into context up front. This mirrors how people use file systems and bookmarks instead of memorizing corpuses.
Three techniques for long-horizon work: compaction (summarize a near-full context and restart with the summary), structured note-taking (persist notes outside the window and re-read them), and sub-agent architectures, where workers explore with tens of thousands of tokens but return distilled summaries of one or two thousand. That last pattern is the one behind Anthropic's multi-agent research system, which they report beat a single-agent baseline by 90.2% on their internal research eval.
A Practical Framework: Write, Select, Compress, Isolate
The cleanest way to organize all of this is the four-operation taxonomy popularized by LangChain. Every technique in both papers is one of these four moves.
Write: Persist Context Outside the Window
Writing means saving information somewhere durable so it survives beyond the current window: scratchpads, NOTES.md files, memory tools, or ACE-style playbooks. The test of a good write strategy is whether the agent can crash, restart with a fresh context, read its own notes, and continue. ACE's contribution here is how to write: itemized entries with incremental updates, so knowledge accumulates instead of getting paraphrased away.
Select: Pull In Only What This Step Needs
Selection covers retrieval in all its forms: RAG, file reads, memory lookups, tool results. The shift Anthropic describes is from pre-inference retrieval ("embed everything, stuff top-k chunks into the prompt") toward agentic search, where the model navigates with lightweight identifiers and loads data on demand. Metadata does real work here: a file's name, location, and timestamp tell the agent whether to open it at all.
Compress: Shrink What's Already There
Compression retains only the tokens still earning their place. The safest version is tool-result clearing: an agent rarely needs the raw output of a tool it called forty turns ago. Full compaction (summarize and restart) is the heavier lever. Tune it the way Anthropic suggests: maximize recall first so nothing critical is lost, then trim for precision. And remember the ACE caveat: compression is for history, not for curated knowledge.
Isolate: Split Context Across Boundaries
Isolation prevents one task's context from polluting another's. Sub-agents are the main pattern: each worker gets a clean window, burns as many tokens as it needs, and hands back only a distilled result. Isolation is also where security lives, which matters more than most write-ups admit; a boundary that keeps context clean can be the same boundary that keeps data contained.
Operation | Core question | Key techniques | Main risk |
Write | What must survive this window? | Scratchpads, memory files, playbooks | Notes nobody re-reads |
Select | What does this step need? | Agentic search, RAG, just-in-time loading | Retrieving too much |
Compress | What has stopped earning its place? | Tool-result clearing, compaction | Brevity bias, context collapse |
Isolate | What must not mix? | Sub-agents, sandboxes, tenancy walls | Losing cross-task coherence |
Context Engineering for Data and Analytics Agents
Most context engineering content uses coding agents as the example. Data and analytics agents are harder, and they make the four operations concrete in a way code doesn't. Here's what each looks like when the agent's job is answering questions about a database. This is our world at Upsolve.ai, so these failure modes come from experience.
Select is one of the most significant steps. The naive approach dumps the full schema into the system prompt. A real production warehouse has hundreds of tables and thousands of columns; that's context rot by design, and the agent starts joining on the wrong keys because three tables all have a customer_id. The fix is a semantic layer: curated definitions of metrics, dimensions, and blessed join paths that the agent selects from, instead of raw DDL. It's Anthropic's "right altitude" principle applied to data: not raw schema (too low), not "you can query the warehouse" (too high), but named business concepts with unambiguous meanings.
Compress or drown. A single SELECT * can return more tokens than the entire rest of the context. Analytics agents need the pattern Anthropic describes Claude Code using: write a targeted query, persist the full result outside the window, and pull only aggregates, samples, or shapes into context. The agent reasons over "23,412 rows, revenue concentrated in two segments," not over 23,412 rows.
Write what the data actually means. The highest-value thing a data agent can persist is tribal knowledge: revenue excludes refunds, the events table is unreliable before March, fiscal year starts in February. This is exactly the ACE playbook pattern, and it's where brevity bias hurts most. A summarizer will compress "exclude test accounts (account_id < 1000) from all revenue queries" into "clean the data first," and every future query is silently wrong.
Isolate along tenant lines. In customer-facing analytics, isolation isn't a performance optimization, it's the product. Customer A's question must never be answered with context from customer B's data: not their schema annotations, not their cached results, not their learned playbook entries. The sub-agent boundary and the security boundary should be the same wall.
Where Context Engineering for Agents Goes Wrong
Three mistakes account for most failures we see:
Treating a bigger context window as the fix. Million-token windows move the cliff; they don't remove the gradient. Recall still degrades with length, and cost scales with every token you didn't need.
Summarizing the wrong things. Compacting redundant history: good. Compacting curated instructions and learned strategies: that's how you get context collapse in slow motion.
Engineering context once, at design time. The defining property of agents is that context changes every step. If your "context strategy" is a static prompt template, you've done prompt engineering and renamed it.
Frequently Asked Questions
What is agentic context engineering?
It's the practice of deciding what information an AI agent sees in its context window at each step of a task: instructions, tools, retrieved data, and history. The term also refers to ACE, a specific Stanford-led framework in which agents incrementally evolve their own context playbooks based on execution feedback.
How is context engineering different from prompt engineering?
Prompt engineering optimizes the wording of instructions, typically written once. Context engineering manages the entire, constantly changing token state of an agent across a multi-step task, including tool outputs, retrieved data, and memory. Anthropic frames it as the natural progression of prompt engineering as systems move from single completions to agents.
What is context rot?
Context rot is the measured decline in a model's ability to accurately recall and use information as its context window fills up. Chroma's benchmark research found the degradation across all models tested, which is why "just add more context" fails as a strategy.
Our context rot deep dive covers the evidence and the fixes.
What are brevity bias and context collapse?
Both are failure modes of letting a model summarize its own context, identified in the ACE paper. Brevity bias is the tendency of summaries to drop specific, high-value detail in favor of concision. Context collapse is the compounding erosion of detail when a context is fully rewritten many times.
Does agentic context engineering replace fine-tuning?
For many adaptation tasks it's now the cheaper first thing to try. The ACE paper positions context adaptation as an alternative to weight updates and shows it beating other context-adaptation methods with far less latency, working from execution feedback alone with no labeled dataset. It doesn't benchmark head-to-head against fine-tuning, though, so fine-tuning still makes sense when the behavior you need can't be reached through context at all.
How do I apply context engineering to a data or analytics agent?
Start with select: replace raw schema dumps with a semantic layer of curated metrics and join paths. Then compress: keep full query results outside the context and load only aggregates or samples. Persist data caveats and business definitions as an incrementally updated playbook, and enforce hard per-tenant isolation so one customer's context never leaks into another's.
Sources
Zhang et al., Stanford / SambaNova / UC Berkeley: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (ICLR 2026). https://arxiv.org/abs/2510.04618
Anthropic: Effective Context Engineering for AI Agents. https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
LangChain: Context Engineering for Agents (write / select / compress / isolate taxonomy). https://www.langchain.com/blog/context-engineering-for-agents
Chroma Research: Context Rot: How Increasing Input Tokens Impacts LLM Performance. https://www.trychroma.com/research/context-rot
Anthropic: How We Built Our Multi-Agent Research System. https://www.anthropic.com/engineering/multi-agent-research-system

Try Upsolve for Embedded Dashboards & AI Insights
Embed dashboards and AI insights directly into your product, with no heavy engineering required.
Fast setup
Built for SaaS products
30‑day free trial






