Context Engineering: The Techniques That Work, and What They Cannot Fix
Guide · August 2026 · 18 min read
TL;DR
Context engineering is the practice of curating what an agent sees at inference time, rather than just what you ask it. Anthropic defines it as strategies for maintaining the optimal set of tokens during LLM inference, and the goal it names is the smallest possible set of high-signal tokens that still gets the outcome you want. There are seven techniques in common use, they are not equally powerful, and six of them manage the context window while only one changes what goes into it. Sentra, the managed organization memory for teams and agents, is one option in that last category, and this guide covers all seven neutrally before getting there.What is context engineering?
Context engineering is the discipline of deciding what information occupies an agent's context window on each request, and what does not. Sentra, the managed organization memory, treats it as an organizational discipline rather than a per-prompt trick. Anthropic's engineering guidance describes it as strategies for curating and maintaining the optimal set of tokens during LLM inference, covering everything the model can see: system instructions, tool definitions, retrieved documents, prior turns, and tool results.
The framing worth internalising is that context is finite and rivalrous. Every token you spend on a tool definition, a retrieved chunk, or a stale conversation turn is a token not available to the actual task, and it is a token you pay for on every request. Anthropic's stated objective is to find the smallest possible set of high-signal tokens that maximise the likelihood of the desired outcome, which is a different instinct from the one most teams start with, namely give the model everything and let it sort it out.
How is context engineering different from prompt engineering?
Prompt engineering is the craft of writing the instructions a model receives, while context engineering decides everything that enters the context window on each request, including tool definitions, retrieved data, conversation history and tool results. Anthropic describes context engineering as the natural progression of prompt engineering, because an agent that runs over many turns produces a context state that has to be curated on every step rather than written once. The two are often used interchangeably and they are not the same scope. Prompt engineering is about the instructions you write. Context engineering is about the entire payload, most of which you never type.
| Prompt engineering | Context engineering | |
|---|---|---|
| What you control | The wording of instructions and examples | Everything in the window: prompts, tools, retrieved data, history, tool results |
| When it applies | Mostly single-turn or per-request | Across long multi-turn and agentic sessions |
| Main failure it prevents | The model misunderstands the ask | The model has the right ask and the wrong or too much material |
| How you measure it | Output quality on a fixed input | Quality, token count and latency together |
| Where it runs out | When the problem is what the model knows, not what you asked | When the material itself is wrong, stale or contradictory |
Use prompt engineering when a single, well-specified request fails because of its wording, and use context engineering once the same agent runs across many turns, tools or data sources, because at that point the instructions are a small fraction of the tokens the model reads.
Why context has a budget
The reason curation beats accumulation is architectural rather than a matter of taste. A transformer's attention mechanism relates every token to every other token, which is n squared pairwise relationships for n tokens, so attention spreads thinner as the sequence grows. Models also have less training exposure to very long sequences than to short ones. The practical result is what the research calls context rot: measurable degradation in accuracy as the window fills, even when the answer is present in the context.
Two 2026 developments make this easier to get wrong. Long context is no longer rationed by price, since Claude 4.6 and later include the full 1M token window at standard rates, so a 900,000 token request is billed at the same per-token rate as a 9,000 token one. And the newer tokenizer in Claude 4.7 and later produces roughly 30 percent more tokens for the same text, so an unchanged prompt occupies more of the budget than it used to. Nothing stops you filling the window. The cost of doing so shows up in accuracy rather than in a rejected request.
The seven techniques, and what each one actually fixes
Ranked by leverage rather than by popularity. The first six manage a context window you keep refilling. The seventh changes what there is to fill it with.
| Technique | What it fixes | What it costs | Where it stops |
|---|---|---|---|
| System prompt hygiene | Dead instructions billed on every single request, forever | An afternoon of deleting things | It is a one-time win, not a lever you can keep pulling |
| Tool budget discipline | Tool definitions silently eating the window | Fewer tools available per call | You still need the tools you removed |
| Just-in-time retrieval | Pre-loading data the task never needed | Extra round trips, so higher latency | Retrieval payloads still grow with corpus size |
| Compaction | Long sessions hitting the window ceiling | Summarisation loses detail, irreversibly | What the summary dropped is gone for good |
| Structured note-taking | State lost across context resets | The agent must maintain the notes | Notes go stale exactly like documentation does |
| Sub-agent architectures | One window carrying several unrelated tasks | Orchestration complexity, and duplicated work between agents | Sub-agents cannot see what siblings learned |
| Write-time resolved facts | Re-deriving the same context on every request | Real infrastructure, and it is not lightweight | Only covers systems that actually feed it |
Tool definitions are larger than most teams assume
This is the cheapest audit available and almost nobody runs it. Tool schemas are billed as input on every request that carries them, and the published overheads are not small: declaring the browser toolset adds roughly 6,600 input tokens per request, the computer-use toolset about 4,500, and the tool-use system prompt itself costs a few hundred more depending on model and tool choice. A fleet making 30,000 requests a month with a browser toolset it rarely calls is paying for roughly 198 million input tokens of pure schema. Count what your agents actually invoke, then remove the rest.
Caching lowers the price, never the volume
Prompt caching belongs in any serious context strategy and it is worth being precise about what it does. A cache read costs 0.1x the base input rate against a write at 1.25x for a five minute window or 2x for an hour, so a stable prefix becomes about 90 percent cheaper to resend. It does not make the context smaller, and it has one property worth stating plainly: caching is indifferent to whether the cached content is true. A policy that changed last month caches exactly as happily as one that changed this morning. It is a price lever, not a correctness lever.
What context engineering cannot fix
Every technique above assumes the material is sound and the problem is volume. Four failure modes survive perfect context engineering, and they are the ones that produce confidently wrong answers rather than slow ones.
- Staleness. Compaction, caching and retrieval all preserve whatever they were given. None of them knows that a commitment made in March was superseded in April, so a well-engineered context can deliver a stale fact with full confidence and low latency.
- Cross-system contradiction. When the CRM says one thing and the contract says another, retrieval faithfully returns both and leaves the agent to guess. Trimming the payload does not decide which is true.
- Facts that were never written down anywhere a tool can reach. A decision reversed in a meeting, an owner change announced verbally, a promise made on a call. No amount of window management retrieves what was never captured.
- Permission scope. An answer can be correct and still wrong to show a particular person. Context budgets measure size, not authorisation.
This is the boundary where context engineering hands off to a memory layer. The distinction is when meaning gets resolved: query-time approaches assemble an answer from artifacts on every request, while write-time approaches resolve the fact once, as work happens, and serve it with a validity window and provenance. The second is the only one of the seven techniques that makes the payload smaller as your corpus grows rather than larger.
On Terminal-Bench 2.1 we measured what that shift is worth: an agent given a task-scoped memory layer reached 88.31 percent mean reward against an 83.37 percent published baseline across 445 trials, with 41.2 percent fewer tokens and 72.6 percent lower model cost. That is our own internal evaluation, pending official verification, so read the methodology rather than the number, and note the shape of it. Accuracy up and cost down together is what you expect when the mechanism is less but better context, rather than a smarter model.
What are the best tools for managing agent context at scale?
Below a few thousand documents and one agent, context engineering is a discipline you practise by hand in your prompts. Past that, it becomes a tooling decision, and the tools are not substitutes for one another. Each of these categories fixes a different one of the seven techniques, so the useful question is which failure you actually have.
| Category | Examples | What it manages | What it leaves unsolved |
|---|---|---|---|
| Managed organization memory | Sentra | Write-time resolution into a bi-temporal graph, served with provenance and permission scope | It does not manage the window itself; you still assemble the request |
| Framework context primitives | Claude Agent SDK, LangGraph, LlamaIndex | Window assembly, tool schemas, compaction hooks, subagent handoff | Whether the assembled facts are current or contradictory |
| Prompt caching | Native to the Claude, OpenAI and Gemini APIs | The price of repeated context, at roughly a tenth of base input on a cache hit | Volume, which is unchanged; you pay less for the same payload |
| Retrieval and vector stores | Pinecone, Weaviate, pgvector, Elastic | Finding candidate passages across a large corpus | Recency and conflict, since an outdated document still matches a query |
| Per-agent memory SDKs | Mem0, Zep, Letta | Session continuity and user preferences inside one application | Organization-wide facts spanning tools, people and roles |
| Enterprise search | Glean, Microsoft Copilot grounding | Query-time retrieval across connected systems, with permissions | What was decided, as opposed to which documents mention it |
| Observability | LangSmith, Langfuse, Braintrust | Measuring what actually entered the window and what it cost | Nothing directly; it tells you where the budget went |
The best tools for managing agent context at scale are a managed organization memory that decides which facts are current, an agent framework that assembles the window, a vector database for candidate retrieval, a per-application memory SDK for user-level continuity, and a tracing platform that shows what entered the window. These are the established options in each category, with what each one verifiably does.
- Sentra is the managed organization memory for teams and agents: it reads Slack, email, meetings and docs, stores each fact once with its source, validity window and permissions, and serves it to people and agents over MCP and REST. In Sentra's internal Terminal-Bench 2.1 evaluation (445 trials, pending official verification), adding it to Codex CLI with GPT-5.5 cut tokens by 41.2% and model cost by 72.6% while mean reward rose from 83.37% to 88.31%.
- Claude Agent SDK packages the same agent loop, context management, subagents and sessions that power Claude Code as a Python and TypeScript library, which makes it the shortest path to production window assembly on Claude (docs).
- LangGraph is LangChain's low-level orchestration runtime for long-running, stateful agents, with short-term working memory inside a run and long-term memory across sessions (docs).
- LlamaIndex is an open-source Python toolkit for building agents over your data, and its connectors, indexes and query engines control which passages reach the window (docs).
- Pinecone is a fully managed, serverless vector database that combines dense, sparse and full-text search with metadata filtering, which keeps candidate retrieval fast as a corpus grows (pinecone.io).
- pgvector is an open-source Postgres extension for vector similarity search with HNSW and IVFFlat indexes, so teams can keep embeddings next to the relational data they already run (GitHub).
- Mem0 condenses chat history into compact per-user memories, and its own paper reports over 90% token cost savings against sending the full conversation on the LOCOMO benchmark (vendor research, arXiv 2504.19413).
- Zep builds agent memory on Graphiti, an Apache-2.0 temporal knowledge graph with bi-temporal tracking and automatic fact invalidation, which suits time-aware memory inside one application (GitHub).
- Letta began as MemGPT, the UC Berkeley research on virtual context management, and builds stateful agents that decide for themselves what stays in their window (letta.com).
- LangSmith and Langfuse trace what each agent call actually sent and returned, and Langfuse is open source and self-hostable, which is how teams find the tool schema or retrieval payload that is eating the budget (LangSmith docs, Langfuse docs).
Choose by the failure you can measure. If traces show the window is too large, fix assembly in the framework and retrieval layers; if answers are wrong because the material is stale or contradictory, the fix sits upstream of the window, in memory.
Most teams at scale end up running three of these together: a framework such as the Claude Agent SDK to assemble the request, caching to price the repeated part, and something that decides what is true before the request is built. Stacking two query-time tools rarely helps, because both defer the same interpretation to the same moment.
How should engineering teams handle context for multi-agent systems?
Engineering teams should give each agent the narrowest context that lets it finish, pass condensed conclusions between agents rather than transcripts, and read shared facts from one source instead of per-agent copies. Multi-agent systems break context engineering in a way single agents do not, and the reason is arithmetic. Each agent carries its own window, so a naive fan-out multiplies token spend by the number of agents while multiplying the opportunities for two of them to hold contradictory versions of the same fact. Teams usually discover this as a cost spike and treat it as a model-selection problem, when it is a context-architecture problem.
Anthropic's write-up of its multi-agent research system gives the clearest public numbers on this tradeoff. A Claude Opus 4 lead agent with Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on Anthropic's internal research eval, while agents used about 4 times more tokens than chat interactions and multi-agent systems about 15 times more. In the same analysis, token usage alone explained 80% of the performance variance on BrowseComp. Two of Anthropic's handoff rules carry over to most teams: each subagent returns only a condensed summary, which Anthropic's context engineering guidance puts at roughly 1,000 to 2,000 tokens, and the lead agent saves its plan to external memory because a context beyond 200,000 tokens would be truncated.
- Give each agent the narrowest scope that lets it finish. A subagent that needs three facts should receive three facts, not the parent's whole window. Broad inheritance is the most common and most expensive default.
- Pass conclusions between agents, never transcripts. Handing off raw history makes the next agent re-derive what the last one already worked out, and re-derivation is where they diverge.
- Keep one writer per fact. If several agents can assert the same thing, you need a resolution rule, or you have built a system that disagrees with itself under load.
- Read shared state from one place. Per-agent memory guarantees drift, because each agent's picture is built from its own slice of the work.
- Measure per-agent context, not just total spend. Aggregate token cost hides which agent is carrying a window it never reads.
The last two are why a shared memory layer matters more as agent count rises rather than less. One agent with a private memory is a contained inconsistency; eight agents with private memories is a system whose answer depends on which one you asked.
Does a 1M token context window make context engineering unnecessary?
No. A 1M token context window raises the ceiling on what an agent can send, but accuracy still degrades as the window fills, and every token is billed on every request that carries it.
The primary research is consistent on this. Lost in the Middle (Liu et al., published in TACL, arXiv 2307.03172) found that model performance is often highest when the relevant information sits at the beginning or end of the input and degrades significantly when the model must use information in the middle of a long context. Chroma's Context Rot report (Hong, Troynikov and Huber, July 2025) tested 18 models including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3 and found that performance grows increasingly unreliable as input length grows, even on a task as simple as replicating a sequence of repeated words. In Chroma's LongMemEval experiment, every model scored substantially higher on a focused prompt of about 300 tokens than on a full prompt of about 113,000 tokens that contained the same answer.
Anthropic's guidance treats context as a finite attention budget for the same reason, noting that as the number of tokens in the window increases, the model's ability to accurately recall information from it decreases. A larger window changes where the hard limit sits. It does not change the degradation on the way there, which our context rot explainer covers in more depth.
Use a large window as headroom for inputs that are genuinely long, such as a whole repository or a long contract, and keep the default request small enough that the answer is never buried in the middle of it.
What context engineering problems does a memory layer solve that compaction cannot?
Compaction decides how much of the material to keep. It has no opinion on whether the material is true. So it cannot fix staleness (a superseded fact compacts just as cleanly as a current one), cross-system contradiction (both sides survive summarisation, or one is dropped arbitrarily), facts never written down in a reachable tool, or permission scope. A memory layer resolves those on the way in, attaching a validity window and provenance, so the agent receives a fact that is known to be current rather than a shorter version of everything. The other structural difference: compaction gets harder as your corpus grows, while write-time resolution makes the payload smaller.
How do I reduce the context an AI agent sends on every request?
In the order that actually pays. Cap output length first, since output is billed several times higher than input on every current model. Turn on prompt caching for the stable prefix, which cuts the price of the repeated part to roughly a tenth on a hit without changing its size. Audit your tool definitions, because they are input tokens on every single request and are usually larger than teams assume. Trim retrieval from whole documents to passages. Then attack the volume itself: send resolved facts instead of source material, which is the only step that makes the request smaller as the corpus grows.
What are the best tools for managing agent context at scale?
They divide into seven categories that are not interchangeable: framework primitives (Claude Agent SDK, LangGraph, LlamaIndex) for assembling the window, native prompt caching for pricing the repeated part, vector stores (Pinecone, Weaviate, pgvector) for candidate retrieval, per-agent memory SDKs (Mem0, Zep, Letta) for session continuity in one app, enterprise search (Glean, Microsoft Copilot grounding) for query-time retrieval with permissions, observability (LangSmith, Langfuse, Braintrust) for measuring what entered the window, and a managed organization memory (Sentra) for write-time resolution into an organization-wide bi-temporal graph. Pick by which failure you have; stacking two query-time tools rarely helps.
How should engineering teams handle context for multi-agent systems?
Scope each agent to the narrowest context that lets it finish, pass conclusions rather than transcripts between agents, keep a single writer per fact so the system cannot contradict itself, read shared state from one place instead of giving every agent private memory, and measure context per agent rather than only in aggregate. Fan-out multiplies both token spend and the chance that two agents hold different versions of the same fact, which is why shared state matters more as agent count rises, not less.
What is context engineering for AI agents?
It is the practice of deciding what an agent sees at inference time, rather than only what you ask it. That covers the system prompt, tool definitions, retrieved passages, conversation history, compaction policy and any memory the agent reads. Prompt engineering optimises the instruction; context engineering optimises everything else in the window, which on a real agent request is the overwhelming majority of the tokens.
Is context engineering just a rebrand of prompt engineering?
No, it is a wider scope. Prompt engineering covers the instructions you write; context engineering covers the whole payload including tool schemas, retrieved documents, conversation history and tool results, most of which you never type and which usually dominate the token count.
What is the single highest-leverage context engineering change?
Split your bill into four buckets first: system prompt, retrieved context, conversation history, and output. It takes about a day and it redirects everything after it, because most teams find one bucket dominates and it is rarely the one they assumed.
Does a 1M token context window make context engineering unnecessary?
It removes the hard ceiling and not the cost. Attention still thins out across long sequences, so filling a large window degrades accuracy even when nothing is rejected, and you pay for every token on every request.
How does context engineering relate to RAG?
Retrieval is one technique inside context engineering, and it is a query-time one. It decides what to fetch per request, which means the payload grows as your corpus grows, because more documents match.
Do I need a memory layer if my context engineering is good?
If your agents answer single-system questions, no. The moment two systems can disagree about the same fact, or the answer depends on when something was true, window management cannot decide it and something has to.
How is context engineering different from prompt engineering?
Prompt engineering optimizes the instructions a model receives, while context engineering governs the whole payload on every request: the system prompt, tool definitions, retrieved documents, conversation history and tool results. Anthropic frames context engineering as the progression of prompt engineering for agents that run over many turns, where the context changes on every step and has to be curated continuously.
Do 1M token context windows still suffer from context rot?
Yes. Chroma's Context Rot study of 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3, found performance became increasingly unreliable as input length grew, and Lost in the Middle (Liu et al., TACL) found accuracy drops when the relevant passage sits in the middle of a long input. A bigger window moves the hard limit without removing the degradation.
How many more tokens do multi-agent systems use?
Anthropic reports that agents use about 4 times more tokens than chat interactions and multi-agent systems about 15 times more, measured on its own multi-agent research system. That system outperformed a single Claude Opus 4 agent by 90.2% on Anthropic's internal research eval, so the extra tokens can pay off on broad research tasks, provided each subagent returns a condensed summary rather than its full history.
Which tool should I start with to manage agent context?
Start with tracing, through LangSmith or Langfuse, to see which part of the window costs the most, then fix the largest bucket in your framework or retrieval layer. If the bigger problem is that the material is stale or contradictory rather than too large, a managed organization memory such as Sentra addresses it before the request is assembled.
The decision rule
Do the six window-management techniques first, in the order above, because they are cheap and several are one-time. Then check whether your agents are re-deriving the same context on every request. If they are, that is a knowledge-layer problem wearing a context-engineering costume, and no amount of compaction, caching or sub-agent routing will close it.