AI glossary

Context Engineering

Context engineering is the practice of deciding what information goes into a large language model’s context window at each step of a task: instructions, examples, retrieved documents, tool definitions, conversation history and the model’s own notes. Where prompt engineering focuses on how to word an instruction, context engineering manages the whole input the model sees, and how that input changes over a long, multi-step job.

Anthropic’s engineering team defines it as “the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference, including all the other information that may land there outside of the prompts.”

Where the term came from

The term spread in June 2025. On 19 June, Shopify CEO Tobi Lütke wrote on X that he preferred “context engineering” over prompt engineering because “it describes the core skill better: the art of providing all the context for the task to be plausibly solvable by the LLM.” Lütke did not claim to have coined it.

Six days later, Andrej Karpathy, a founding member of OpenAI and former director of AI at Tesla, added “+1 for ‘context engineering’ over ‘prompt engineering’”. His framing became the most quoted definition: in every industrial-strength LLM app, context engineering is “the delicate art and science of filling the context window with just the right information for the next step.” He also named the trade-off: with too little or badly formatted context the model underperforms, while with too much or irrelevant context “the LLM costs might go up and performance might come down.”

Within months, LangChain (July 2025), Anthropic (September 2025), OpenAI’s cookbook and Google’s Agent Development Kit team (December 2025) had all published guidance under the same name.

Context engineering vs prompt engineering

Prompt engineering Context engineering
Unit of work A single instruction or prompt Everything in the context window, over many steps
Main question How do I phrase this? What should the model see right now, and what should it not?
Typical setting One-off chat or a fixed template AI agents and multi-step LLM applications
Inputs managed System prompt, examples System prompt, tools, retrieved data, history, memory, outputs of other agents

The two are not rivals. A well-written prompt is still part of the context; context engineering is the larger discipline that decides what surrounds it. Karpathy’s point was that “prompt” makes people think of a short task description, while real applications assemble far more than that before every model call.

Why it matters: long context is not free

Modern models accept context windows of a million tokens or more, which might suggest you can simply include everything. Measured evidence says otherwise:

  • Lost in the middle. Liu et al. (“Lost in the Middle”, arXiv July 2023, published in TACL 2024) found that performance “is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades” when the model must use information in the middle of a long context.
  • Context rot. Chroma’s July 2025 report tested 18 models from Anthropic, OpenAI, Google and Alibaba and concluded that “models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.” Even a single distracting passage reduced accuracy, and models did much better on focused prompts than on the same task buried in a full conversation history.

Anthropic describes the same constraint as an “attention budget”: every token added competes for the model’s limited attention, so more context can mean worse answers, not better ones. Every token also costs money and latency, which is why context design connects directly to prompt caching and API cost control.

Core techniques

LangChain groups context engineering into four strategies, and most published techniques fit one of them:

  • Write: save information outside the context window so the agent can use it later, such as a scratchpad or long-term memory. Anthropic calls one form of this “structured note-taking”: the agent regularly writes notes to memory outside the context window.
  • Select: pull only the relevant information in. This includes retrieval-augmented generation, choosing which tools to expose, and what Anthropic calls “just-in-time” retrieval, where the agent keeps lightweight references (file paths, queries, links) and loads the data only when needed.
  • Compress: keep only the tokens the task requires. The main method is compaction, which Anthropic describes as summarizing a conversation that is nearing the context limit and restarting with the summary. OpenAI’s Agents SDK cookbook offers trimming (keep the last N turns) and summarization as the two basic options.
  • Isolate: split the work so no single context has to hold everything. In sub-agent architectures, specialized agents work in their own clean context and return condensed results to a coordinating agent.

Standards such as the Model Context Protocol (MCP) are part of the same picture: they decide how tools and external data are described to the model, and every tool definition takes up space in the context.

Common failure modes

  • Overloading: pasting entire documents or full chat histories “just in case”, which triggers the degradation measured in the context-rot research.
  • Context poisoning: an error or hallucination gets written into the history or a summary and is then repeated as if it were fact in later steps. OpenAI’s cookbook names this as a risk of summarization.
  • Lost instructions: key constraints placed in the middle of a long input, where models are least likely to use them.
  • Tool sprawl: exposing dozens of overlapping tools, so the model wastes context and picks the wrong one.

FAQ

What is context engineering in simple terms?

It is deciding what an AI model gets to see before it answers: which instructions, documents, tools and past messages go into its context window, and which are left out or summarized.

Who coined the term context engineering?

It was popularized in June 2025 by Shopify CEO Tobi Lütke and Andrej Karpathy, both writing on X. Lütke said he liked the term rather than claiming to have invented it.

Is context engineering replacing prompt engineering?

Not replacing, but containing it. The prompt is one part of the context. Context engineering adds everything else the model sees, which matters most for agents that run over many steps.

Why not just use a model with a huge context window?

Because models use long context unevenly. Research shows accuracy falls as inputs grow and when key information sits in the middle, and every extra token adds cost and latency.

What are the main context engineering techniques?

Writing information to external memory, selecting only relevant information (including RAG), compressing history through summarization or compaction, and isolating work in sub-agents with their own context.

For a practical build, see How to Build an AI Agent, and for the cost side, Prompt Caching Explained.