{"id":943,"date":"2026-10-09T12:00:13","date_gmt":"2026-10-09T12:00:13","guid":{"rendered":"https:\/\/imcodinggenius.com\/?p=943"},"modified":"2026-10-09T12:00:13","modified_gmt":"2026-10-09T12:00:13","slug":"context-engineering-for-ai-agents-4-techniques-to-cut-token-costs","status":"publish","type":"post","link":"https:\/\/imcodinggenius.com\/?p=943","title":{"rendered":"Context Engineering for AI Agents: 4 Techniques to Cut Token Costs"},"content":{"rendered":"<p>There\u2019s a failure mode in agent systems that never shows up as an error. The agent works well for twenty turns, then gets subtly worse \u2014 forgets a constraint from earlier, repeats a tool call it already made, contradicts something it said. No exception, no alert, no obvious cause.<\/p>\n<p>What happened is that context filled up and something had to go. The question is whether you decided what went, or whether a default did.<\/p>\n<h2 class=\"wp-block-heading\">Where the tokens actually go<\/h2>\n<p>Before optimizing, measure. A long-running agent\u2019s context is usually four things:<\/p>\n<p><strong>Tool definitions.<\/strong> Every tool the agent can see costs tokens for its name, description, and JSON schema. A well-documented MCP tool runs 150\u2013400 tokens. Connect twelve MCP servers averaging twenty tools each and you\u2019re at 240 tools \u2014 call it 30,000 tokens of schema, resident in every single request, before the user has said a word.<\/p>\n<p>This is the most common and most invisible cost. It scales with what you <em>connected<\/em>, not with what the agent <em>uses<\/em>.<\/p>\n<p><strong>Conversation history.<\/strong> Grows linearly with turns. Everyone expects this one.<\/p>\n<p><strong>Tool results.<\/strong> The variance killer. Most calls return a few hundred tokens. Then one query returns 400KB of JSON, and unless something intervenes, that payload sits in context for the rest of the run \u2014 re-sent on every subsequent turn.<\/p>\n<p><strong>System prompt and instructions.<\/strong> Fixed and usually the smallest piece, which is ironic given how much time teams spend editing it.<\/p>\n<p>The distribution matters strategically: two of the four are <em>structural<\/em> (tool definitions, tool results) and can be fixed once, architecturally. The other two are linear and can only be managed. Fix the structural ones first \u2014 they\u2019re where the large, cheap wins are.<\/p>\n<h2 class=\"wp-block-heading\">Four techniques<\/h2>\n<h3 class=\"wp-block-heading\">Deferred tool loading<\/h3>\n<p>Don\u2019t put all 240 tool schemas in the prompt. Give the model a compact index \u2014 server names and one-line summaries \u2014 and load full definitions only for the server it decides to use.<\/p>\n<p>Typical saving is 60\u201380% of tool-definition tokens on runs that touch two or three servers, which is most runs.<\/p>\n<p>The cost: an extra round trip when the agent picks a new server, and slightly worse tool discovery. The model can\u2019t reason about a tool whose schema it hasn\u2019t seen, so if your agent genuinely needs to compare capabilities across many servers, the index has to be good enough to route on. Write those one-liners carefully.<\/p>\n<h3 class=\"wp-block-heading\">Large tool response offloading<\/h3>\n<p>Set a threshold \u2014 5,000 tokens is a reasonable default. Under it, results go into context normally. Over it, write the payload to the sandbox filesystem and put a summary plus a path into context instead. The agent can then grep, slice, or parse the file with code if it needs detail.<\/p>\n<p>This converts a catastrophic cost into a small one, and it fits how the data is usually used. An agent that fetched 10,000 rows almost never needs all 10,000 in its reasoning; it needs an aggregate or a handful of matches.<\/p>\n<p>The cost: the agent needs sandbox access, and it needs to be prompted to actually query the file rather than guessing from the summary. Without that nudge it will sometimes answer from the summary alone \u2014 confidently and wrongly.<\/p>\n<h3 class=\"wp-block-heading\">Code mode<\/h3>\n<p>Instead of the model making five sequential tool calls and receiving five results into context, let it write one script that makes all five calls, joins the data, and returns only the final answer.<\/p>\n<p>The intermediate payloads never enter context at all. For genuinely multi-step data work \u2014 cross-referencing two systems, aggregating across pages \u2014 this is the difference between a run that fits and one that doesn\u2019t.<\/p>\n<p>The cost: harder to debug, since the reasoning is inside a script rather than visible as discrete steps. It also requires the agent to be a competent enough programmer for the task, which is model-dependent. Use it for data-shaping work, not for decisions you\u2019ll need to audit step by step.<\/p>\n<h3 class=\"wp-block-heading\">Sub-agents<\/h3>\n<p>Delegate a bounded subtask to a fresh agent with its own clean context. It burns thirty turns exploring, and returns one paragraph. The parent\u2019s context grows by that paragraph.<\/p>\n<p>This is the most powerful technique available and the one people reach for last. It\u2019s also the one that composes with everything else: a sub-agent can use deferred loading, offloading, and code mode inside its own run.<\/p>\n<p>The cost: total token spend usually goes <em>up<\/em> even as the parent\u2019s context stays small, because you\u2019re paying for the sub-agent\u2019s full transcript. You\u2019re trading tokens for quality and run length. That\u2019s normally the right trade, but watch the bill, and watch sub-agent count for delegation loops.<\/p>\n<h3 class=\"wp-block-heading\">And compaction, as a floor<\/h3>\n<p>When context approaches the limit despite all of the above, summarize older turns rather than dropping them. Truncation discards the decision that explains current state; summarization keeps a lossy version of it.<\/p>\n<p>Compaction is a safety net, not a strategy. If it\u2019s firing regularly, one of the four techniques above isn\u2019t doing its job. Treat every compaction event as a signal worth investigating rather than a feature working as intended.<\/p>\n<h2 class=\"wp-block-heading\">The routing problem underneath<\/h2>\n<p>All of this assumes you can see what tools exist before deciding what to load \u2014 which means tool definitions need to come from somewhere queryable, not from files committed next to each agent.<\/p>\n<p>This is where MCP\u2019s design pays off. If your servers are registered centrally, an agent can enumerate available capability cheaply, pull schemas on demand, and pick tools at runtime. TrueFoundry\u2019s <a href=\"https:\/\/www.truefoundry.com\/docs\/ai-gateway\/mcp\/mcp-overview\" target=\"_blank\" rel=\"noopener\">MCP Gateway<\/a> is one implementation of that registry pattern \u2014 servers register once, authentication and per-user OAuth stay at the gateway, and agents select individual tools from the catalogue rather than embedding endpoints and credentials. The side benefit is that tools flagged destructive once, centrally, inherit an approval gate in every agent that uses them.<\/p>\n<p>The related point is that these techniques are runtime behavior, not application logic, which means they belong in the harness rather than in your agent. TrueFoundry\u2019s <a href=\"https:\/\/www.truefoundry.com\/docs\/agent-platform\/agent-harness\/overview\" target=\"_blank\" rel=\"noopener\">agent harness documentation<\/a> enumerates them as configuration toggles \u2014 deferred tool loading, large tool responses, code mode, subagents, compaction \u2014 which is a useful checklist regardless of what you run on. If you\u2019re implementing these yourself, that list is roughly the scope.<\/p>\n<h2 class=\"wp-block-heading\">Instrument before you optimize<\/h2>\n<p>Track two things per turn: <strong>active context size<\/strong> and <strong>context composition<\/strong> \u2014 how much is tool definitions versus history versus results.<\/p>\n<p>Size tells you how close to the wall you are and when compaction will fire. Composition tells you which technique to apply. An agent at 80% context that\u2019s mostly tool definitions needs deferred loading. One that\u2019s mostly a single enormous tool result needs offloading. One that\u2019s mostly history needs sub-agents. Same symptom, three different fixes, and no way to choose without the breakdown.<\/p>\n<p>Most teams have neither number. Getting the first one is usually an afternoon, and it will immediately explain at least one bug you\u2019ve been carrying for a month.<\/p>\n<h2 class=\"wp-block-heading\">Order of operations<\/h2>\n<p><strong>Measure composition.<\/strong> One afternoon. Do this first; it determines everything after.<\/p>\n<p><strong>Deferred tool loading.<\/strong> Biggest ratio of saving to effort if you have many MCP servers.<\/p>\n<p><strong>Large response offloading.<\/strong> Kills the worst tail case.<\/p>\n<p><strong>Sub-agents<\/strong> for anything with a bounded, delegable subtask.<\/p>\n<p><strong>Code mode<\/strong> for multi-step data work specifically.<\/p>\n<p><strong>Compaction<\/strong> as a floor, with alerting when it fires.<\/p>\n<p>Skip straight to step 4 and you\u2019ll get a real improvement and still be paying 30,000 tokens a request for tools nobody calls.<\/p>\n<p>The post <a href=\"https:\/\/www.thecrazyprogrammer.com\/2026\/10\/context-engineering-for-ai-agents-4-techniques-to-cut-token-costs.html\">Context Engineering for AI Agents: 4 Techniques to Cut Token Costs<\/a> appeared first on <a href=\"https:\/\/www.thecrazyprogrammer.com\/\">The Crazy Programmer<\/a>.<\/p>","protected":false},"excerpt":{"rendered":"<p>There\u2019s a failure mode in agent systems that never shows up as an error. The agent works well for twenty turns, then gets subtly worse \u2014 forgets a constraint from earlier, repeats a tool call it already made, contradicts something it said. No exception, no alert, no obvious cause. What &#8230; <\/p>\n<div><a class=\"more-link bs-book_btn\" href=\"https:\/\/imcodinggenius.com\/?p=943\">Read More<\/a><\/div>\n","protected":false},"author":0,"featured_media":944,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[],"class_list":["post-943","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-development"],"_links":{"self":[{"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=\/wp\/v2\/posts\/943"}],"collection":[{"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=943"}],"version-history":[{"count":0,"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=\/wp\/v2\/posts\/943\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=\/wp\/v2\/media\/944"}],"wp:attachment":[{"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=943"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=943"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/imcodinggenius.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=943"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}