直接工具调用会为每个定义和结果消耗上下文。智能体通过编写代码来调用工具,能够更好地扩展。以下是其在 MCP 中的运作方式。
模型上下文协议(MCP)是一个用于将 AI 智能体连接到外部系统的开放标准。传统上,将智能体连接到工具和数据需要为每一对组合进行定制集成,这造成了碎片化和重复劳动,使得构建真正互联的系统难以扩展。MCP 提供了一种通用协议——开发者只需在其智能体中实现一次 MCP,即可解锁整个集成生态系统。
自 2024 年 11 月推出 MCP 以来,其采用速度非常快:社区已构建了数千个 MCP 服务器,所有主流编程语言都有可用的 SDK,并且行业已将 MCP 作为连接智能体与工具和数据的事实标准。
如今,开发者通常会构建能够访问数十个 MCP 服务器上成百上千个工具的智能体。然而,随着连接工具数量的增长,预先加载所有工具定义并通过上下文窗口传递中间结果,会拖慢智能体的速度并增加成本。
在这篇博客中,我们将探讨代码执行如何使智能体更高效地与 MCP 服务器交互,从而在消耗更少模型 token 的同时处理更多工具。
工具消耗过多模型 token 会降低智能体效率
随着 MCP 使用规模的扩大,有两种常见模式会增加智能体的成本和延迟:
- 工具定义使上下文窗口过载;
- 中间工具结果消耗额外的模型 token。
1. 工具定义使上下文窗口过载
大多数 MCP 客户端会预先将所有工具定义直接加载到上下文中,并使用直接工具调用语法将其暴露给模型。这些工具定义可能如下所示:
gdrive.getDocument
Description: Retrieves a document from Google Drive
Parameters:
documentId (required, string): The ID of the document to retrieve
fields (optional, string): Specific fields to return
Returns: Document object with title, body content, metadata, permissions, etc. salesforce.updateRecord
Description: Updates a record in Salesforce
Parameters:
objectType (required, string): Type of Salesforce object (Lead, Contact, Account, etc.)
recordId (required, string): The ID of the record to update
data (required, object): Fields to update with their new values
Returns: Updated record object with confirmation 工具描述会占用更多上下文窗口空间,从而增加响应时间和成本。当智能体连接到数千个工具时,它们需要在读取请求之前处理数十万个模型 token。
2. 中间工具结果消耗额外的模型 token
大多数 MCP 客户端允许模型直接调用 MCP 工具。例如,你可以对智能体说:"从 Google Drive 下载我的会议记录,并将其附加到 Salesforce 的潜在客户中。"
模型会发出类似这样的调用:
TOOL CALL: gdrive.getDocument(documentId: "abc123")
→ returns "Discussed Q4 goals...\n[full transcript text]"
(loaded into model context)
TOOL CALL: salesforce.updateRecord(
objectType: "SalesMeeting",
recordId: "00Q5f000001abcXYZ",
data: { "Notes": "Discussed Q4 goals...\n[full transcript text written out]" }
)
(model needs to write entire transcript into context again) 每一个中间结果都必须经过模型处理。在这个例子中,完整的通话记录会经过模型两次。对于一个两小时的销售会议,这可能意味着要额外处理 50,000 个 token。更大的文档甚至可能超出上下文窗口的限制,导致工作流程中断。
在处理大型文档或复杂数据结构时,模型在工具调用之间复制数据时更容易出错。

使用 MCP 执行代码可提高上下文效率
随着代码执行环境在智能体中变得越来越普遍,一种解决方案是将 MCP 服务器作为代码 API 呈现,而不是直接进行工具调用。智能体随后可以编写代码来与 MCP 服务器交互。这种方法解决了两个挑战:智能体可以只加载它们需要的工具,并在执行环境中处理数据,然后再将结果传回给模型。
有多种方法可以实现这一点。一种方法是生成一个文件树,包含来自已连接 MCP 服务器的所有可用工具。以下是使用 TypeScript 的实现:
servers
├── google-drive
│ ├── getDocument.ts
│ ├── ... (other tools)
│ └── index.ts
├── salesforce
│ ├── updateRecord.ts
│ ├── ... (other tools)
│ └── index.ts
└── ... (other servers) 然后每个工具对应一个文件,类似于:
// ./servers/google-drive/getDocument.ts
import { callMCPTool } from "../../../client.js";
interface GetDocumentInput {
documentId: string;
}
interface GetDocumentResponse {
content: string;
}
/* Read a document from Google Drive */
export async function getDocument(input: GetDocumentInput): Promise<GetDocumentResponse> {
return callMCPTool<GetDocumentResponse>('google_drive__get_document', input);
}
我们上面提到的 Google Drive 到 Salesforce 的例子变成了如下代码:
// Read transcript from Google Docs and add to Salesforce prospect
import * as gdrive from './servers/google-drive';
import * as salesforce from './servers/salesforce';
const transcript = (await gdrive.getDocument({ documentId: 'abc123' })).content;
await salesforce.updateRecord({
objectType: 'SalesMeeting',
recordId: '00Q5f000001abcXYZ',
data: { Notes: transcript }
});
智能体通过探索文件系统来发现工具:列出 `./servers/` 目录以查找可用的服务器(如 google-drive 和 salesforce),然后读取它需要的特定工具文件(如 `getDocument.ts` 和 `updateRecord.ts`)以了解每个工具的接口。这使得智能体只需加载当前任务所需的定义。这将 token 使用量从 150,000 个 token 减少到 2,000 个 token——节省了 98.7% 的时间和成本。
Cloudflare 也发布了类似的研究成果,将使用 MCP 执行代码称为“代码模式”。其核心见解是一致的:大语言模型擅长编写代码,开发者应善用这一优势,构建能够更高效地与 MCP 服务器交互的智能体。
使用 MCP 执行代码的优势
使用 MCP 执行代码,智能体能够按需加载工具、在数据到达模型前进行过滤,以及通过单一步骤执行复杂逻辑,从而更高效地利用上下文。此外,采用这种方法还能带来安全性和状态管理方面的好处。
渐进式披露
模型非常擅长浏览文件系统。将工具以代码形式呈现在文件系统中,可以让模型按需读取工具定义,而无需一次性读取所有内容。
或者,也可以在服务器中添加一个 search_tools 工具来查找相关定义。例如,在使用上述假设的 Salesforce 服务器时,智能体会搜索“salesforce”,并仅加载当前任务所需的那些工具。在 search_tools 工具中加入一个 detail level 参数,允许智能体选择所需的详细程度(例如仅名称、名称与描述,或包含模式的完整定义),这也有助于智能体节省上下文并高效地查找工具。
上下文高效的工具结果
处理大型数据集时,智能体可以在返回结果之前,在代码中对数据进行过滤和转换。考虑获取一个包含 10,000 行的电子表格:
// Without code execution - all rows flow through context
TOOL CALL: gdrive.getSheet(sheetId: 'abc123')
→ returns 10,000 rows in context to filter manually
// With code execution - filter in the execution environment
const allRows = await gdrive.getSheet({ sheetId: 'abc123' });
const pendingOrders = allRows.filter(row =>
row["Status"] === 'pending'
);
console.log(`Found ${pendingOrders.length} pending orders`);
console.log(pendingOrders.slice(0, 5)); // Only log first 5 for review 智能体看到的是 5 行,而不是 10,000 行。类似的模式也适用于聚合、跨多个数据源的连接,或提取特定字段——所有这些都不会使上下文窗口膨胀。
更强大且上下文高效的控制流
循环、条件判断和错误处理可以使用熟悉的代码模式来完成,而不是通过链式调用单个工具。例如,如果你需要在 Slack 中发送一个部署通知,智能体可以编写:
let found = false;
while (!found) {
const messages = await slack.getChannelHistory({ channel: 'C123456' });
found = messages.some(m => m.text.includes('deployment complete'));
if (!found) await new Promise(r => setTimeout(r, 5000));
}
console.log('Deployment notification received'); 这种方法比在智能体循环中交替进行 MCP 工具调用和 sleep 命令要高效得多。
此外,能够编写出可执行的条件分支树还能节省“首 token 生成时间”:智能体无需等待模型评估 if 语句,而是让代码执行环境来完成这一操作。
隐私保护操作
当智能体通过 MCP 执行代码时,中间结果默认保留在执行环境中。这样一来,智能体只能看到你明确记录或返回的内容,这意味着你不希望与模型共享的数据可以在工作流中流转,而无需进入模型的上下文窗口。
对于更敏感的工作负载,智能体框架可以自动对敏感数据进行 token 化处理。例如,假设你需要将电子表格中的客户联系信息导入 Salesforce。智能体会编写如下代码:
const sheet = await gdrive.getSheet({ sheetId: 'abc123' });
for (const row of sheet.rows) {
await salesforce.updateRecord({
objectType: 'Lead',
recordId: row.salesforceId,
data: {
Email: row.email,
Phone: row.phone,
Name: row.name
}
});
}
console.log(`Updated ${sheet.rows.length} leads`); MCP 客户端会拦截数据,并在数据到达模型之前对个人身份信息进行 token 化处理:
// What the agent would see, if it logged the sheet.rows:
[
{ salesforceId: '00Q...', email: '[EMAIL_1]', phone: '[PHONE_1]', name: '[NAME_1]' },
{ salesforceId: '00Q...', email: '[EMAIL_2]', phone: '[PHONE_2]', name: '[NAME_2]' },
...
] 然后,当数据在另一个 MCP 工具调用中被共享时,会通过 MCP 客户端中的查找表进行反 token 化处理。真实的电子邮件地址、电话号码和姓名会从 Google Sheets 流向 Salesforce,但绝不会经过模型。这可以防止智能体意外记录或处理敏感数据。你还可以利用此功能定义确定性的安全规则,选择数据可以流向何处以及来自何处。
状态持久化与技能
具备文件系统访问权限的代码执行允许智能体在多次操作之间维持状态。智能体可以将中间结果写入文件,从而能够恢复工作并跟踪进度:
const leads = await salesforce.query({
query: 'SELECT Id, Email FROM Lead LIMIT 1000'
});
const csvData = leads.map(l => `${l.Id},${l.Email}`).join('\n');
await fs.writeFile('./workspace/leads.csv', csvData);
// Later execution picks up where it left off
const saved = await fs.readFile('./workspace/leads.csv', 'utf-8'); 智能体还可以将自己的代码持久化为可复用的函数。一旦智能体为某个任务开发出可运行的代码,它就可以保存该实现以供将来使用:
// In ./skills/save-sheet-as-csv.ts
import * as gdrive from './servers/google-drive';
export async function saveSheetAsCsv(sheetId: string) {
const data = await gdrive.getSheet({ sheetId });
const csv = data.map(row => row.join(',')).join('\n');
await fs.writeFile(`./workspace/sheet-${sheetId}.csv`, csv);
return `./workspace/sheet-${sheetId}.csv`;
}
// Later, in any agent execution:
import { saveSheetAsCsv } from './skills/save-sheet-as-csv';
const csvPath = await saveSheetAsCsv('abc123'); 这与“技能”的概念紧密相关——技能是包含可复用指令、脚本和资源的文件夹,用于提升模型在特定任务上的表现。在这些已保存的函数中添加一个 SKILL.md 文件,就能创建一个结构化的技能,供模型引用和使用。随着时间的推移,这能让你的智能体构建一个更高级能力的工具箱,不断进化出高效工作所需的脚手架。
需要注意的是,代码执行本身会引入其自身的复杂性。运行智能体生成的代码需要一个安全的执行环境,配备适当的沙箱机制、资源限制和监控措施。这些基础设施要求增加了运维开销和安全考量,而直接的工具调用则可以避免这些问题。代码执行的优势——降低 token 成本、减少延迟以及改进工具组合能力——需要与这些实现成本进行权衡。
总结
MCP 为智能体连接众多工具和系统提供了一个基础协议。然而,一旦连接了过多的服务器,工具定义和结果可能会消耗过多的 token,从而降低智能体的效率。
尽管这里的许多问题看似新颖——上下文管理、工具组合、状态持久化——但它们在软件工程中已有成熟的解决方案。代码执行将这些既定模式应用于智能体,使其能够利用熟悉的编程结构更高效地与 MCP 服务器进行交互。如果您实施这种方法,我们鼓励您与 MCP 社区分享您的发现。
致谢
本文由 Adam Jones 和 Conor Kelly 撰写。感谢 Jeremy Fox、Jerome Swannack、Stuart Ritchie、Molly Vorwerck、Matt Samuels 和 Maggie Vo 对本文草稿提供的反馈意见。
Direct tool calls consume context for each definition and result. Agents scale better by writing code to call tools instead. Here's how it works with MCP.
The Model Context Protocol (MCP) is an open standard for connecting AI agents to external systems. Connecting agents to tools and data traditionally requires a custom integration for each pairing, creating fragmentation and duplicated effort that makes it difficult to scale truly connected systems. MCP provides a universal protocol—developers implement MCP once in their agent and it unlocks an entire ecosystem of integrations.
Since launching MCP in November 2024, adoption has been rapid: the community has built thousands of MCP servers, SDKs are available for all major programming languages, and the industry has adopted MCP as the de-facto standard for connecting agents to tools and data.
Today developers routinely build agents with access to hundreds or thousands of tools across dozens of MCP servers. However, as the number of connected tools grows, loading all tool definitions upfront and passing intermediate results through the context window slows down agents and increases costs.
In this blog we'll explore how code execution can enable agents to interact with MCP servers more efficiently, handling more tools while using fewer tokens.
Excessive token consumption from tools makes agents less efficient
As MCP usage scales, there are two common patterns that can increase agent cost and latency:
- Tool definitions overload the context window;
- Intermediate tool results consume additional tokens.
1. Tool definitions overload the context window
Most MCP clients load all tool definitions upfront directly into context, exposing them to the model using a direct tool-calling syntax. These tool definitions might look like:
gdrive.getDocument
Description: Retrieves a document from Google Drive
Parameters:
documentId (required, string): The ID of the document to retrieve
fields (optional, string): Specific fields to return
Returns: Document object with title, body content, metadata, permissions, etc. salesforce.updateRecord
Description: Updates a record in Salesforce
Parameters:
objectType (required, string): Type of Salesforce object (Lead, Contact, Account, etc.)
recordId (required, string): The ID of the record to update
data (required, object): Fields to update with their new values
Returns: Updated record object with confirmation Tool descriptions occupy more context window space, increasing response time and costs. In cases where agents are connected to thousands of tools, they’ll need to process hundreds of thousands of tokens before reading a request.
2. Intermediate tool results consume additional tokens
Most MCP clients allow models to directly call MCP tools. For example, you might ask your agent: "Download my meeting transcript from Google Drive and attach it to the Salesforce lead."
The model will make calls like:
TOOL CALL: gdrive.getDocument(documentId: "abc123")
→ returns "Discussed Q4 goals...\n[full transcript text]"
(loaded into model context)
TOOL CALL: salesforce.updateRecord(
objectType: "SalesMeeting",
recordId: "00Q5f000001abcXYZ",
data: { "Notes": "Discussed Q4 goals...\n[full transcript text written out]" }
)
(model needs to write entire transcript into context again) Every intermediate result must pass through the model. In this example, the full call transcript flows through twice. For a 2-hour sales meeting, that could mean processing an additional 50,000 tokens. Even larger documents may exceed context window limits, breaking the workflow.
With large documents or complex data structures, models may be more likely to make mistakes when copying data between tool calls.

Code execution with MCP improves context efficiency
With code execution environments becoming more common for agents, a solution is to present MCP servers as code APIs rather than direct tool calls. The agent can then write code to interact with MCP servers. This approach addresses both challenges: agents can load only the tools they need and process data in the execution environment before passing results back to the model.
There are a number of ways to do this. One approach is to generate a file tree of all available tools from connected MCP servers. Here's an implementation using TypeScript:
servers
├── google-drive
│ ├── getDocument.ts
│ ├── ... (other tools)
│ └── index.ts
├── salesforce
│ ├── updateRecord.ts
│ ├── ... (other tools)
│ └── index.ts
└── ... (other servers) Then each tool corresponds to a file, something like:
// ./servers/google-drive/getDocument.ts
import { callMCPTool } from "../../../client.js";
interface GetDocumentInput {
documentId: string;
}
interface GetDocumentResponse {
content: string;
}
/* Read a document from Google Drive */
export async function getDocument(input: GetDocumentInput): Promise<GetDocumentResponse> {
return callMCPTool<GetDocumentResponse>('google_drive__get_document', input);
}
Our Google Drive to Salesforce example above becomes the code:
// Read transcript from Google Docs and add to Salesforce prospect
import * as gdrive from './servers/google-drive';
import * as salesforce from './servers/salesforce';
const transcript = (await gdrive.getDocument({ documentId: 'abc123' })).content;
await salesforce.updateRecord({
objectType: 'SalesMeeting',
recordId: '00Q5f000001abcXYZ',
data: { Notes: transcript }
});
The agent discovers tools by exploring the filesystem: listing the ./servers/ directory to find available servers (like google-drive and salesforce), then reading the specific tool files it needs (like getDocument.ts and updateRecord.ts) to understand each tool's interface. This lets the agent load only the definitions it needs for the current task. This reduces the token usage from 150,000 tokens to 2,000 tokens—a time and cost saving of 98.7%.
Cloudflare published similar findings, referring to code execution with MCP as “Code Mode." The core insight is the same: LLMs are adept at writing code and developers should take advantage of this strength to build agents that interact with MCP servers more efficiently.
Benefits of code execution with MCP
Code execution with MCP enables agents to use context more efficiently by loading tools on demand, filtering data before it reaches the model, and executing complex logic in a single step. There are also security and state management benefits to using this approach.
Progressive disclosure
Models are great at navigating filesystems. Presenting tools as code on a filesystem allows models to read tool definitions on-demand, rather than reading them all up-front.
Alternatively, a search_tools tool can be added to the server to find relevant definitions. For example, when working with the hypothetical Salesforce server used above, the agent searches for "salesforce" and loads only those tools that it needs for the current task. Including a detail level parameter in the search_tools tool that allows the agent to select the level of detail required (such as name only, name and description, or the full definition with schemas) also helps the agent conserve context and find tools efficiently.
Context efficient tool results
When working with large datasets, agents can filter and transform results in code before returning them. Consider fetching a 10,000-row spreadsheet:
// Without code execution - all rows flow through context
TOOL CALL: gdrive.getSheet(sheetId: 'abc123')
→ returns 10,000 rows in context to filter manually
// With code execution - filter in the execution environment
const allRows = await gdrive.getSheet({ sheetId: 'abc123' });
const pendingOrders = allRows.filter(row =>
row["Status"] === 'pending'
);
console.log(`Found ${pendingOrders.length} pending orders`);
console.log(pendingOrders.slice(0, 5)); // Only log first 5 for review The agent sees five rows instead of 10,000. Similar patterns work for aggregations, joins across multiple data sources, or extracting specific fields—all without bloating the context window.
More powerful and context-efficient control flow
Loops, conditionals, and error handling can be done with familiar code patterns rather than chaining individual tool calls. For example, if you need a deployment notification in Slack, the agent can write:
let found = false;
while (!found) {
const messages = await slack.getChannelHistory({ channel: 'C123456' });
found = messages.some(m => m.text.includes('deployment complete'));
if (!found) await new Promise(r => setTimeout(r, 5000));
}
console.log('Deployment notification received'); This approach is more efficient than alternating between MCP tool calls and sleep commands through the agent loop.
Additionally, being able to write out a conditional tree that gets executed also saves on “time to first token” latency: rather than having to wait for a model to evaluate an if-statement, the agent can let the code execution environment do this.
Privacy-preserving operations
When agents use code execution with MCP, intermediate results stay in the execution environment by default. This way, the agent only sees what you explicitly log or return, meaning data you don’t wish to share with the model can flow through your workflow without ever entering the model's context.
For even more sensitive workloads, the agent harness can tokenize sensitive data automatically. For example, imagine you need to import customer contact details from a spreadsheet into Salesforce. The agent writes:
const sheet = await gdrive.getSheet({ sheetId: 'abc123' });
for (const row of sheet.rows) {
await salesforce.updateRecord({
objectType: 'Lead',
recordId: row.salesforceId,
data: {
Email: row.email,
Phone: row.phone,
Name: row.name
}
});
}
console.log(`Updated ${sheet.rows.length} leads`); The MCP client intercepts the data and tokenizes PII before it reaches the model:
// What the agent would see, if it logged the sheet.rows:
[
{ salesforceId: '00Q...', email: '[EMAIL_1]', phone: '[PHONE_1]', name: '[NAME_1]' },
{ salesforceId: '00Q...', email: '[EMAIL_2]', phone: '[PHONE_2]', name: '[NAME_2]' },
...
] Then, when the data is shared in another MCP tool call, it is untokenized via a lookup in the MCP client. The real email addresses, phone numbers, and names flow from Google Sheets to Salesforce, but never through the model. This prevents the agent from accidentally logging or processing sensitive data. You can also use this to define deterministic security rules, choosing where data can flow to and from.
State persistence and skills
Code execution with filesystem access allows agents to maintain state across operations. Agents can write intermediate results to files, enabling them to resume work and track progress:
const leads = await salesforce.query({
query: 'SELECT Id, Email FROM Lead LIMIT 1000'
});
const csvData = leads.map(l => `${l.Id},${l.Email}`).join('\n');
await fs.writeFile('./workspace/leads.csv', csvData);
// Later execution picks up where it left off
const saved = await fs.readFile('./workspace/leads.csv', 'utf-8'); Agents can also persist their own code as reusable functions. Once an agent develops working code for a task, it can save that implementation for future use:
// In ./skills/save-sheet-as-csv.ts
import * as gdrive from './servers/google-drive';
export async function saveSheetAsCsv(sheetId: string) {
const data = await gdrive.getSheet({ sheetId });
const csv = data.map(row => row.join(',')).join('\n');
await fs.writeFile(`./workspace/sheet-${sheetId}.csv`, csv);
return `./workspace/sheet-${sheetId}.csv`;
}
// Later, in any agent execution:
import { saveSheetAsCsv } from './skills/save-sheet-as-csv';
const csvPath = await saveSheetAsCsv('abc123'); This ties in closely to the concept of Skills, folders of reusable instructions, scripts, and resources for models to improve performance on specialized tasks. Adding a SKILL.md file to these saved functions creates a structured skill that models can reference and use. Over time, this allows your agent to build a toolbox of higher-level capabilities, evolving the scaffolding that it needs to work most effectively.
Note that code execution introduces its own complexity. Running agent-generated code requires a secure execution environment with appropriate sandboxing, resource limits, and monitoring. These infrastructure requirements add operational overhead and security considerations that direct tool calls avoid. The benefits of code execution—reduced token costs, lower latency, and improved tool composition—should be weighed against these implementation costs.
Summary
MCP provides a foundational protocol for agents to connect to many tools and systems. However, once too many servers are connected, tool definitions and results can consume excessive tokens, reducing agent efficiency.
Although many of the problems here feel novel—context management, tool composition, state persistence—they have known solutions from software engineering. Code execution applies these established patterns to agents, letting them use familiar programming constructs to interact with MCP servers more efficiently. If you implement this approach, we encourage you to share your findings with the MCP community.
Acknowledgments
This article was written by Adam Jones and Conor Kelly. Thanks to Jeremy Fox, Jerome Swannack, Stuart Ritchie, Molly Vorwerck, Matt Samuels, and Maggie Vo for feedback on drafts of this post.