它在编程方面达到了前沿水平,定价为每百万输入 token 0.50 美元、每百万输出 token 2.50 美元,成为智能与成本的全新最优组合。我们还发布了一份关于如何训练该模型的技术报告。
前沿水平的编程智能
我们正在快速提升模型质量。Composer 2 在我们衡量的所有基准测试中均取得了大幅改进,包括 Terminal-Bench 2.01 和 SWE-bench Multilingual:
| 模型 | CursorBench | Terminal-Bench 2.0 | SWE-bench Multilingual |
|---|---|---|---|
| Composer 2 | 61.3 | 61.7 | 73.7 |
| Composer 1.5 | 44.2 | 47.9 | 65.9 |
| Composer 1 | 38.0 | 40.0 | 56.9 |
这些质量提升源于我们首次进行的持续预训练,这为扩展强化学习提供了远更强大的基础。
在此基础之上,我们通过强化学习对长周期编程任务进行训练。Composer 2 能够解决需要数百个动作的挑战性任务。
立即体验 Composer 2
Composer 2 定价为每百万输入 token 0.50 美元、每百万输出 token 2.50 美元。
此外还有一个智能水平相同但速度更快的变体,定价为每百万输入 token 1.50 美元、每百万输出 token 7.50 美元,其成本低于其他快速模型²。我们正将快速版本设为默认选项。完整详情请参阅当前的 Composer 模型文档。
在个人套餐中,Composer 使用量计入第一方模型池,并包含慷慨的免费额度。立即在 Cursor 或我们新界面的早期 alpha 版本中体验 Composer 2。
- Terminal-Bench 2.0 是由 Laude Institute 维护的终端使用智能体评估基准。Anthropic 模型分数使用 Claude Code 测试框架,OpenAI 模型分数使用 Simple Codex 测试框架。我们的 Cursor 分数使用官方 Harbor 评估框架(Terminal-Bench 2.0 的指定测试框架)及默认基准设置计算得出。我们对每个模型-智能体组合运行了 5 次迭代,并报告平均值。有关该基准的更多详情,请访问官方 Terminal Bench 网站。对于 Composer 2 以外的其他模型,我们取官方排行榜分数与在我们基础设施中运行记录分数之间的最高值。↩
- 所有模型的每秒 token 数(TPS)均来自 2026 年 3 月 18 日 Cursor 流量的快照。Composer 与 GPT 模型的 token 大小相近。Anthropic 的 token 大约小 15%,TPS 数值已做归一化处理以反映这一差异。同样,非 Anthropic 模型的输出 token 价格也按约 15% 的相同比例进行了调整。实际速度可能因供应商容量及后续优化而有所变化。↩
Composer 2 技术报告
Sasha Rush
推出 Composer 2.5
It's frontier-level at coding and priced at $0.50/M input and $2.50/M output tokens, making it a new, optimal combination of intelligence and cost. We also released a technical report on how we trained it.
Frontier-level coding intelligence
We're rapidly improving the quality of our model. Composer 2 delivers large improvements on all benchmarks we measure, including Terminal-Bench 2.01 and SWE-bench Multilingual:
| Model | CursorBench | Terminal-Bench 2.0 | SWE-bench Multilingual |
|---|---|---|---|
| Composer 2 | 61.3 | 61.7 | 73.7 |
| Composer 1.5 | 44.2 | 47.9 | 65.9 |
| Composer 1 | 38.0 | 40.0 | 56.9 |
These quality improvements come from our first continued pretraining run, which provides a far stronger base to scale our reinforcement learning.
From this base, we train on long-horizon coding tasks through reinforcement learning. Composer 2 is able to solve challenging tasks requiring hundreds of actions.
Try Composer 2
Composer 2 is priced at $0.50/M input and $2.50/M output tokens.
There is also a faster variant with the same intelligence at $1.50/M input and $7.50/M output tokens, which has a lower cost than other fast models2. We're making fast the default option. See the current Composer model docs for full details.
On individual plans, Composer usage is part of the First-party models pool with generous usage included. Try Composer 2 today in Cursor or in the early alpha of our new interface.
- Terminal-Bench 2.0 is an agent evaluation benchmark for terminal use maintained by the Laude Institute. Anthropic model scores use the Claude Code harness and OpenAI model scores use the Simple Codex harness. Our Cursor score was computed using the official Harbor evaluation framework (the designated harness for Terminal-Bench 2.0) with default benchmark settings. We ran 5 iterations per model-agent pair and report the average. More details on the benchmark can be found at the official Terminal Bench website. For other models besides Composer 2, we took the max score between the official leaderboard score and the score recorded running in our infrastructure. ↩
- Tokens per second (TPS) for all models are from a snapshot of Cursor traffic on March 18th, 2026. Token sizing for Composer and GPT models are similar. Anthropic tokens are ~15% smaller and the TPS number is normalized to reflect that. Similarly, output token price for non-Anthropic models was scaled to match the same ~15% change. Speed may vary depending on provider capacity and improvements over time. ↩
A technical report on Composer 2
Sasha Rush
Introducing Composer 2.5