推出 Grok 4.5 真实世界的工程卓越性 训练 Grok 4.5 比 Flash 模型更快 擅长办公任务 定价 快速上手
今天,我们正式发布 Grok 4.5,这是 SpaceXAI 最智能的模型,专为在编程、智能体任务和知识工作领域表现出色而打造。它是我们有史以来最强的模型,并且是与 Cursor 共同训练的。
真实世界的工程卓越性
Grok 4.5 在涵盖编程、科学、工程和数学知识的数据集上进行了训练。凭借智能且高效的推理能力,Grok 4.5 在真实工程任务中表现出色,并在这些任务上超越了同类领先模型。
DeepSWE 1.0 DeepSWE 1.1 SWE Marathon Terminal Bench 2.1 SWE Bench Pro
竞品数据均来自各开发者公开发布的系统卡或基准排行榜
模型得分的基准柱状图。DeepSWE 1.0(由 Datacurve 创建,AA 使用各模型提供商的测试框架运行):Fable (max) 66.1%,GPT 5.5 (xhigh) 64.31%,Grok 4.5 62.0%,Opus 4.8 (max) 55.75%,Opus 4.7 (max) 40.12%。DeepSWE 1.1(由 Datacurve 使用 mini-swe-agent 测试框架运行):Fable (max) 70%,GPT 5.5 (xhigh) 67%,Opus 4.8 (max) 59%,Grok 4.5 53%,GLM 5.2 44%。SWE Marathon 解决率(pass@1):Grok 4.5 29.0%,Opus 4.8 (max) 26.0%,Fable (max) 24.0%,Opus 4.7 (max) 16.0%。Terminal Bench 2.1:Fable (max) 84.3%,GPT 5.5 (xhigh) 83.4%,Grok 4.5 83.3%,Opus 4.8 (max) 78.9%,Opus 4.7 (max) 78.9%。SWE Bench Pro 解决率:Fable (max) 80.4%,Opus 4.8 (max) 69.2%,Grok 4.5 64.7%,Opus 4.7 (max) 64.3%,GLM 5.2 62.1%,GPT 5.5 (xhigh) 58.6%。竞品数据均来自各开发者公开发布的系统卡或基准排行榜。
训练 Grok 4.5
Grok 4.5 在数万块 NVIDIA GB300 GPU 上进行了训练,并采用了专为大规模运行设计的训练与稳定性技术。除了原始 token 数量之外,我们在数据过滤和整理方面投入了大量精力:包括去重、质量评分和领域聚焦选择,以确保数据混合保持高覆盖度和高信号质量。
我们在强化学习上进行了规模化扩展,并高度关注每个 token 的智能水平。我们的 RL 训练覆盖数十万个任务,以多步骤软件工程及其他技术工作为核心,采用自动化评分与基于模型的评分。我们的训练栈专为高度异步训练而设计,因此智能体 rollout 可以持续运行数小时,同时在数万块 GPU 上持续进行学习。其结果是,在真实的工程与智能体任务中,推理变得更加智能且高效。
仅凭一条提示词构建而成
Grok 4.5 在编程方面能力极强,从具有挑战性的 Rust 和 C/C++ 任务,到从提示词到生产环境的端到端应用构建,均能胜任。以下是该模型仅凭一条提示词构建的一些示例。即使规格说明极为简略,Grok 4.5 也能高度熟练地创建设计精良、功能完整的端到端应用。
太阳系
制作一个精美的宇宙与太阳系模拟。应具备可调节时间的加速功能、逼真的运动、轨道和恒星。使用 threejs。HUD 应设计精良,符合现代设计原则。
制作一个精美的宇宙与太阳系模拟。应具备可调节时间的加速功能、逼真的运动、轨道和恒星。使用 threejs。HUD 应设计精良,符合现代设计原则。
app.localhost — 宇宙
比闪电模型更快
Grok 4.5 以 80 TPS 的快速模型速度提供服务。结合在相同任务上比最新领先模型高出两倍的 token 效率,该模型能以更快的速度和更低的成本为您提供智能结果。
Token 效率
每个 SWE Bench Pro 任务的平均输出 token 数
Grok 4.5 0
Opus 4.8 (max) 0
4.2 倍更少的 token
0 70k token
Token 效率,每个 SWE Bench Pro 任务的平均输出 token 数——Grok 4.5 平均使用 15,954 个输出 token 即可解决任务,约为 Opus 4.8 (max) 的 67,020 个 token 的 4.2 倍。
擅长办公任务
Grok 4.5 现已成为 Grok Build 中的默认模型。除了出色的编程能力外,Grok Build 还能够构建复杂的 Excel 模型,这些模型涉及从网络搜索资料、多工作表公式使用,甚至还能为将来参考留下便签或注释。
在 PowerPoint 和 Word 中,Grok 4.5 同样细致入微。该模型能够使用原生 PowerPoint 形状构建复杂图表,设计直观的幻灯片内容,并在 Word 中撰写清晰的文稿。
规划一份包含 5 张幻灯片的季度业务回顾
自动保存
Q3 回顾.pptx
搜索
开始 插入 设计 切换 动画 审阅
批注 共享
新建幻灯片 布局 B I U 排列 设计灵感 加载项 Grok
1
季度回顾 十月 · 财年 26
Q3 业务回顾
收入、利润率、管线——以及我们下一步的投资方向
01 收入 02 利润率 03 管线 04 展望
第 1 张幻灯片,共 5 张 · 英语(美国)100%
详细了解我们的 Word、PowerPoint 和 Excel 插件。
定价
与其他领先模型相比,Grok 4.5 的交付成本极具竞争力。Grok 4.5 的定价为每百万输入 token 2 美元,每百万输出 token 6 美元。该模型在 token 效率上还达到了同类领先模型的大约 2 倍,能够以不到一半的步骤数完成任务。总体而言,Grok 4.5 在单位时间和成本内提供了最高的智能水平。
快速开始
Grok 4.5 现已可在 Grok Build、所有套餐的 Cursor 以及 SpaceXAI 控制台中使用。只需获取一个 API 密钥,通过几行代码即可开始使用:
curl -s https://api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.5",
"input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
}'
创建 API 密钥
立即通过 SpaceXAI API 开始使用 Grok 4.5 进行构建。
API 文档
阅读文档,将 Grok 4.5 集成到你的技术栈中。
在 Grok Build 中免费试用
我们限时提供 Grok 4.5 在 Grok Build 和 Cursor 中的免费使用。立即访问 x.ai/cli 开始体验。
Introducing Grok 4.5 Real-world engineering excellence Training Grok 4.5 Faster than flash models Excels at Office work Pricing Getting started
Today, we're launching Grok 4.5, SpaceXAI's smartest model built to excel at coding, agentic tasks, and knowledge work. It's our strongest model ever and was trained alongside Cursor.
Real-world engineering excellence
Grok 4.5 was trained on datasets spanning knowledge in coding, science, engineering, and math. With both intelligent and efficient reasoning, Grok 4.5 excels at real engineering tasks and exceeds comparable leading models at these tasks.
DeepSWE 1.0 DeepSWE 1.1 SWE Marathon Terminal Bench 2.1 SWE Bench Pro
Competitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards
Benchmark bar charts of model scores. DeepSWE 1.0 (eval created by Datacurve, run with each model provider’s harnesses by AA): Fable (max) 66.1%, GPT 5.5 (xhigh) 64.31%, Grok 4.5 62.0%, Opus 4.8 (max) 55.75%, Opus 4.7 (max) 40.12%. DeepSWE 1.1 (mini-swe-agent harness run by Datacurve): Fable (max) 70%, GPT 5.5 (xhigh) 67%, Opus 4.8 (max) 59%, Grok 4.5 53%, GLM 5.2 44%. SWE Marathon resolution rate (pass@1): Grok 4.5 29.0%, Opus 4.8 (max) 26.0%, Fable (max) 24.0%, Opus 4.7 (max) 16.0%. Terminal Bench 2.1: Fable (max) 84.3%, GPT 5.5 (xhigh) 83.4%, Grok 4.5 83.3%, Opus 4.8 (max) 78.9%, Opus 4.7 (max) 78.9%. SWE Bench Pro resolve rate: Fable (max) 80.4%, Opus 4.8 (max) 69.2%, Grok 4.5 64.7%, Opus 4.7 (max) 64.3%, GLM 5.2 62.1%, GPT 5.5 (xhigh) 58.6%. Competitor figures are drawn from the respective developers’ published system cards or benchmark leaderboards.
Training Grok 4.5
Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs, with training and stability techniques designed for large-scale runs. Beyond raw token volume, we invested heavily in data filtering and curation: deduplication, quality scoring, and domain-focused selection so that the data mixture stayed high-coverage and high-signal.
We scaled reinforcement learning with a strong focus on per-token intelligence. Our RL training covers hundreds of thousands of tasks, centered on multi-step software engineering and other technical work, with automated and model-based grading. Our stack is built for highly asynchronous training, so agentic rollouts can run for many hours while learning continues across tens of thousands of GPUs. The result is more intelligent and efficient reasoning on real engineering and agentic tasks.
Built with one prompt
Grok 4.5 is incredibly capable at coding, from challenging Rust and C/C++ tasks to end-to-end app building from prompt to production. Below are some examples built by the model with one prompt. Grok 4.5 is highly proficient at creating well-designed, end-to-end functional apps even with minimal specification.
Solar system
Make a beautiful simulation of the universe and solar system. should be sped up with adjustable time, realistic motion, orbits, stars. use threejs. Make the HUD well styled and conform to modern design principles.
Make a beautiful simulation of the universe and solar system. should be sped up with adjustable time, realistic motion, orbits, stars. use threejs. Make the HUD well styled and conform to modern design principles.
app.localhost — cosmos
Faster than flash models
Grok 4.5 is served at fast-model speeds of 80 TPS. Combined with twice greater token efficiency than the latest leading models at the same tasks, the model delivers intelligent results to you more quickly and at far lower costs.
Token efficiency
avg. output tokens per SWE Bench Pro task
Grok 4.5 0
Opus 4.8 (max)0
4.2×fewer tokens
0 70 k tokens
Token efficiency, average output tokens per SWE Bench Pro task — Grok 4.5 resolves tasks with 15,954 output tokens on average, about 4.2× fewer than Opus 4.8 (max) at 67,020
Excels at Office work
Grok 4.5 is now the default model in Grok Build. In addition to its coding proficiency, Grok Build is capable of building complex Excel models that involve research from the web, multi-sheet formula use, and even leaves stickies or notes behind for future reference.
In PowerPoint and Word, Grok 4.5 is similarly meticulous. The model is capable of using native PowerPoint shapes to build complex diagrams, designing intuitive slide content, and writing clear prose in Word.
Outline a 5-slide quarterly business review
AutoSave
Q3 Review.pptx
Search
Home Insert Design Transitions Animations Review
Comments Share
New Slide Layout B I U Arrange Design Ideas Add-ins Grok
1
Quarterly review Oct · FY26
Q3 Business Review
Revenue, margins, pipeline — and where we invest next
01 Revenue 02 Margins 03 Pipeline 04 Outlook
Slide 1 of 5 · English (US)100%
Learn more about our plugins for Word, PowerPoint, and Excel.
Pricing
Grok 4.5 is delivered at an incredibly competitive cost compared to other leading models. Grok 4.5 is priced at $2 per million input tokens and $6 per million output tokens. The model also achieves roughly 2x the token efficiency of comparable leading models, solving tasks in under half the number of steps. Overall, Grok 4.5 delivers the highest intelligence per unit of time and cost.
Getting started
Grok 4.5 is available today in Grok Build, in Cursor on all plans, and from the SpaceXAI console. Simply grab an API key and get started in a few lines of code:
curl -s https://api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.5",
"input": "Find and fix the bug, then explain it: function median(a){a.sort();return a[a.length/2]}"
}'
Create an API Key
Start building with Grok 4.5 today via the SpaceXAI API.
API Docs
Read the docs and integrate Grok 4.5 into your stack.
Try it in Grok Build for free
We’re offering free Grok 4.5 usage for a limited time inGrok Buildand Cursor. Get started today atx.ai/cli.