我们最新的 Gemini 模型为智能体工作流和网络安全带来了下一代智能。
Tulsee Doshi 产品管理高级总监
Raluca Ada Popa
Google DeepMind Gemini 安全负责人
x.com Facebook LinkedIn Mail 复制链接
继三周前发布的 3.7 Flash 之势,并在短短六周内完成我们的第三次 Flash 版本发布,今天我们推出 Gemini 3.8——这是我们迄今最强的推理与编码模型,同时保持了与 3.7 相同的速度和低成本。Gemini 3.8 提供两个版本:
Gemini 3.8 Flash:我们最智能的“多面手”模型,在软件工程、智能体任务以及专业领域的复杂多步推理方面,相较 3.7 Flash 实现了显著提升。该模型以与 3.7 Flash 相同的首发价格 1 提供,即每百万输入 token 0.75 美元、每百万输出 token 3.75 美元。Gemini 3.8 Flash Cyber:我们最强的网络安全模型,在漏洞检测和自动修复方面具备前沿水平,通过我们全新的 Fairwind 计划向受信任的防御者开放。
虽然这两款发布针对不同的部署环境而设计,但它们的底层都由相同的基础智能驱动,并通过长时运行的智能体循环进一步加速——这种循环旨在递归评估并优化底层模型。这一共享核心在编码和推理能力上的显著提升,得益于多项创新,其中包括在要求极高的网络安全领域所进行的严格训练。
Gemini 3.8 Flash:为长时程编码与自主智能体而生
Gemini 3.8 Flash 相比 3.7 Flash 带来了显著提升,其性能往往接近成本更高的前沿模型。
在 DeepSWE v1.1(长时程软件工程)基准上,3.8 Flash 在端到端自主解决复杂工程问题方面超越了大多数规模更大的前沿模型,而成本仅为其一小部分。
此外,3.8 Flash 在专业知识领域展现出了关键企业级自主运行所需的可靠性。在需要高级分析与报告能力的量化及专业领域,3.8 Flash 在 Vals Finance Agent V2 和 Harvey 的 Legal Agent Benchmark 等基准上均优于 3.7 Flash 及其他前沿模型。3.8 Flash 在 HLE-Verified 上还取得了 54.9% 的成绩,展示了其在 STEM、人文学科和专业领域处理多步推理的能力。
这些性能提升源于一个核心设计选择:3.8 Flash 更加“卖力”。在处理复杂任务时,它表现出更强的勤勉度——执行额外的推理步骤,并迭代式地调用工具。有时,模型可能会使用更多 token 来最大化性能,尤其是在较高努力水平下。
对于算力效率是首要约束条件的应用,开发者可以利用较低的努力水平来最小化 token 开销,或者继续依赖 Gemini 3.7 Flash——该版本在效率优先的工作负载上仍获得全面支持。
Gemini 3.8 Flash 在 Google Antigravity 中仅凭一个简单的提示词,配合一条循环指令,就构建出了这款游戏。该游戏利用谜题、环境叙事以及由 Nano Banana 生成的纹理,打造出一个沉浸式 3D 关卡,让你扮演一位在城堡中穿行的巫师。
Gemini 3.8 Flash 在 Google Antigravity 中仅凭单个提示词,就构建出一个功能完整的 DOS 版 Google Maps,该版本完全可玩,包含地点查询、路线规划和街景功能。
探索由 Gemini 3.8 Flash 在 Google Antigravity 中构建的著名地理景点地形图,它使用了来自美国地质调查局的真实数据集,可呈现实时剖面、二维投影和科学解释。
Hardware Anatomy 是一个交互式 3D 可视化工具,由 Gemini 3.8 Flash 在 Google AI Studio 中构建,可为硬件设备生成符合物理比例的拆解图,并渲染出逼真的 Three.js 效果。它会自动将设备分解为多个图层,你可以通过一个“拆解”滑块来展开并检视这些图层。
Gemini 3.8 Flash Cyber:专家级网络性能
Gemini 3.8 Flash Cyber 通过 Fairwind 计划向一组受信任的防御者开放,凭借 Flash 的速度和成本优势,支持快速迭代,在当今复杂的网络安全环境中提供了决定性优势。
自主漏洞发现
在行业标准漏洞发现基准 CyberGym 上,Gemini 3.8 Flash Cyber 在自主漏洞发现方面展现出前沿级性能。它超越了 3.5 Flash Cyber 以及规模显著更大的前沿模型。
为了更好地反映现实世界的防御需求——这些需求并不仅限于 CyberGym 中那样的 C/C++ 代码库——我们还使用一个全面的内部基准对 Gemini 3.8 Flash Cyber 进行了评估,该基准要求模型在横跨 20 种编程语言的复杂代码库中发现各种漏洞。在此基准上,该模型相较我们之前的模型实现了显著飞跃,成功率超过 70%。
自动化漏洞修复
在开发 Gemini 3.8 Flash Cyber 时,我们特别专注于为防御者配备专家级能力,使其在与攻击者的对抗中占据优势。正因如此,我们从一开始就投入于漏洞修复,并将其优先级置于漏洞利用等进攻性能力之上。
CWE-Bench 由 Collinear 运营,是一个极具挑战性的外部补丁能力评测基准。在该基准上,Gemini 3.8 Flash Cyber 处于帕累托前沿:其 pass@1 达到 47.2%,而领先的前沿模型为 47.8%,但成本却显著更低。
实际影响:保护 Google 的代码安全
我们已经在使用 Gemini 3.8 Flash Cyber 来保护 Google 内部的代码安全。例如:
Chrome 安全团队发现,3.8 Flash Cyber 针对 Chrome 漏洞生成的正确补丁数量,是那些体积大得多的最佳商业模型的 2.6 倍。Wiz 发现,Gemini 3.8 Flash Cyber 在其内部渗透测试基准上的召回率高出 7.5-9.7%,而成本仅为其他领先前沿模型的 1/2.3 至 1/5.2。Google 的云漏洞研究团队利用 3.8 Flash Cyber 模型,在不到 2 小时内发现了一个关键的基础性漏洞,而这类漏洞的研究与发现通常需要数月时间。
我们的 Fairwind 计划合作伙伴的评价
以安全为核心理念构建
根据我们的前沿安全框架,3.8 Flash 内置了针对化学、生物、放射性和核(CBRN)以及网络攻击等领域的滥用防护措施,同时支持有益的使用场景。3.8 Flash Cyber 在网络安全方面配备了更为宽松的缓解措施,因此仅向需要更全面网络能力的受信任防御者提供。
Gemini 3.8 系列模型在提示词注入鲁棒性方面也取得了显著进步(据 Gray Swan 评测),为 Gemini 模型用户抵御与提示词注入相关的恶意攻击提供了保护。
Gemini 3.8 Flash 与 Cyber:即刻开始使用
开发者:使用 3.8 Flash 进行构建,在 Google Antigravity 中探索智能体优先的工作流,或通过 Google AI Studio 和 Android Studio 在 Gemini API 中即刻开始构建,也可在 Stitch 中生成 UI。从我们的开发者文档开始上手。企业用户:在 Gemini Enterprise 中访问 3.8 Flash。消费者:Google AI Pro 和 Ultra 订阅用户可在 Gemini 应用、Google 搜索中的 AI Mode 以及 Google Sheets 中的 Gemini 中使用 3.8 Flash。Cyber:通过我们全新的 Fairwind Program,我们为受信任的政府机构、关键基础设施运营商及软件维护者提供 Gemini 3.8 Flash Cyber 的优先访问权限。申请访问。
1
introductory价格有效期至 2026 年 12 月 31 日。自 2027 年 1 月 1 日起,将适用每 1M 输入 token $1.50、每 1M 输出 token $7.50 的价格。
相关报道
安全与保障 ### 面向政府与企业的主动式网络防御 作者:Four Flynn AI ### 我们于 2026 年 8 月发布的最新 AI 资讯 作者:Google 团队新闻 Gemini 模型 ### 借助 Gemini 推出智能体视频理解功能 作者:Rohan Doshi & Mario Lučić 开发者工具 ### Gemini Omni 1.1 Flash 让你以更强的掌控力进行构建 作者:Anish Nangia & Alisa Fortin Gemini 模型 ### 借助 Gemini 3.5 Transcribe 实现智能转写 作者:Diego Melendo Casado & Luke Leonhard Gemini 模型 ### “全栈”AI 究竟意味着什么? 作者:Lindsey Lanquist
Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity.
Tulsee Doshi Senior Director, Product Management
Raluca Ada Popa
Gemini Security Lead, Google DeepMind
x.com Facebook LinkedIn Mail Copy link
Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. Gemini 3.8 introduces 2 variants:
Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price 1 as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program.
While tailored for different deployment environments, both of today's releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity.
Gemini 3.8 Flash: built for long-horizon coding and autonomous agents
Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models.
On DeepSWE v1.1 (Long-Horizon Software Engineering) 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.
Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains. In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM, humanities, and professional fields.
These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.
For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.
Gemini 3.8 Flash built this game with a simple prompt using a looping instruction in Google Antigravity. The game uses puzzles, environmental storytelling, and textures generated with Nano Banana to create an immersive 3D level in which you play a wizard navigating a castle.
Gemini 3.8 Flash builds a fully functional DOS version of Google Maps in a single prompt in Google Antigravity that is fully playable with locations, directions, and Street View.
Explore realtime cross-sections, 2D projections, scientific explanations in a topographic map of famous geographical sites built with Gemini 3.8 Flash in Google Antigravity using real datasets from the U.S. Geological Survey.
Hardware Anatomy is an interactive 3D visualizer built with Gemini 3.8 Flash in Google AI Studio that generates realistic Three.js renderings of physically-proportioned teardowns for hardware devices. It automatically decomposes devices into layers you can explode and inspect with a deconstruction slider.
Gemini 3.8 Flash Cyber: expert cyber performance
Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program, provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration.
Autonomous vulnerability discovery
On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models.
To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%.
Automated patching
With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation.
CWE-Bench, run by Collinear, is a challenging external benchmark for patching capabilities. On this benchmark, Gemini 3.8 Flash Cyber is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost.
Real-world impact: securing Google’s code
We’re already using Gemini 3.8 Flash Cyber to secure code across Google. For example:
The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger. Wiz found that Gemini 3.8 Flash Cyber achieves +7.5-9.7% higher recall on their internal penetration testing benchmark for a 2.3-5.2x lower cost compared to other leading frontier models. Google’s Cloud Vulnerability Research team leveraged the 3.8 Flash Cyber model to find a critical foundational vulnerability in less than 2 hours, a vulnerability for which research and discovery usually takes months.
What our Fairwind Program partners are saying
Built with safety in mind
3.8 Flash ships with safeguards against misuse in the domains of Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense, while enabling beneficial use cases, as per our Frontier Safety Framework. 3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders who require a more comprehensive set of cyber capabilities.
Gemini 3.8 models have also made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection related malicious attacks.
Gemini 3.8 Flash and Cyber: get started today
Developers: Build with 3.8 Flash and explore agent-first workflows in Google Antigravity or start building today in the Gemini API via Google AI Studio and Android Studio, or generate UIs in Stitch. Get started with our developer docs. Enterprises: Access 3.8 Flash in Gemini Enterprise. Consumers: 3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app, AI Mode in Google Search and Gemini in Google Sheets. Cyber: Through our new Fairwind Program, we’re providing trusted government authorities, as well as critical infrastructure operators and software maintainers with prioritized access to Gemini 3.8 Flash Cyber. Apply for access.
1
Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
Related stories
Safety & Security ### Proactive cyber defense for governments and enterprises By Four Flynn AI ### The latest AI news we announced in August 2026 By News from Google Team Gemini models ### Introducing agentic video understanding with Gemini By Rohan Doshi & Mario Lučić Developer tools ### Gemini Omni 1.1 Flash lets you build with more control By Anish Nangia & Alisa Fortin Gemini models ### Intelligent transcription with Gemini 3.5 Transcribe By Diego Melendo Casado & Luke Leonhard Gemini models ### What does “full-stack” AI actually mean? By Lindsey Lanquist