我们正在将 Gemini 扩展为一个世界模型,使其能够通过模拟世界的各个方面来制定计划并想象新的体验。
过去十年间,我们为现代人工智能时代奠定了大量基础,从开创所有大语言模型所依赖的 Transformer 架构,到开发能够像 AlphaGo 和 AlphaZero 一样学习和规划的智能体系统。
我们将这些技术应用于量子计算、数学、生命科学和算法发现领域,并取得了突破。同时,我们继续深化基础研究的广度与深度,致力于为通用人工智能(AGI)所需的下一项重大突破进行发明创造。
正因如此,我们正致力于将我们最优秀的多模态基础模型 Gemini 2.5 Pro 扩展为一个“世界模型”,使其能够像大脑一样,通过理解和模拟世界的各个方面来制定计划并想象新的体验。
我们朝着这个方向已经努力了一段时间,从开创性地训练智能体掌握围棋和星际争霸等复杂游戏,到构建能够根据单张图像提示生成可交互的 3D 模拟环境的 Genie 2。
我们已经可以看到这些能力正在显现的证据:Gemini 利用世界知识和推理来表征与模拟自然环境的能力;Veo 对直觉物理学的深刻理解;以及 Gemini Robotics 教导机器人抓取物体、遵循指令并即时调整的方式。
将 Gemini 打造成一个世界模型,是开发一种更通用、更有用的新型人工智能——通用 AI 助手——的关键一步。这是一种智能的 AI,能够理解你所处的上下文,并能代表你在任何设备上进行规划和采取行动。
将 Project Astra 的实时能力引入我们的产品
我们的最终愿景是将 Gemini 应用转变为一个通用 AI 助手,它将为我们执行日常任务、处理琐碎行政事务并呈现令人愉悦的新推荐——从而提升我们的生产力,丰富我们的生活。
这一切始于我们去年在研究原型项目 Astra 中首次探索的能力,例如视频理解、屏幕共享和记忆功能。
过去一年里,我们一直在将此类能力整合到 Gemini Live 中,让更多人今天就能体验。我们持续不懈地改进,并在前沿领域探索新的创新。例如,我们升级了语音输出,通过原生音频使其更加自然;改进了记忆功能;并增加了计算机控制能力。
我们现在正从受信任的测试者那里收集关于这些能力的反馈,并致力于将它们引入 Gemini Live、搜索中的新体验、面向开发者的 Live API,以及眼镜等新形态设备。
在这一过程的每一步,安全与责任都是我们工作的核心。我们最近开展了一项大型研究项目,探讨先进 AI 助手相关的伦理问题,这项工作持续为我们的研究、开发和部署提供指导。
构建能为你多任务处理的 AI
我们还通过 Project Mariner 探索智能体能力如何帮助人们进行多任务处理。这是一个研究原型,从浏览器入手,探索人机交互的未来。
自去年十二月推出 Project Mariner 以来,我们一直与一组受信任的测试者密切合作,收集反馈并改进其实验性能力。
Project Mariner 现在包含一个智能体系统,可以同时完成多达十项不同的任务。这些智能体可以帮助你查找信息、预订、购物、做研究等等——所有这些都能同时进行。
更新后的 Project Mariner 已面向美国的 Google AI Ultra 订阅用户开放。我们正在将其计算机使用能力引入 Gemini API,并计划在今年内将其更多能力带到 Google 产品中。在搜索和 Gemini 应用中,了解更多关于我们智能体能力的信息。
凭借这一点以及我们所有开创性的工作,我们正在构建更加个性化、主动且强大的 AI,以丰富我们的生活,加速科学进步的节奏,并开启一个发现与奇迹的新黄金时代。
We’re extending Gemini to become a world model that can make plans and imagine new experiences by simulating aspects of the world.
Over the last decade, we’ve laid a lot of the foundations for the modern AI era, from pioneering the Transformer architecture on which all large language models are based, to developing agent systems that can learn and plan like AlphaGo and AlphaZero.
We’ve applied these techniques to make breakthroughs in quantum computing, mathematics, life sciences and algorithmic discovery. And we continue to double down on the breadth and depth of our fundamental research, working to invent the next big breakthroughs necessary for artificial general intelligence (AGI).
This is why we’re working to extend our best multimodal foundation model, Gemini 2.5 Pro, to become a “world model” that can make plans and imagine new experiences by understanding and simulating aspects of the world, just as the brain does.
We’ve been taking strides in this direction for a while, from our pioneering work training agents to master complex games like Go and StarCraft, to building Genie 2, which is capable of generating 3D simulated environments that you can interact with, from a single image prompt.
Already, we can see evidence of these capabilities emerging in Gemini’s ability to use world knowledge and reasoning to represent and simulate natural environments, Veo’s deep understanding of intuitive physics, and the way Gemini Robotics teaches robots to grasp, follow instructions and adjust on the fly.
Making Gemini a world model is a critical step in developing a new, more general and more useful kind of AI —a universal AI assistant. This is an AI that’s intelligent, understands the context you are in, and that can plan and take action on your behalf, across any device.
Bringing Project Astra’s live capabilities into our products
Our ultimate vision is to transform the Gemini app into a universal AI assistant that will perform everyday tasks for us, take care of our mundane admin and surface delightful new recommendations — making us more productive and enriching our lives.
This starts with the capabilities we first explored last year in our research prototype Project Astra, such as video understanding, screen sharing and memory.
Over the past year, we’ve been integrating capabilities like these into Gemini Live for more people to experience today. We continue to relentlessly improve and explore new innovations at the frontier. For example, we upgraded voice output to be more natural with native audio, we’ve improved memory and added computer control.
We’re now gathering feedback about these capabilities from trusted testers and are working to bring them to Gemini Live, to new experiences in Search, the Live API for developers and new form factors, like glasses.
Through every step of this process, safety and responsibility are central to our work. We recently conducted a large research project, exploring the ethical issues surrounding advanced AI assistants, and this work continues to inform our research, development and deployment.
Building AI that can multitask for you
We’ve also been exploring how agentic capabilities can help people multitask, with Project Mariner. This is a research prototype that explores the future of human-agent interaction, starting with browsers.
Since launching Project Mariner last December, we’ve been working closely with a group of trusted testers to gather feedback and improve its experimental capabilities.
Project Mariner now includes a system of agents that can complete up to ten different tasks at a time. These agents can help you look up information, make bookings, buy things, do research and more — all at the same time.
The updated Project Mariner is available to Google AI Ultra subscribers in the U.S. We're bringing its computer use capabilities into the Gemini API, and we’re planning to bring more of its capabilities to Google products throughout the year. Read more about our agentic capabilities in Search and the Gemini app.
With this, and all our groundbreaking work, we’re building AI that’s more personal, proactive and powerful, enriching our lives, advancing the pace of scientific progress and ushering in a new golden age of discovery and wonder.