Topic · 主题全部主题 →

Google / Gemini

Google 与 DeepMind 的 AI 动态:Gemini 系列、Veo 视频模型、研究成果与产品生态的持续追踪。

2,182条收录
296条精选

精选归档 · 第 8 页

141160 条 · 共 296

5月30日

星期六 · 4 条
03:08
01:38
Google Blog:AI(RSS)精选
AI 评分 74/100
Gemini Omni 与 Gemini 3.5 的 11 个实战展示

Google 在 2026 年 Google I/O 大会上发布了新一代多模态模型 Gemini Omni 与 Gemini 3.5,并同步提供了 11 个视频,集中演示了这两款模型在实际场景中的能力。

另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Google 官方放出的这组视频演示,直接展示了 Gemini Omni 和 3.5 的实际表现,比参数和 benchmark 更直观,做多模态应用的可以逐帧研究。

5月29日

星期五 · 4 条
15:21
IT之家(RSS)精选
AI 评分 70/100
谷歌 DeepMind CEO 哈萨比斯:AGI 最快三年内到来,研发速度远超预期

谷歌 DeepMind 首席执行官德米斯·哈萨比斯预测,AGI 研发速度远超预期,最快可能在 2029 年至 2030 年前后出现。作为 AlphaGo、AlphaFold 的主导者,他认为当前 AI 智能体是未来更强智能的预演,随着多模态和自主决策能力成熟,三年内迎来 AGI 关键突破已非科幻。但他同时警示,全球社会对 AGI 到来的准备严重不足,必须提前建立规则与防护机制。


推荐理由:哈萨比斯作为造出 AlphaFold 的诺贝尔奖得主,三年内 AGI 的判断不是空话,他同时强调社会完全没准备好,这种紧迫感比单纯的时间表更值得看。
05:12
Google Research:Blog(网页)精选
AI 评分 79/100
创新时代:Google Research 在 I/O 2026

Google Research 在 I/O 2026 大会上展示了其在多个前沿领域的技术进展,包括应用AI、基础机器学习算法以及量子AI等。本次大会的核心主题是展示其在将科学发现与研究成果转化为现实世界影响方面的持续努力。

另有 1 家信源报道Google Blog:AI(RSS)
推荐理由:Google 把研究成果直接发 Nature,ERA 和 Co-Scientist 这套工具让 AI 从写诗进化到做实验,健康 AI 的临床验证数据也很扎实,搞科研的可以蹲一下访问资格。
02:41
Google Developers Blog(RSS)精选
AI 评分 73/100
使用 Google Pay & Wallet Developer MCP server 加速你的集成工作流

Google 推出 Google Pay & Wallet Developer MCP server,这是一款开放标准工具,旨在将 AI 开发助手和 IDE 安全连接到实时的 API 与账户上下文。开发者无需离开开发环境,即可搜索官方文档、验证 Wallet pass 定义、检查集成状态以及管理商户账户。该集成旨在通过减少上下文切换并提供实时、可靠的 AI 支持来减少开发摩擦,从而加速开发工作流。


推荐理由:这是 Google 为支付场景做的 MCP 服务器,把文档和账户操作直接塞进 IDE,减少上下文切换,做 Google Pay 集成的开发者可以试试看。

5月28日

星期四 · 5 条
23:41
Google Developers Blog(RSS)精选
AI 评分 64/100
社区如何利用Tunix和TPU训练Gemma学会"思考"

Google在Kaggle举办的Tunix黑客马拉松,挑战开发者利用TPU和有限算力,将小型基础模型转变为通用推理引擎。获胜团队通过多阶段后训练流程实现了这一目标,该流程结合了监督微调(SFT)与GRPO、SimPO等先进对齐技术。比赛结果表明,社区能够借助开源资源成功训练出高能力的结构化推理模型。


推荐理由:Google 官方比赛总结,证明用 Kaggle TPU 和开源工具就能把 Gemma 训练出不错推理能力,对想自己微调模型的小团队是个实用参考。
05:52
Google Gemini@GeminiApp精选
AI 评分 77/100
Gemini Omni轻松转换视频视觉风格Easily transform your videos into new visual styles with Gemini Omni.Just upload a video or photo and ask Gemini to apply a look or style to your final output.使用 Gemini Omni 轻松将您的视频转换为新的视觉风格。 只需上传视频或照片,并要求 Gemini 为您的最终输出应用某种外观或风格。
另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Gemini 终于把图像风格迁移做到视频上了,并且直接集成到 Omni 里,不需要任何剪辑软件,对短视频创作者是个小但实用的更新。
01:39
Google Developers Blog(RSS)精选
AI 评分 66/100
Google Pay 最新更新

Google Pay 正向"智能体商务"演进,推出了通用商务协议和新的 MCP 服务器,允许 AI 智能体管理集成与分析趋势。Android 平台更新引入了动态回调以支持快速结账,并通过 WebView 将支付功能扩展至社交媒体应用。此外,平台还推出了跨设备生物认证和新的交易信号,旨在帮助商家减少流程摩擦。


推荐理由:Google Pay 往 agentic commerce 迈了一大步,新的通用协议和 MCP server 让 AI agent 能直接管支付和分析,做 agent 或支付的开发者都得看看。
01:34
Google Research:Blog(网页)精选
AI 评分 70/100
通过零信任聚合实现的隐私分析

Google Research 推出了一种新的隐私分析解决方案。该方案结合了一种新的密码学安全聚合协议与可信执行环境(TEE)的透明性,旨在实现前沿的隐私与安全保证。其核心是基于零信任原则,通过密码学与硬件保护的结合,确保系统仅能获取群体的匿名化聚合洞察。


推荐理由:Google 的隐私聚合新方案把多轮交互砍成一次提交,对做设备端联邦分析的人来说是工程上的一大步,而且结合 TEE 做双层防护,这个思路值得抄。
00:35
Chubby♨️@kimmonismus精选
AI 评分 80/100
与Google搜索产品副总裁Robby Stein的访谈:AI原生搜索时代I sat down with Robby Stein (@rmstein), Google’s VP of Product for Search, at @Google I/O.Robby is one of the most interesting product leaders in tech: he helped build Instagram Stories, Reels and Close Friends, and now leads core Google Search products including AI Overviews, AI Mode, Lens and ranking.We talked about one of the biggest shifts in the history of the web:Google Search becoming AI-native.Topics we covered:• AI Mode and whether it is an evolution of Search or a reinvention of it • how Google breaks complex questions into multiple searches behind the scenes • why AI search is much more expensive to run than traditional search • whether Google’s TPUs and infrastructure give it an advantage no one else can match • why Search volume is growing instead of being cannibalized by AI • the tension between great AI answers and traffic for publishers • how Google decides which sources and links to show • what a better internet could look like if AI Search works as intendedThe big question behind the whole conversation: If Google gives you the answer directly, what happens to the link-based web?A small caveat: sadly the microphones didnt work properly. Therefore the audio quality in this episode isn't perfect due to a recording issue - we appreciate your understanding.本文记录了与Google搜索产品副总裁Robby Stein在Google I/O的访谈,核心探讨Google Search向"AI原生"模式的重大转变。讨论话题包括AI Mode是进化还是重塑、如何将复杂问题拆解为多轮搜索、AI搜索的高运行成本、Google TPU及基础设施的优势、AI时代搜索量不减反增的原因,以及优质AI回答与出版商流量之间的张力。访谈还涉及Google决定展示哪些信息源与链接的逻辑,并围绕一个核心问题展开:如果Google直接给出答案,传统的基于链接的网页生态将走向何方?
另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Google 搜索 VP 首次拆解 AI Mode 背后的成本逻辑、流量分配和 TPU 优势,比 I/O 演讲深得多,做搜索和内容生态的都值得听。

5月27日

星期三 · 2 条
18:20
HuggingFace Daily Papers(社区热门论文)精选
AI 评分 72/100
Gemini Embedding 2:来自Gemini的原生多模态嵌入模型

Google DeepMind推出Gemini Embedding 2,这是一款原生多模态嵌入模型,支持在统一表示空间中嵌入视频、音频、图像和文本。该模型利用Gemini的多模态能力,通过大规模对比学习实现SOTA性能。在关键基准上表现优异:MSCOCO取得62.9 R@1,Vatex取得68.8 NDCG@10,MTEB multilingual达到69.9,MTEB Code达到84.0,超越了专用模型。其统一能力使其适用于RAG、推荐与搜索等下游任务,并在天文学、生物科学、艺术和烹饪等专业领域展现出强大的零样本性能。


推荐理由:Google 把多模态嵌入统一到一个模型里了,文本、代码、跨模态检索全面刷榜,做 RAG 和搜索的该认真看看了。
05:28
Google AI@GoogleAI精选
AI 评分 75/100
Gemini Omni 视频提示词使用指南http://x.com/i/article/2059377716965888000Mastering Gemini Omni: The Ultimate Video Prompting GuideLast week, we introduced Gemini Omni—our newest model designed to create anything from any input, starting with video.You can experience the speed and creativity of Gemini Omni Flash today across @geminiapp, @GoogleFlow, @GoogleFlowMusic, and on @YouTube Shorts and Create.To help you push the boundaries of what’s possible, here are five tips to get the most out of Gemini Omni’s advanced video generation capabilities.1. Leverage Real-World KnowledgeYou don’t need to over-explain the world to Gemini Omni. It’s built with Gemini’s deep understanding of history, science, and culture, so it can reliably create outputs that look, feel, and move realistically. Skip the granular descriptions. Use cultural touchstones, historical eras, or scientific terms directly in your prompt.Example Prompts:• [The video shows items of the alphabet. An unusual item starting with each letter is shown sitting on a table (like a Capybara for C, disco globe for D and Lava Lamp for L). All 26 letters must be represented by 26 items with matching lower thirds displaying the letter. Only one item and lower third at a time. Each lower third must look like a black marker written on a slip of paper in the bottom left. Rapid fire, roughly 9 frames per item at 24FPS. Last frame is a slip of paper "THE END." The whole video is accompanied by calm smooth music]• [Astronaut's POV on Mars]• [A marble rolling fast on a chain reaction style track, continuous smooth shot]2. Take Control of Text RenderingGemini Omni not only has advanced text rendering capabilities, it even allows you seamlessly integrate text into your visuals. You can specify typography, spatial placement, animation styles, and complex visual effects like double exposures all perfectly synced to the action in your video.Example Prompts:• [word by word, one word on the screen at a time: did, you, know, that, this, model, can, do, pretty, good, text!? Each word appears with a different animated style, perfect pacing to a rhythm, sizzle reel]• [Overlay motion-tracked, minimalist text commentary onto the physical environment of the video. This text represents [the subject] deadpan, immediate inner monologue that’s observant, slightly absurd, and life-contemplating. Think “intrusive thoughts.” Clean, white, lowercase sans-serif text (like Helvetica or Inter). The text hovers in 3D space, connected to the subjects being commented on via ultra-thin, crisp, white leader lines]3. Direct Your Camera Like a ProThink like a cinematographer. Gemini Omni responds incredibly well to precise videography directions, camera types, and framing instructions. Try integrating these terms into your next prompt:Example prompts:• Shots & Angles: "One continuous shot", "oner", "static", "locked off", or "fixed angle."• Camera Movements: "Push in", "punch in", "pan left", or "dolly zoom."• Camera Styles: "Natural smartphone zoom", "vintage film camera", or "grainy webcam style."4. Edit Iteratively (and keep what works)Every great video is made in the edit. With Gemini Omni, you don't need to rewrite your entire prompt from scratch to fix a single mistake. Ask for specific, targeted updates, like changing a background or swapping a caption. Omni will preserve the core structure of your video across multiple amends, letting you focus only on what needs tweaking.Example prompts:• [Transport the violin to a new environment]• [Make the violin invisible]• [Change the camera angle so it’s looking over the violinist’s shoulder]5. Change the Action on the FlyWant to alter a character's pacing or emotion mid-scene? You can directly prompt Gemini Omni to modify how a subject moves or interacts with their environment without breaking the continuity of the character model.Example prompts:• [Make the character walk on their tiptoes]• [Speed up the pacing]• [Have them leap into the air]Start CreatingThe director’s chair is yours. Try out these prompting techniques with Gemini Omni Flash, and tag @GoogleAI to show us what you create!Google 发布了其多模态模型 Gemini Omni 的视频生成功能使用指南。该模型可通过 Gemini 应用、Google Flow 等平台体验。指南包含五项提示词技巧:利用模型已有的现实世界知识进行简洁描述;精确控制文本在视频中的渲染与排版;使用专业镜头指令(如推拉摇移)像电影摄影师一样调度画面;通过迭代编辑高效修改视频;以及在生成中直接调整角色的动作节奏或情绪。其核心在于通过精准的提示词引导模型生成复杂且可控的视频内容。另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Google 官方放出的视频提示技巧,没有废话全是可复制的 prompt,想玩 Gemini Omni 的创作者可以直接抄作业。

5月26日

星期二 · 2 条
03:58
Chubby♨️@kimmonismus精选
AI 评分 79/100
苹果据称正使用定制版1.2T参数Google模型重塑下一代SiriApple isn’t just adding Gemini to Siri. It is reportedly using a custom 1.2T-parameter Google model as the brain behind parts of the next Siri overhaul (Reuters).That is no small size. The question, therefore, is how Apple’s Gemini will perform and how fast it runs. In particular, simple queries are expected to run locally.Gemini 3.5 Flash is estimated to have around 300 billion parameters; Apple’s model will thus be significantly larger. However, the question also remains whether size actually pays off in this context. Apple’s model must deliver answers to everyday queries quickly-and be fast enough while doing so.Anyways, next months will be exciting in so many ways:-WWDC / Apple Intelligence with Gemini • GPT-5.6 • Sonnet4.8/Opus 4.8 (probably) -Gemini 3.5 pro (confirmed)Super hyped!据报道,苹果为改造下一代Siri,正使用一个定制版、参数规模达1.2T的Google大模型作为其核心,这显著大于预估约300B参数的Gemini 3.5 Flash。该模型将驱动Siri的部分功能,其中简单查询预期会在本地设备运行。苹果面临的关键挑战是确保该大模型能够足够快速地响应日常问题。此外,下个月AI领域预计将有多项重要发布,包括WWDC上的Apple Intelligence与Gemini整合、GPT-5.6、可能的Sonnet 4.8/Opus 4.8,以及已确认的Gemini 3.5 Pro。
另有 7 家信源报道IT之家(RSS)TechCrunch:AI(RSS)Apple:Newsroom(RSS)公众号:数字生命卡兹克The Verge:AI(RSS)X:Kim (@kimmonismus)X:Testing Catalog (@testingcatalog)
推荐理由:Apple 把 1.2T 参数的定制 Gemini 塞进 Siri,简单查询还打算跑本地,这比单纯集成 Gemini 激进得多,WWDC 见真章。

5月25日

星期一 · 1 条
18:58
The Decoder:AI News(RSS)精选
AI 评分 72/100
Google DeepMind 的 AlphaProof Nexus 以几百美元的成本解决数十年未解的数学问题

Google DeepMind 的 AlphaProof Nexus 自主解决了 9 个开放的 Erdős 问题,其中包括两个困扰数学界 56 年的难题。其推理成本低至每个问题仅需几百美元。系统通过 Lean 编译器验证每个证明步骤,而非使用 OpenAI 的自然语言方法。当前的整体问题解决成功率为 2.5%。

另有 1 家信源报道IT之家(RSS)
推荐理由:AlphaProof Nexus 花几百美元就解决了数学家 56 年没做出来的问题,虽然成功率只有 2.5%,但这条路证明形式化验证+强化学习是走得通的,做推理的该盯着看了。