精选归档 · 第 8 页
第 141–160 条 · 共 296 条
03:08
参与我们的 I/O 2026 测验:该测验由 Google AI Studio 氛围编程生成Google 使用其开发工具 Google AI Studio,通过氛围编程(vibe coding)方式,创建了一个关于 Google I/O 2026 主要公告的在线测验。
推荐理由:Google 用 AI Studio 自己 vibe code 了个 I/O 测验,是想展示普通人也玩得转,但 quiz 本身信息量不大,想体验 vibe coding 的可以顺手玩玩。
01:38
Gemini Omni 与 Gemini 3.5 的 11 个实战展示Google 在 2026 年 Google I/O 大会上发布了新一代多模态模型 Gemini Omni 与 Gemini 3.5,并同步提供了 11 个视频,集中演示了这两款模型在实际场景中的能力。
另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Google 官方放出的这组视频演示,直接展示了 Gemini Omni 和 3.5 的实际表现,比参数和 benchmark 更直观,做多模态应用的可以逐帧研究。
15:21
谷歌 DeepMind CEO 哈萨比斯:AGI 最快三年内到来,研发速度远超预期谷歌 DeepMind 首席执行官德米斯·哈萨比斯预测,AGI 研发速度远超预期,最快可能在 2029 年至 2030 年前后出现。作为 AlphaGo、AlphaFold 的主导者,他认为当前 AI 智能体是未来更强智能的预演,随着多模态和自主决策能力成熟,三年内迎来 AGI 关键突破已非科幻。但他同时警示,全球社会对 AGI 到来的准备严重不足,必须提前建立规则与防护机制。
推荐理由:哈萨比斯作为造出 AlphaFold 的诺贝尔奖得主,三年内 AGI 的判断不是空话,他同时强调社会完全没准备好,这种紧迫感比单纯的时间表更值得看。
05:12
Google Research:Blog(网页)精选
创新时代:Google Research 在 I/O 2026Google Research 在 I/O 2026 大会上展示了其在多个前沿领域的技术进展,包括应用AI、基础机器学习算法以及量子AI等。本次大会的核心主题是展示其在将科学发现与研究成果转化为现实世界影响方面的持续努力。
另有 1 家信源报道Google Blog:AI(RSS)
推荐理由:Google 把研究成果直接发 Nature,ERA 和 Co-Scientist 这套工具让 AI 从写诗进化到做实验,健康 AI 的临床验证数据也很扎实,搞科研的可以蹲一下访问资格。
02:41
Google Developers Blog(RSS)精选
使用 Google Pay & Wallet Developer MCP server 加速你的集成工作流Google 推出 Google Pay & Wallet Developer MCP server,这是一款开放标准工具,旨在将 AI 开发助手和 IDE 安全连接到实时的 API 与账户上下文。开发者无需离开开发环境,即可搜索官方文档、验证 Wallet pass 定义、检查集成状态以及管理商户账户。该集成旨在通过减少上下文切换并提供实时、可靠的 AI 支持来减少开发摩擦,从而加速开发工作流。
推荐理由:这是 Google 为支付场景做的 MCP 服务器,把文档和账户操作直接塞进 IDE,减少上下文切换,做 Google Pay 集成的开发者可以试试看。
23:41
Google Developers Blog(RSS)精选
社区如何利用Tunix和TPU训练Gemma学会"思考"Google在Kaggle举办的Tunix黑客马拉松,挑战开发者利用TPU和有限算力,将小型基础模型转变为通用推理引擎。获胜团队通过多阶段后训练流程实现了这一目标,该流程结合了监督微调(SFT)与GRPO、SimPO等先进对齐技术。比赛结果表明,社区能够借助开源资源成功训练出高能力的结构化推理模型。
推荐理由:Google 官方比赛总结,证明用 Kaggle TPU 和开源工具就能把 Gemma 训练出不错推理能力,对想自己微调模型的小团队是个实用参考。
01:39
Google Developers Blog(RSS)精选
Google Pay 最新更新Google Pay 正向"智能体商务"演进,推出了通用商务协议和新的 MCP 服务器,允许 AI 智能体管理集成与分析趋势。Android 平台更新引入了动态回调以支持快速结账,并通过 WebView 将支付功能扩展至社交媒体应用。此外,平台还推出了跨设备生物认证和新的交易信号,旨在帮助商家减少流程摩擦。
推荐理由:Google Pay 往 agentic commerce 迈了一大步,新的通用协议和 MCP server 让 AI agent 能直接管支付和分析,做 agent 或支付的开发者都得看看。
01:34
Google Research:Blog(网页)精选
通过零信任聚合实现的隐私分析Google Research 推出了一种新的隐私分析解决方案。该方案结合了一种新的密码学安全聚合协议与可信执行环境(TEE)的透明性,旨在实现前沿的隐私与安全保证。其核心是基于零信任原则,通过密码学与硬件保护的结合,确保系统仅能获取群体的匿名化聚合洞察。
推荐理由:Google 的隐私聚合新方案把多轮交互砍成一次提交,对做设备端联邦分析的人来说是工程上的一大步,而且结合 TEE 做双层防护,这个思路值得抄。
18:20
HuggingFace Daily Papers(社区热门论文)精选
Gemini Embedding 2:来自Gemini的原生多模态嵌入模型Google DeepMind推出Gemini Embedding 2,这是一款原生多模态嵌入模型,支持在统一表示空间中嵌入视频、音频、图像和文本。该模型利用Gemini的多模态能力,通过大规模对比学习实现SOTA性能。在关键基准上表现优异:MSCOCO取得62.9 R@1,Vatex取得68.8 NDCG@10,MTEB multilingual达到69.9,MTEB Code达到84.0,超越了专用模型。其统一能力使其适用于RAG、推荐与搜索等下游任务,并在天文学、生物科学、艺术和烹饪等专业领域展现出强大的零样本性能。
推荐理由:Google 把多模态嵌入统一到一个模型里了,文本、代码、跨模态检索全面刷榜,做 RAG 和搜索的该认真看看了。
05:28
Google AI@GoogleAI精选 Gemini Omni 视频提示词使用指南http://x.com/i/article/2059377716965888000Mastering Gemini Omni: The Ultimate Video Prompting GuideLast week, we introduced Gemini Omni—our newest model designed to create anything from any input, starting with video.You can experience the speed and creativity of Gemini Omni Flash today across @geminiapp, @GoogleFlow, @GoogleFlowMusic, and on @YouTube Shorts and Create.To help you push the boundaries of what’s possible, here are five tips to get the most out of Gemini Omni’s advanced video generation capabilities.1. Leverage Real-World KnowledgeYou don’t need to over-explain the world to Gemini Omni. It’s built with Gemini’s deep understanding of history, science, and culture, so it can reliably create outputs that look, feel, and move realistically. Skip the granular descriptions. Use cultural touchstones, historical eras, or scientific terms directly in your prompt.Example Prompts:• [The video shows items of the alphabet. An unusual item starting with each letter is shown sitting on a table (like a Capybara for C, disco globe for D and Lava Lamp for L). All 26 letters must be represented by 26 items with matching lower thirds displaying the letter. Only one item and lower third at a time. Each lower third must look like a black marker written on a slip of paper in the bottom left. Rapid fire, roughly 9 frames per item at 24FPS. Last frame is a slip of paper "THE END." The whole video is accompanied by calm smooth music]• [Astronaut's POV on Mars]• [A marble rolling fast on a chain reaction style track, continuous smooth shot]2. Take Control of Text RenderingGemini Omni not only has advanced text rendering capabilities, it even allows you seamlessly integrate text into your visuals. You can specify typography, spatial placement, animation styles, and complex visual effects like double exposures all perfectly synced to the action in your video.Example Prompts:• [word by word, one word on the screen at a time: did, you, know, that, this, model, can, do, pretty, good, text!? Each word appears with a different animated style, perfect pacing to a rhythm, sizzle reel]• [Overlay motion-tracked, minimalist text commentary onto the physical environment of the video. This text represents [the subject] deadpan, immediate inner monologue that’s observant, slightly absurd, and life-contemplating. Think “intrusive thoughts.” Clean, white, lowercase sans-serif text (like Helvetica or Inter). The text hovers in 3D space, connected to the subjects being commented on via ultra-thin, crisp, white leader lines]3. Direct Your Camera Like a ProThink like a cinematographer. Gemini Omni responds incredibly well to precise videography directions, camera types, and framing instructions. Try integrating these terms into your next prompt:Example prompts:• Shots & Angles: "One continuous shot", "oner", "static", "locked off", or "fixed angle."• Camera Movements: "Push in", "punch in", "pan left", or "dolly zoom."• Camera Styles: "Natural smartphone zoom", "vintage film camera", or "grainy webcam style."4. Edit Iteratively (and keep what works)Every great video is made in the edit. With Gemini Omni, you don't need to rewrite your entire prompt from scratch to fix a single mistake. Ask for specific, targeted updates, like changing a background or swapping a caption. Omni will preserve the core structure of your video across multiple amends, letting you focus only on what needs tweaking.Example prompts:• [Transport the violin to a new environment]• [Make the violin invisible]• [Change the camera angle so it’s looking over the violinist’s shoulder]5. Change the Action on the FlyWant to alter a character's pacing or emotion mid-scene? You can directly prompt Gemini Omni to modify how a subject moves or interacts with their environment without breaking the continuity of the character model.Example prompts:• [Make the character walk on their tiptoes]• [Speed up the pacing]• [Have them leap into the air]Start CreatingThe director’s chair is yours. Try out these prompting techniques with Gemini Omni Flash, and tag @GoogleAI to show us what you create!译Google 发布了其多模态模型 Gemini Omni 的视频生成功能使用指南。该模型可通过 Gemini 应用、Google Flow 等平台体验。指南包含五项提示词技巧:利用模型已有的现实世界知识进行简洁描述;精确控制文本在视频中的渲染与排版;使用专业镜头指令(如推拉摇移)像电影摄影师一样调度画面;通过迭代编辑高效修改视频;以及在生成中直接调整角色的动作节奏或情绪。其核心在于通过精准的提示词引导模型生成复杂且可控的视频内容。另有 17 家信源报道IT之家(RSS)Google DeepMind:Blog(RSS)Google Developers Blog(RSS)Google Blog:AI(RSS)X:Gemini (@GeminiApp)X:Google AI (@GoogleAI)X:Jeff Dean (@JeffDean)X:Logan Kilpatrick (@OfficialLoganK)Hacker News 热门(buzzing.cc 中文翻译)X:Google DeepMind (@GoogleDeepMind)X:阿易 AI Notes (@AYi_AInotes)The Verge:AI(RSS)X:Ethan Mollick (@emollick)X:Sundar Pichai (@sundarpichai)X:Google AI for Developers (@googleaidevs)X:Rohan Paul (@rohanpaul_ai)The Decoder:AI News(RSS)
推荐理由:Google 官方放出的视频提示技巧,没有废话全是可复制的 prompt,想玩 Gemini Omni 的创作者可以直接抄作业。
18:58
The Decoder:AI News(RSS)精选
Google DeepMind 的 AlphaProof Nexus 以几百美元的成本解决数十年未解的数学问题Google DeepMind 的 AlphaProof Nexus 自主解决了 9 个开放的 Erdős 问题,其中包括两个困扰数学界 56 年的难题。其推理成本低至每个问题仅需几百美元。系统通过 Lean 编译器验证每个证明步骤,而非使用 OpenAI 的自然语言方法。当前的整体问题解决成功率为 2.5%。
另有 1 家信源报道IT之家(RSS)
推荐理由:AlphaProof Nexus 花几百美元就解决了数学家 56 年没做出来的问题,虽然成功率只有 2.5%,但这条路证明形式化验证+强化学习是走得通的,做推理的该盯着看了。