Topic · 主题全部主题 →

多模态

文本之外的能力:视觉理解、图文混合、音视频输入输出的模型与产品进展。

3,705条收录
405条精选

精选归档 · 第 19 页

361380 条 · 共 405

4月3日

星期五 · 1 条
01:09
Artificial Analysis@ArtificialAnlys精选
Google发布Gemma 4多模态开源模型系列Google has released Gemma 4, a new family of multimodal open-weight models including Gemma 4 E2B, Gemma 4 E4B, Gemma 4 31B and Gemma 4 26B A4B@GoogleDeepMind’s new Gemma 4 family introduces four multimodal models supporting text, image, and video inputs. We evaluated Gemma 4 31B (dense) and Gemma 4 26B A4B (MoE), both with a 256k context window, while the other two smaller models support up to 128k. With 31B and 26B parameters respectively, both evaluated models can run on a single H100.On GPQA Diamond, our scientific reasoning evaluation, Gemma 4 31B (Reasoning) scores 85.7%, the second highest result we have recorded for an open-weights model with fewer than 40B parameters, just behind Qwen3.5 27B (Reasoning, 85.8%). It reaches this score using only ~1.2M output tokens, fewer than Qwen3.5 27B (~1.5M) and Qwen3.5 35B A3B (~1.6M). Gemma 4 26B A4B (Reasoning) scores 79.2%, ahead of gpt-oss-120B (high, 76.2%) but behind Qwen3.5 9B (Reasoning, 80.6%).We are now running the Artificial Analysis Intelligence Index on all four Gemma 4 models and will share a full update once those results are complete.Google DeepMind推出Gemma 4系列四款多模态开源模型,支持文本、图像及视频输入。31B(密集架构)与26B A4B(MoE架构)拥有256k上下文窗口,可在单张H100运行;另两款较小模型支持128k上下文。GPQA Diamond测试中,Gemma 4 31B(Reasoning)获85.7%,仅次于Qwen3.5 27B,但输出token仅约1.2M,效率更优;26B A4B(Reasoning)得分79.2%,超越gpt-oss-120B。
另有 2 家信源报道X:Artificial Analysis (@ArtificialAnlys)X:Jeff Dean (@JeffDean)
推荐理由:Google发布多模态开源模型Gemma 4,单卡H100可跑且科学推理能力突出

4月2日

星期四 · 3 条
08:00
Hugging Face:Blog(RSS)精选
AI 评分 88/100
Welcome Gemma 4: 设备端的 Frontier 多模态智能

Google 正式发布了 Gemma 4,这是一款前沿的多模态人工智能模型,其核心特点是能够在设备端本地运行。该模型通过开源方式发布,旨在推动人工智能技术的进步与民主化。Gemma 4 的“在设备端”能力意味着数据处理可在本地完成,无需持续连接云端,这有望提升响应速度、增强隐私保护并实现离线使用。此举是 Google 通过开源和开放科学来普及人工智能的持续努力的一部分。


推荐理由:前沿多模态模型开源,设备端可运行,降低AI部署门槛。
00:00
智谱:研究(网页内嵌数据)精选
GLM-5V-Turbo发布:多模态Coding基座模型

智谱发布GLM-5V-Turbo多模态Coding基座模型,原生支持图像、视频、设计稿理解及画框、截图、读网页等工具调用,上下文窗口达200k。采用新一代CogViT视觉编码器与30+任务协同强化学习,在保持纯文本编程能力的同时强化GUI Agent能力。与Claude Code、AutoClaw等框架深度协同,支持"图像即代码"前端复刻及GUI自主探索,提供开箱即用的官方Skills。


推荐理由:智谱发布多模态Coding基座GLM-5V-Turbo,深度适配Claude Code等Agent

3月31日

星期二 · 1 条
23:10
Hugging Face:Blog(RSS)精选
AI 评分 70/100
Granite 4.0 3B Vision:面向企业文档的紧凑型多模态智能

IBM Granite团队发布了Granite 4.0 3B Vision模型,这是一个专为企业文档处理设计的紧凑型多模态大语言模型。该模型参数为30亿,具备视觉理解能力,能够同时处理文本和图像信息,特别针对报告、表格、图表等企业文档进行优化。其紧凑尺寸旨在降低部署和运行成本,使企业能够在资源受限的环境中高效实现文档智能分析、信息提取和知识管理。模型已在Hugging Face平台发布。


推荐理由:IBM 推出轻量级多模态模型,企业文档场景可直接落地部署

3月30日

星期一 · 1 条
04:00
Qwen:Blog Retrieval(API)精选
Qwen3.5-Omni:全面扩展,迈向原生全模态 AGI

Qwen Studio 发布,集成聊天机器人、图像视频理解、图像生成、文档处理、网页搜索、工具使用及 Artifacts 功能,提供全模态 AI 一站式解决方案。

另有 1 家信源报道Qwen:Blog Retrieval(API)
推荐理由:阿里发布Qwen3.5-Omni多模态模型,迈向原生全模态AGI

3月29日

星期日 · 1 条
22:32
Gary Marcus:The Road to AI We Can Trust(RSS)精选
当前前沿模型视觉理解的幻象

当前前沿多模态大模型在标准胸部X光问答基准测试中,无需访问任何图像即可获得顶级排名。这一反常现象暴露出模型视觉理解能力的严重缺陷,表明其性能可能依赖数据偏见或文本线索而非真实的图像解析能力。研究揭示了现有视觉语言模型评估体系的深层漏洞,指出所谓"视觉理解"可能只是缺乏真实感知能力的幻觉。


推荐理由:揭示多模态基准测试漏洞,医学AI应用需警惕数据泄露风险

3月26日

星期四 · 1 条
00:00

3月25日

星期三 · 1 条
00:00
Google Research:Blog(网页)精选
Vibe Coding XR:基于 XR Blocks 与 Gemini 加速 AI + XR 原型开发

Google XR 团队推出 Vibe Coding XR 工作流,结合 Gemini Canvas 与开源框架 XR Blocks,利用长上下文推理能力将自然语言提示在 60 秒内转化为可交互、支持物理效果的 WebXR 应用。该方案基于 WebXR、three.js 和 LiteRT.js 构建,支持手势交互与深度感知,可在桌面模拟环境或 Android XR 头显中实时预览。已展示的应用包括几何可视化数学辅导和交互式物理实验室,用户可通过捏合等手势操作 3D 对象,快速验证空间交互设计。


推荐理由:Google推出Vibe Coding XR,用自然语言快速生成可交互的Android XR空间应用。

3月24日

星期二 · 1 条
08:00
Google Developers Blog(RSS)精选
AI 评分 71/100
跳跃即玩:利用Gemini与MediaPipe进行开发

该工作流通过Gemini Canvas,借助高级提示词快速原型化MediaPipe Pose Landmarker等体感游戏机制。开发者可在Google AI Studio中优化原型,采用低延迟的“轻量”模型和稳定的追踪点(如肩部关节点)以确保游戏响应灵敏。最后,流程利用Gemini Code Assist将实验性代码重构为模块化、可用于生产的应用程序,使其能够支持多种多模态输入,从而显著简化了体感控制游戏的开发过程。


推荐理由:开发者可快速上手AI游戏开发,优化性能并部署生产应用。

3月20日

星期五 · 1 条
19:48
Artificial Analysis@ArtificialAnlys精选
Mistral发布开源模型Small 4,支持混合推理与图像理解Mistral has released Mistral Small 4, an open weights model with hybrid reasoning and image input, scoring 27 on the Artificial Analysis Intelligence Index@MistralAI's Small 4 is a 119B mixture-of-experts model with 6.5B active parameters per token, supporting both reasoning and non-reasoning modes.In reasoning mode, Mistral Small 4 scores 27 on the Artificial Analysis Intelligence Index, a 12-point improvement from Small 3.2 (15) and now among the most intelligent models Mistral has released, surpassing Mistral Large 3 (23) and matching the proprietary Magistral Medium 1.2 (27). However, it lags open weights peers with similar total parameter counts such as gpt-oss-120B (high, 33), NVIDIA Nemotron 3 Super 120B A12B (Reasoning, 36), and Qwen3.5 122B A10B (Reasoning, 42).Key takeaways:➤ Reasoning and non-reasoning modes in a single model: Mistral Small 4 supports configurable hybrid reasoning with reasoning and non-reasoning modes, rather than the separate reasoning variants Mistral has released previously with their Magistral models. In reasoning mode, the model scores 27 on the Artificial Analysis Intelligence Index. In non-reasoning mode, the model scores 19, a 4-point improvement from its predecessor Mistral Small 3.2 (15)➤ More token efficient than peers of similar size: At ~52M output tokens, Mistral Small 4 (Reasoning) uses fewer tokens to run the Artificial Analysis Intelligence Index compared to reasoning models such as gpt-oss-120B (high, ~78M), NVIDIA Nemotron 3 Super 120B A12B (Reasoning, ~110M), and Qwen3.5 122B A10B (Reasoning, ~91M). In non-reasoning mode, the model uses ~4M output tokens➤ Native support for image input: Mistral Small 4 is a multimodal model, accepting image input as well as text. On our multimodal evaluation, MMMU-Pro, Mistral Small 4 (Reasoning) scores 57%, ahead of Mistral Large 3 (56%) but behind Qwen3.5 122B A10B (Reasoning, 75%). Neither gpt-oss-120B nor NVIDIA Nemotron 3 Super 120B A12B support image input. All models support text output only➤ Improvement in real-world agentic tasks: Mistral Small 4 scores an Elo of 871 on GDPval-AA, our evaluation based on OpenAI's GDPval dataset that tests models on real-world tasks across 44 occupations and 9 major industries, with models producing deliverables such as documents, spreadsheets, and diagrams in an agentic loop. This is more than double the Elo of Small 3.2 (339) and close to Mistral Large 3 (880), but behind gpt-oss-120B (high, 962), NVIDIA Nemotron 3 Super 120B A12B (Reasoning, 1021), and Qwen3.5 122B A10B (Reasoning, 1130)➤ Lower hallucination rate than peer models of similar size: Mistral Small 4 scores -30 on AA-Omniscience, our evaluation of knowledge reliability and hallucination, where scores range from -100 to 100 (higher is better) and a negative score indicates more incorrect than correct answers. Mistral Small 4 scores ahead of gpt-oss-120B (high, -50), Qwen3.5 122B A10B (Reasoning, -40), and NVIDIA Nemotron 3 Super 120B A12B (Reasoning, -42)Key model details:➤ Context window: 256K tokens (up from 128K on Small 3.2)➤ Pricing: $0.15/$0.6 per 1M input/output tokens➤ Availability: Mistral first-party API only. At native FP8 precision, Mistral Small 4's 119B parameters require ~119GB to self-host the weights (more than the 80GB of HBM3 memory on a single NVIDIA H100)➤ Modality: Image and text input with text output only➤ Licensing: Apache 2.0 licenseMistral发布开源权重模型Mistral Small 4,采用119B参数MoE架构(每token激活6.5B参数),支持可切换的推理/非推理模式及图像输入。推理模式在Artificial Analysis Intelligence Index获27分,超越Mistral Large 3,但低于gpt-oss-120B等竞品。模型token效率优于同类,幻觉率更低(AA-Omniscience -30分),支持256K上下文窗口,采用Apache 2.0许可证。

推荐理由:Mistral 开源 Small 4,支持混合推理与多模态,Agent 任务表现大幅提升

3月19日

星期四 · 2 条
11:12
Demis Hassabis@demishassabis精选
用 @stitchbygoogle 即可 vibe design 惊艳界面You can vibe design some incredible interfaces with @stitchbygoogleGoogle 发布 vibe design 平台 Stitch,支持自然语言描述直接生成高保真界面和交互原型,可通过语音实时调整布局。目前仅面向 18 岁以上用户,在 Gemini 支持的英语国家开放。

Google Labs: Introducing the new @stitchbygoogle, Google’s vibe design platform that transforms natural language into high-fidelity d...


推荐理由:Google推出AI设计工具Stitch,自然语言生成界面并支持语音协作,顺应Vibe Design趋势
04:00
Qwen:Blog Retrieval(API)精选
Qwen3.5-Max-Preview 现已上线 Arena

Qwen3.5-Max-Preview 已登陆 LMSYS Chatbot Arena。Qwen Studio 提供聊天机器人、图像与视频理解、图像生成、文档处理、网页搜索、工具调用及 artifacts 等全栈功能。


推荐理由:阿里 Qwen3.5-Max 预览版上线 Arena,支持多模态理解与工具调用

3月17日

星期二 · 1 条
20:33
Hugging Face:Blog(RSS)精选
AI 评分 83/100
Holotron-12B - 高吞吐计算机使用智能体

H公司发布了多模态计算机使用模型Holotron-12B。该模型基于NVIDIA开源的Nemotron-Nano-12B-VL模型,使用专有数据混合进行训练,专注于在交互环境中高效感知、决策和行动。其采用混合状态空间模型与注意力机制架构,在单张H100 GPU上实现了比前代Holo2-8B高2倍以上的吞吐量,在100并发基准测试中达到每秒8900个token。在WebVoyager基准测试中,性能从基线的35.1%提升至80.5%,在定位和导航基准上也显著提升。模型已通过NVIDIA开放模型许可在Hugging Face发布。


推荐理由:高效推理的计算机使用代理模型,适合生产部署,开发者可直接试用。

3月13日

星期五 · 1 条
08:00
Claude Platform:开发者版本说明(RSS)精选
AI 评分 69/100
Claude Opus 4.6 和 Sonnet 4.6 的 1M token 上下文窗口正式可用

Claude Opus 4.6 和 Sonnet 4.6 的 1M token 上下文窗口现已正式可用,按标准定价计费,超过 200k token 的请求无需 beta header 即可自动生效。同时,所有支持模型的专用 1M 速率限制已移除,改用标准账户限制;使用 1M 上下文窗口时,每请求的媒体上限从 100 提升至 600 张图片或 PDF 页面。


推荐理由:Claude 的百万 token 上下文终于从 beta 转正,200k token 以上请求不再需要 beta 头,媒体文件限制一口气提到 600 页,做长文档处理的开发者可以直接切生产。

3月12日

星期四 · 2 条
00:00
Claude:Blog(网页)精选
Claude 新增交互式图表、图解与可视化功能

Claude 推出可视化功能测试版,支持在对话中实时生成交互式图表、图解等视觉内容,无需代码即可随对话调整修改。该功能不同于可下载的 Artifacts,以内联临时形式辅助理解当前话题,默认向所有套餐用户开启。同时 Claude 还新增食谱、天气等主题格式,并支持在对话内直接交互 Figma、Canva 和 Slack 等应用。


推荐理由:Claude推出对话内交互式图表功能,实时生成可视化助力理解

3月10日

星期二 · 1 条
18:00
OpenAI:官网动态(RSS · 排除企业/客户案例)精选
ChatGPT 推出数学与科学学习新方式

ChatGPT 新增数学与科学交互式可视化解释功能,支持实时探索公式、变量及概念,帮助学生更直观地理解理科知识。


推荐理由:ChatGPT 新增数学与科学可视化交互功能,提升学习体验

3月9日

星期一 · 1 条
00:00
Runway:News(网页)精选
Runway 推出 Characters:单图实时生成可对话虚拟角色 API

Runway 推出 Characters API,基于 GWM-1 世界模型,支持用单张图片零微调生成实时可对话虚拟角色。支持自定义外观风格、声音、性格及知识库,具备自然表情、眼神、口型同步和手势。面向客户支持、培训教育和品牌营销等企业场景,已获 BBC 等采用。开发者可通过 API 集成,消费者也可在网页端体验预设角色。


推荐理由:Runway推出实时视频Agent,单图生成可对话数字人,拓展AI交互形态

3月4日

星期三 · 1 条
01:00