GPT-5.6 模型家族现已登陆软件开发智能体 Kiro,包含 Sol、Terra 和 Luna 三款模型。在 Terminal-Bench 2.1 测试中,GPT-5.6 Terra 在 Kiro 内完成任务成本降低约 82%。该更新由 OpenAI 与 AWS 合作优化,旨在以更少迭代和更高 token 价值帮助开发者产出更高质量的代码。
GPT-5.6 模型家族现已登陆软件开发智能体 Kiro,包含 Sol、Terra 和 Luna 三款模型。在 Terminal-Bench 2.1 测试中,GPT-5.6 Terra 在 Kiro 内完成任务成本降低约 82%。该更新由 OpenAI 与 AWS 合作优化,旨在以更少迭代和更高 token 价值帮助开发者产出更高质量的代码。
Ox Alpha is currently the most popular model on @OpenRouter and is free to use! A reasoning model designed for coding, s...
Qwen's newest model is in the Arena under the name "paloma". Its frontend coding capabilities feel like they might be on...
名为 Ox Alpha 的神秘新 AI 模型在 OpenRouter 上线,被描述为面向编程、持续智能体工作和生产负载的推理模型,Stripe CEO Patrick Collison 称其“非常出色”。该模型由匿名第三方提供,引发关于其背后开发者的广泛猜测,焦点集中在中国的 Z.ai(GLM 模型)或微软未发布的 MAI 版本上。
OpenRouter 于 8 月 20 日上线匿名 AI 模型 Ox Alpha,免费开放 1 周。该模型在 DeepSWE 十项测试任务中得分约 80%,超过 Claude Fable 5 的 65% 和 GPT-5.6 Sol 的 52%。其上下文窗口达 1,048,576 个 token,支持文本、图像和视频多模态输入,分词器指纹与 GLM-5.3 高度吻合,线索指向智谱。
SITUATION EXPLAINED: A stealth model on OpenRouter is beating Fable and Sol on coding, and nobody knows who made it. • O...
Ox Alpha 在一次提示下生成 64,745 个输出 token,单次响应产出完整 Three.js dreamcore 3D 场景的全部程序化代码,无需外部素材。该模型支持 1M 上下文窗口、约 131k 最大输出 token,可处理文本+图像+视频输入并输出文本,定位为面向编码、长程智能体软件工程及生产负载的推理模型。aimlapi 正免费提供 Ox Alpha。
NEW OX ALPHA FEELS CRAZY 🤯 We asked Ox Alpha to build a dreamcore 3D world in three.js — one prompt, one shot. No asset...
A new Kimi model, likely K3.1, is now being tested on the Code @arena under the name "korrine" K3 was tested on the Aren...
Ox Alpha 是一款面向编码、持续性智能体工作和生产负载的推理模型,现已免费开放使用。该模型支持 1,048,576 token 的上下文窗口,最大输出 131,072 token。
Ornith-1.5 发布,将 Ornith-1.0 的自我构建框架扩展为完整自我优化闭环:模型自主提出任务、生成任务专属脚手架并产出解决方案用于强化学习。
另有 2 家信源报道X:Rohan Paul (@rohanpaul_ai)IT之家(RSS)Aloha! 🌺Introducing Ornith-1.5, a family of open-source LLMs spanning 9B Dense, 35B MoE, and 397B MoE, trained with sel...
Grok 4.6 旗舰模型现已上线 Amazon Bedrock,支持 500K-token 上下文窗口、文本与图像输入,以及 Low 至 xHigh 四档可配置推理强度。该模型面向长时运行智能体、编程与知识工作,提供客户端工具调用、结构化输出及流式响应。定价为全球区域每百万 token 输入 $2、输出 $6、缓存读取 $0.50,覆盖所有提供 Bedrock 的 AWS 区域。
智谱AI发布新基础模型GLM-5.3的API,聚焦复杂编程、网络安全防御与长周期任务。该模型在Artificial Analysis Intelligence Index上得分60,进入全球前沿模型梯队,与Anthropic Claude Fable 5、OpenAI GPT-5.6 Sol等闭源模型并列,并与月之暗面Kimi K3持平。
同一事件,精选展示《GLM-5.3上线:AA智能指数60分并列开源第一,成本更低》GLM-5.3 API is now live. - Built for coding, defensive cybersecurity, and long-horizon agentic tasks - Priced the same a...
Qwen3.8-27B 开源模型发布,27B 参数即可在单张 24GB 消费级显卡本地运行,编码与 Agent 能力超越 Claude Opus 4.6(SWE-bench Pro 61.7 vs 53.4,OSWorld 84.3 vs 72.7)。
We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With ...
同一事件,精选展示《通义千问开源 Qwen3.8 系列模型》Google AI 发布本周回顾:Gemini 3.7 Flash 作为最强编码与智能体模型,已上线 Gemini API 及 Google Workspace 等平台。
New on DeepInfra: Qwen3.8-2.4T-A95B 🚀 @Alibaba_Qwen's latest sparse MoE — 2.4T total params, 95B active, 512 experts. B...
🚀 Day-0 Support! @Alibaba_Qwen has open-sourced Qwen3.8-2.4T-A95B — and it’s now live on SiliconFlow. ⚡ With 2.4T param...
另有 4 家信源报道X:Rohan Paul (@rohanpaul_ai)IT之家(RSS)X:硅基流动 SiliconFlow (@SiliconFlowAI)X:Clément Delangue(Hugging Face CEO) (@ClementDelangue)阿里巴巴发布Qwen3.8-27B,一款27B参数的开源多模态稠密模型,采用Apache 2.0许可,支持Transformers、vLLM、SGLang及本地量化部署。
We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With ...
同一事件,精选展示《通义千问开源 Qwen3.8 系列模型》Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous age...
Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous age...
Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B activ...
Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous age...
Now available: @Alibaba_Qwen 3.8-2.4T-A95B from @alibaba_cloud on DigitalOcean Serverless Inference via NVIDIA HGX™ B300...
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved throug...
通义千问发布 Qwen3.8-27B,一款仅 27B 参数的原生多模态密集模型,综合表现超越 Qwen3.7-Plus,在真实编码与办公工作流中尤为突出。该模型原生支持 262K 上下文,可通过 YaRN 轻松扩展至 1M token,采用 Apache 2.0 许可证开源。此外,Qwen3.8-2.4T-A95B(Max 级)的开放权重也已于近期发布。
同一事件,精选展示《通义千问开源 Qwen3.8 系列模型》Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved throug...
Qwen3.8-2.4T-A95B is now live on Fireworks with Day-0 support! This 2.4T parameter MoE model is built for autonomous age...
阿里通义千问发布 Qwen3.8-27B,仅 27B 参数的开源多模态模型,宣称在智能体编程、计算机使用、浏览器任务及长周期专业工作中达到 SOTA,多项基准超越 Claude Opus 4.6 Max。支持图像、视频、可配置推理,原生 262K 上下文窗口可扩展至 1M,Apache 2.0 协议可自托管。基准结果为官方自测,尚待独立验证。
同一事件,精选展示《通义千问开源 Qwen3.8 系列模型》智谱发布 GLM-5.3,基于 743B 基座模型后训练,在漏洞发现等网络防御测试中超越 Anthropic Mythos 5。其 Terminal Bench 3.0 从 4.6 升至 28.3,DeepSWE 达 66.9;CyberGym 得分 84.5%,但 ExploitBench 仅 54.4%,落后于 Mythos 5 的 78.0%。
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved throug...
同一事件,精选展示《GLM-5.3 发布:编程能力开源第一,并涌现网络安全能力》智谱发布 GLM-5.3,与上一代 GLM-5.2 共享相同基座,全部提升来自扩展的后训练,并称其为最强开源权重编程模型,智能体任务进步最大。该模型经安全数据训练,在 269 个项目中找到 2,436 个漏洞,部分漏洞存在长达 40 年。GLM-5.3 现可通过 GLM Coding Plan 使用,权重将在安全审查完成后于两周内开源。
同一事件,精选展示《GLM-5.3 发布:编程能力开源第一,并涌现网络安全能力》