Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text ...
Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text ...
OpenRouter 于 8 月 20 日上线匿名 AI 模型 Ox Alpha,免费开放 1 周。该模型在 DeepSWE 十项测试任务中得分约 80%,超过 Claude Fable 5 的 65% 和 GPT-5.6 Sol 的 52%。其上下文窗口达 1,048,576 个 token,支持文本、图像和视频多模态输入,分词器指纹与 GLM-5.3 高度吻合,线索指向智谱。
阿里发布第三代图像模型 Qwen-Image-3.0 系列,旗舰版 Pro 在 Artificial Analysis 图像编辑榜位列第六、文生图榜第九,较上代分别提升 83 和 48 Elo 分。该系列支持 4.5k token 提示词、10 像素级精确文字渲染及 12 种语言,Pro 版定价 $0.04/张(1K 分辨率),已在阿里云 Model Studio 上线。
I’m happy to report that Grok 4.6 is just as good as Kimi K3 for everything related to the kernel and modding. I was tes...
Google 发布 Gemini 3.7 Flash,在 Artificial Analysis 智能指数上得分 56,较 3.6 Flash 提升 4 分,并登上智能 vs. 单任务耗时帕累托前沿。
同一事件,精选展示《Google DeepMind 推出 Gemini 3.7 Flash:面向编程与智能体的最强工作模型》韩国AI实验室Upstage发布旗舰推理模型Solar Pro 4,Artificial Analysis智能指数达42分,较Solar Pro 3的14分大幅提升27分,与Inkling(xhigh,42分)持平。
另有 1 家信源报道IT之家(RSS)Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price...
同一事件,精选展示《xAI 发布 Grok 4.6,强化长时运行智能体能力》Grok 4.6 is now out 🚀🚀🚀 Smart, fast & amazing bang for buck!
微软发布面向 GitHub Copilot 的代码模型 MAI Code 1.1 Flash,称其代码质量更优、token 效率提升 25%、成本仅为 6 月前代的四分之一。
另有 1 家信源报道IT之家(RSS)蚂蚁集团发布开源推理模型 Ling 3.0 Tiny,在 Artificial Analysis 智能指数上得 25 分,以 7.9B 总参数和 1.3B 激活参数(MoE)扩展了智能与激活参数的帕累托前沿,性能接近 gpt-oss-120b(24 分)但参数少 15 倍。
同一事件,精选展示《Ling-3.0-tiny 正式开源:1.3B 激活参数如何进入真实任务》Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4 in the Text Arena (1498 pts), and has reshaped the Pareto front...
We've got a change at the top. Meta's muse-spark 1.2 reaches 2nd place in the sidequest-bench.
Kimi K3 以 57 分继续领跑开源权重模型,仅领先 Qwen3.8 Max 一分,后者计划下周开源权重。Qwen3.8 Max 在 GDPval-AA 上以 1739 Elo 反超 Kimi K3(1685),且激活参数更小(95B vs 2.8T 总参数)。两者能力可视为持平,但 Kimi K3 单任务成本 $0.86,比 Qwen3.8 Max 的 $1.14 低 25%。
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader...
Muse Spark 1.2 just cracked the top 5 on the Vals Index, at just $0.69 per test. This is 3x cheaper than Kimi and 10x or...
One of the things the pelican benchmark is still useful for is visually representing (to a tiny extent) the improvements...
Meta 发布 Muse Spark 1.2,在 Artificial Analysis 智能指数中得分 54,较 1.1 版提升 3 分,与 Grok 4.5 并列美国实验室第三。
另有 2 家信源报道The Decoder:AI News(RSS)X:Artificial Analysis (@ArtificialAnlys)We ran a test between the new Qwen3.8-Max, Opus 5 and GPT-5.6 Sol. 3 models. same prompt. one-shot with the /design comm...
Celeris-1 is now ranked as the #1 fastest AI model according to @ArtificialAnlys. 2,038 output tokens per second, #1 of ...
Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with ...
DeepSeek V4-Flash isn’t just cheaper per token. It reportedly completes the same benchmark tasks as Fable 5 at 105× lowe...
While DeepSeek V4-Flash is significantly cheaper on price per token, this can be misleading if the overall cost per task...
Exciting news: DeepSeek-V4-Flash-High by @deepseek_ai has reshaped the Pareto Frontier in the Frontend Code Arena, with ...
DeepSeek-V4-Flash 正式版 API 上线公测,在 Artificial Analysis 智能指数(AII)中获 50 分,较 4 月版提升 10 分,比 V4 Pro 高 6 分,仅比 GPT-5.6 Luna 低 1 分。
同一事件,精选展示《DeepSeek V4 Flash 0731 开源,登顶开源模型前三》DeepSeek 发布开源模型 DeepSeek V4 Flash 0731,在 Artificial Analysis 智能指数上得分 50,位列开源模型前三。该模型采用 MIT 许可,总参数 284B(激活 13B),FP4/FP8 混合精度约 167GB,与 V4 Flash 架构和定价一致,并已上线官方 API。
另有 14 家信源报道X:DeepSeek (@deepseek_ai)X:阿易 AI Notes (@AYi_AInotes)MarkTechPost(RSS)公众号:数字生命卡兹克X:Emad Mostaque (@EMostaque)X:Rohan Paul (@rohanpaul_ai)Simon Willison 博客IT之家(RSS)公众号:卡尔的AI沃茨X:硅基流动 SiliconFlow (@SiliconFlowAI)X:X.PIN (@thexpin)X:Kim (@kimmonismus)X:Elvis Saravia (@omarsar0, DAIR.AI)Hacker News 热门(buzzing.cc 中文翻译)Mureka V9在Artificial Analysis音乐双榜均列第二,器乐榜距榜首Suno V5.5仅1 Elo分。该模型由Skywork AI推出,可生成2-4分钟全曲,支持多语言歌词及最多12音轨分离,现已在Mureka应用和API上线。
BREAKING: Grok 4.5 outperforms GPT-5.6 Terra across almost all shared benchmarks on AskClash, leading in ACB, GPQA, SWE-...
DeepSeek V4 Flash 0731 在 Artificial Analysis 智能指数上得 50 分,较 4 月版提升 10 分,领先 V4 Pro 6 分,并进入智能与成本帕累托前沿。
同一事件,精选展示《DeepSeek V4 Flash 0731 开源,登顶开源模型前三》This is insane! not kidding at all. DeepSeek v4 *Flash* super close to Opus 4.8, Insane upgrade - DeepSWE 54.4%, outperf...
同一事件,精选展示《DeepSeek V4 Flash 0731 开源,登顶开源模型前三》MiniMax H3在Artificial Analysis视频编辑排行榜位列第一,文生视频第二、图生视频第三。该模型支持多模态输入,可生成5-15秒24fps带原生音频的片段,2K定价$7.80/分钟。MiniMax计划以社区许可开源权重,允许年收入低于$20M的组织商用,若发布将成为最强开源视频模型。
另有 4 家信源报道X:MiniMax (@MiniMax_AI)MiniMax:Blog(网页)X:X.PIN (@thexpin)X:Testing Catalog (@testingcatalog)Kimi K3 weights have been released! Kimi K3 is now the leading open weights model at 57 in the Artificial Analysis Intel...
Anthropic 发布 Claude Opus 5,定位为接近 Claude Fable 5 前沿智能但价格减半的模型,目前在 Artificial Analysis 排行榜上领先所有模型。该模型定价与 Opus 4.8 相同,并提供成本翻倍的“快速模式”。因通用能力提升,Opus 5 在网络安全漏洞发现上接近 Mythos 5,但在漏洞利用上仍大幅落后。
另有 16 家信源报道X:Claude (@claudeai)X:Thariq (@trq212)X:Yuchen Jin (@Yuchenj_UW)Anthropic:Newsroom(网页)X:Kim (@kimmonismus)MarkTechPost(RSS)TechCrunch:AI(RSS)公众号:卡尔的AI沃茨X:Boris Cherny (@bcherny)X:Rohan Paul (@rohanpaul_ai)X:小北 (@frxiaobei)X:AI Safety Memes (@AISafetyMemes)The Verge:AI(RSS)X:Claude Devs (@ClaudeDevs)X:Testing Catalog (@testingcatalog)IT之家(RSS)Anthropic 发布 Claude Opus 5,已全量上线,支持 100 万上下文,知识截止至 2026 年 5 月。在 ARC-AGI-3 上得分 30.2%,接近 GPT-5.6 Sol(7.8%)的四倍;输入每百万 Token 5 美元、输出每百万 Token 25 美元,价格与 Opus 4.8 相同,仅为 Fable 5 的一半。
同一事件,精选展示《Anthropic 发布 Claude Opus 5》Anthropic 推出旗舰模型 Claude Opus 5,在智能体编程和知识工作基准测试中领先,并在 ARC-AGI-3 上取得 30.2% 的分数,是 GPT-5.6 Sol 的近四倍。
另有 16 家信源报道X:Claude (@claudeai)X:Thariq (@trq212)X:Yuchen Jin (@Yuchenj_UW)Anthropic:Newsroom(网页)X:Kim (@kimmonismus)MarkTechPost(RSS)TechCrunch:AI(RSS)公众号:卡尔的AI沃茨X:Boris Cherny (@bcherny)X:Rohan Paul (@rohanpaul_ai)X:小北 (@frxiaobei)X:AI Safety Memes (@AISafetyMemes)The Verge:AI(RSS)X:Claude Devs (@ClaudeDevs)X:Testing Catalog (@testingcatalog)IT之家(RSS)Google has released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Both halve time per task relative to their predecessors ...
月之暗面发布 Kimi K3,参数量达 2.8 万亿,支持百万上下文,将于 7 月 27 日全面开源。该模型在 AA 智能分数中排名第三,并在 BrowseComp、Automation Bench 和 SpreadsheetBench 2 三项 Agent 评测中取得第一。
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This ...