Artificial Analysis@ArtificialAnlys
68AI 编辑部评分,满分 100
2026-08-08 07:57· 17分钟前
AI 导读

蚂蚁集团发布Ling 3.0 Flash,124B总参数、5B激活的开源推理模型,在Artificial Analysis智能指数上得分38,较上代Ling 2.6 Flash提升24分,并以更少激活参数追平MiMo-V2.5和Qwen3.6 27B。该模型位于开源权重模型智能-总参数帕累托前沿,262K上下文窗口,API定价每百万输入token $0.075、输出$0.22,采用MIT许可。

Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models

@AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52.

Key results:

➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters.

➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545)

➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer

➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it.

Additional model details:

➤ Size: 124B total parameters, 5B active

➤ Context window: 262K

➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount

➤ License: MIT

➤ Providers: inclusionAI first-party API and DeepInfra third-party API

来源:Artificial Analysis · x.com

Artificial Analysis · @ArtificialAnlys · X·2026-08-08 07:57·17分钟前
AI 导读

蚂蚁集团发布Ling 3.0 Flash,124B总参数、5B激活的开源推理模型,在Artificial Analysis智能指数上得分38,较上代Ling 2.6 Flash提升24分,并以更少激活参数追平MiMo-V2.5和Qwen3.6 27B。该模型位于开源权重模型智能-总参数帕累托前沿,262K上下文窗口,API定价每百万输入token $0.075、输出$0.22,采用MIT许可。

Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models

@AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52.

Key results:

➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters.

➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545)

➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer

➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it.

Additional model details:

➤ Size: 124B total parameters, 5B active

➤ Context window: 262K

➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount

➤ License: MIT

➤ Providers: inclusionAI first-party API and DeepInfra third-party API

来源:Artificial Analysis· x.com