Artificial Analysis@ArtificialAnlys
67AI 编辑部评分,满分 100

韩国Upstage发布Solar Pro 4,智能指数42分

2026-08-13 01:20· 47分钟前
AI 导读

韩国AI实验室Upstage发布旗舰推理模型Solar Pro 4,Artificial Analysis智能指数达42分,较Solar Pro 3的14分大幅提升27分,与Inkling(xhigh,42分)持平。

Korean AI lab Upstage has released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, a significant increase from Solar Pro 3's 14

Solar Pro 4 is @upstageai's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026. At 42 on the Intelligence Index it sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43), and shows a 27-point increase over Solar Pro 3. Pricing increases to $0.30/$1.20/$0.06 per 1M input/output/cache hit tokens from Solar Pro 3's $0.15/$0.60/$0.02 via Upstage's first-party API.

Key results:

➤ Solar Pro 4's largest improvements on Solar Pro 3 are on agentic and long context work. Terminal-Bench v2.1 improves from 12% to 57%, AA-LCR from 31% to 71%, and τ3-Banking from 9% to 23%. GDPval-AA v2 shows strong progress on real-world agentic tasks, where Solar Pro 3 scored an Elo of 498, well below the human baseline of 1000, Solar Pro 4 scores 1276.

➤ AA-Omniscience improvement from -53 to -1 was driven by abstaining on more questions. Solar Pro 4 attempts only 41% of questions against 92% for Solar Pro 3, and its hallucination rate is 24%, higher than Command A+ (14%) and MiniMax-M3 (18%), and a vast improvement from Solar Pro 3's 88%. AA-Omniscience Accuracy remains unchanged at 19%.

➤ Solar Pro 4 is more token efficient than Solar Pro 3, though still verbose for its intelligence level. It uses 43k output tokens per Intelligence Index task, around 17% fewer than Solar Pro 3's 52k.

➤ The intelligence gain comes with a hit to latency. Solar Pro 4 takes 9.5 minutes to complete an average Intelligence Index task, against 6.9 minutes for Solar Pro 3, despite using fewer output tokens per task.

➤ Pricing is $0.30/$1.20 per 1M input/output tokens. This is in line with @MiniMax_AI's first-party pricing for MiniMax-M3, which scores 3 points higher at 45, and is more expensive than @deepseek_ai's DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at $0.14/$0.28 and 52 on the Intelligence Index. Cache hits are priced at $0.06 per 1M, an 80% discount on input token price.

Additional model details:

➤ Context window: 512K tokens

➤ Max output tokens: 128K

➤ Modalities: Text input and output only

➤ Pricing: $0.30 / $1.20 / $0.06 per 1M input/output/cache hit tokens

➤ Inference providers at time of launch: Upstage first-party API, OpenRouter, TimelyRouter

来源:Artificial Analysis · x.com

韩国Upstage发布Solar Pro 4,智能指数42分

Artificial Analysis · @ArtificialAnlys · X·2026-08-13 01:20·47分钟前
AI 导读

韩国AI实验室Upstage发布旗舰推理模型Solar Pro 4,Artificial Analysis智能指数达42分,较Solar Pro 3的14分大幅提升27分,与Inkling(xhigh,42分)持平。

Korean AI lab Upstage has released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, a significant increase from Solar Pro 3's 14

Solar Pro 4 is @upstageai's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026. At 42 on the Intelligence Index it sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43), and shows a 27-point increase over Solar Pro 3. Pricing increases to $0.30/$1.20/$0.06 per 1M input/output/cache hit tokens from Solar Pro 3's $0.15/$0.60/$0.02 via Upstage's first-party API.

Key results:

➤ Solar Pro 4's largest improvements on Solar Pro 3 are on agentic and long context work. Terminal-Bench v2.1 improves from 12% to 57%, AA-LCR from 31% to 71%, and τ3-Banking from 9% to 23%. GDPval-AA v2 shows strong progress on real-world agentic tasks, where Solar Pro 3 scored an Elo of 498, well below the human baseline of 1000, Solar Pro 4 scores 1276.

➤ AA-Omniscience improvement from -53 to -1 was driven by abstaining on more questions. Solar Pro 4 attempts only 41% of questions against 92% for Solar Pro 3, and its hallucination rate is 24%, higher than Command A+ (14%) and MiniMax-M3 (18%), and a vast improvement from Solar Pro 3's 88%. AA-Omniscience Accuracy remains unchanged at 19%.

➤ Solar Pro 4 is more token efficient than Solar Pro 3, though still verbose for its intelligence level. It uses 43k output tokens per Intelligence Index task, around 17% fewer than Solar Pro 3's 52k.

➤ The intelligence gain comes with a hit to latency. Solar Pro 4 takes 9.5 minutes to complete an average Intelligence Index task, against 6.9 minutes for Solar Pro 3, despite using fewer output tokens per task.

➤ Pricing is $0.30/$1.20 per 1M input/output tokens. This is in line with @MiniMax_AI's first-party pricing for MiniMax-M3, which scores 3 points higher at 45, and is more expensive than @deepseek_ai's DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at $0.14/$0.28 and 52 on the Intelligence Index. Cache hits are priced at $0.06 per 1M, an 80% discount on input token price.

Additional model details:

➤ Context window: 512K tokens

➤ Max output tokens: 128K

➤ Modalities: Text input and output only

➤ Pricing: $0.30 / $1.20 / $0.06 per 1M input/output/cache hit tokens

➤ Inference providers at time of launch: Upstage first-party API, OpenRouter, TimelyRouter

来源:Artificial Analysis· x.com