SemiAnalysis · @SemiAnalysis_ · X·2026-09-08 08:00·38分钟前
SemiAnalysis@SemiAnalysis_
45AI 编辑部评分,满分 100
2026-09-08 08:00· 38分钟前

Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse. (1/5)🧵

来源:SemiAnalysis· x.com