Ox Alpha 模型 DeepSWE 测试接近 GPT-5.6 Sol mid

Chubby♨️ · @kimmonismus · X·2026-08-23 06:35·1天前
AI 导读

神秘模型 Ox Alpha 在 DeepSWE 测试中取得约 63% 成绩,表现与 GPT-5.6 Sol mid 相当。若该模型确为 GLM-5.3 Flash,能在 DGX Spark 上本地运行,将极具性价比。用户评价其代码质量好、处理复杂任务能力强,但速度偏慢且偶尔遗留死代码。

Chubby♨️@kimmonismus
35AI 编辑部评分,满分 100

Ox Alpha 模型 DeepSWE 测试接近 GPT-5.6 Sol mid

2026-08-23 06:35· 1天前
AI 导读

神秘模型 Ox Alpha 在 DeepSWE 测试中取得约 63% 成绩,表现与 GPT-5.6 Sol mid 相当。若该模型确为 GLM-5.3 Flash,能在 DGX Spark 上本地运行,将极具性价比。用户评价其代码质量好、处理复杂任务能力强,但速度偏慢且偶尔遗留死代码。

Ox Alpha has now been fully tested against the DeepSWE set by @davis7 . Overall, it's performing more or less on par with GPT-5.6 Sol mid.

But he makes a good point. If it really is GLM-5.3 Flash that's now performing at the level of 5.6 Sol mid and could run locally on a DGX Spark, I wouldn't just be satisfied, it would be an absolutely fantastic deal.

5.6 Sol mid running locally, with the only cost being power consumption, for all sorts of tasks would be incredibly great. This model 24/7 hermes local: game changer.

Ben DavisActual DeepSWE run on the ox alpha mystery model is done. Ended at ~63% NOT the 80% my first subset test got, which makes way more sense. I've been using this t...

来源:Chubby♨️· x.com