dex@dexhorthy
38AI 编辑部评分,满分 100
2026-08-08 00:25· 35分钟前
AI 导读

新一期播客深入解析华盛顿大学推出的 SlopCodeBench 基准,该基准要求模型随时间演化代码库,并衡量代码复杂度增长。目前最强模型(Fable、Sol、Kimi K3)通过率仅 33%。主持人 Dex Horthy 与 Vaib 探讨了基准机制及前沿模型测试发现。

🦄New Ep alert - If you've been following along, I'm obsessed with a new benchmark out of UW called SlopCodeBench - the best models in the world (Fable, Sol, Kimi K3) top out at a 33% pass rate.

What makes SCB cool is that it forces a model to evolve a codebase over time

AND it measures how that code grows in complexity as the model tries to solve progressively harder problems

I sat down with @vaibcode to go deep on how the benchmark works and my findings from the latest family of frontier models.

we do 🦄 AI that works live every tuesday at 10:15am PT - come hang for the next one!

来源:dex · x.com

dex · @dexhorthy · X·2026-08-08 00:25·35分钟前
AI 导读

新一期播客深入解析华盛顿大学推出的 SlopCodeBench 基准,该基准要求模型随时间演化代码库,并衡量代码复杂度增长。目前最强模型(Fable、Sol、Kimi K3)通过率仅 33%。主持人 Dex Horthy 与 Vaib 探讨了基准机制及前沿模型测试发现。

🦄New Ep alert - If you've been following along, I'm obsessed with a new benchmark out of UW called SlopCodeBench - the best models in the world (Fable, Sol, Kimi K3) top out at a 33% pass rate.

What makes SCB cool is that it forces a model to evolve a codebase over time

AND it measures how that code grows in complexity as the model tries to solve progressively harder problems

I sat down with @vaibcode to go deep on how the benchmark works and my findings from the latest family of frontier models.

we do 🦄 AI that works live every tuesday at 10:15am PT - come hang for the next one!

来源:dex· x.com