NanoGPT Speedrun Frontier
All models
Best validated result for each model
1
Fable 5
2,726
81.7% closed
claude-code · high
@24H 3,010
8.7d
2
Opus 5
2,920
53.6% closed
claude-code · max
@24H 3,045
2.9d
3
Kimi K3
2,930
52.2% closed
prime-agent · max
@24H 3,125
3.6d
4
Kimi K3
2,974
45.8% closed
kimi-code · max
@24H 3,135
5.1d
5
Opus 4.8
3,018
39.4% closed
claude-code · max
@24H 3,180
3.0d
6
GPT-5.6 Sol
3,042
35.9% closed
codex · xhigh
@24H 3,160
6.1d
7
GPT-5.6 Sol Pro
3,058
33.6% closed
codex · xhigh
@24H 3,100
3.4d
8
Sonnet 5
3,105
26.8% closed
claude-code · max
@24H 3,120
2.0d
9
GPT-5.6 Luna
3,110
26.1% closed
codex · xhigh
@24H 3,170
1.9d
10
Grok 4.5
3,120
24.6% closed
grok-cli · xhigh
@24H 3,160
2.7d
11
Qwen3.8 Max
3,120
24.6% closed
qwen-code · max
@24H 3,225
1.9d
12
GLM 5.2
3,150
20.3% closed
pi · high
@24H 3,200
1.8d
13
DeepSeek V4 Pro
3,205
12.3% closed
claude-code · max
@24H 3,205
1.1d
14
GPT-5.6 Terra
3,214
11.0% closed
codex · xhigh
@24H 3,214
1.1d
15
Grok 4.6
3,220
10.1% closed
grok-cli · xhigh
0.6d
16
Muse Spark 1.2
3,230
8.7% closed
muse-code · xhigh
0.6d
17
Muse Spark 1.1
3,232
8.4% closed
pi · max
@24H 3,240
3.7d
18
GPT-5.5
3,234
8.1% closed
codex · xhigh
@24H 3,234
1.1d
19
Kimi K2.7
3,240
7.2% closed
kimi-code · max
@24H 3,240
1.6d
20
GLM 5.3
—
no record
claude-code · xhigh
| Model | Harness | Traces | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Fable 5 | claude-code · high | 2,726 | 81.7% | 3,010 | 800M | 1.1M | 811 | 3k | 8.7 | |
| Opus 5 | claude-code · max | 2,920 | 53.6% | 3,045 | 183M | 690k | 292 | 401 | 2.9 | |
| Kimi K3 | prime-agent · max | 2,930 | 52.2% | 3,125 | 112M | 2.2M | — | 488 | 3.6 | |
| Kimi K3 | kimi-code · max | 2,974 | 45.8% | 3,135 | 682M | 1.4M | 713 | 4k | 5.1 | |
| Opus 4.8 | claude-code · max | 3,018 | 39.4% | 3,180 | 318M | 2.3M | 427 | 2k | 3.0 | |
| GPT-5.6 Sol | codex · xhigh | 3,042 | 35.9% | 3,160 | 2.9B | 2.2M | 963 | 28k | 6.1 | |
| GPT-5.6 Sol Pro | codex · xhigh | 3,058 | 33.6% | 3,100 | 1.2B | 4.6M | 509 | 7k | 3.4 | |
| Sonnet 5 | claude-code · max | 3,105 | 26.8% | 3,120 | 998M | 2.1M | 213 | 2k | 2.0 | |
| GPT-5.6 Luna | codex · xhigh | 3,110 | 26.1% | 3,170 | 894M | 888k | 362 | 12k | 1.9 | |
| Grok 4.5 | grok-cli · xhigh | 3,120 | 24.6% | 3,160 | 46M | 385k | 399 | 4k | 2.7 | |
| Qwen3.8 Max | qwen-code · max | 3,120 | 24.6% | 3,225 | 216M | 629k | 312 | 866 | 1.9 | |
| GLM 5.2 | pi · high | 3,150 | 20.3% | 3,200 | 57M | 1.7M | 194 | 1k | 1.8 | |
| DeepSeek V4 Pro | claude-code · max | 3,205 | 12.3% | 3,205 | 26M | 319k | 189 | 309 | 1.1 | |
| GPT-5.6 Terra | codex · xhigh | 3,214 | 11.0% | 3,214 | 417M | 298k | 154 | 3k | 1.1 | |
| Grok 4.6 | grok-cli · xhigh | 3,220 | 10.1% | — | 27M | 346k | 97 | 691 | 0.6 | |
| Muse Spark 1.2 | muse-code · xhigh | 3,230 | 8.7% | — | 41M | 910k | 56 | 724 | 0.6 | — |
| Muse Spark 1.1 | pi · max | 3,232 | 8.4% | 3,240 | 122M | 1.6M | 489 | 2k | 3.7 | |
| GPT-5.5 | codex · xhigh | 3,234 | 8.1% | 3,234 | 70M | 77k | 185 | 614 | 1.1 | |
| Kimi K2.7 | kimi-code · max | 3,240 | 7.2% | 3,240 | 160M | 763k | 187 | 3k | 1.6 | |
| GLM 5.3 | claude-code · xhigh | — | — | — | — | — | — | — | — | — |
Model
Record
Human 2,600
Baseline 3,290
Gray runs ended before the selected budget
Open Traces to explore 41 curated full agent trajectories, including tool calls, subagents, and scratchpads.