Step 3.7 Flash hits #2 on Claw-Eval General for autonomous agents.
We’re seeing strong performance across multi-step execution and robustness in long-horizon tasks, ranking just behind Claude Opus 4.6.
Promising signals for real-world agent workloads.