Gemini 3.7 Flash official released. Surprisingly solid jump over 3.6!
FrontierCode: 43.6% vs. 34.4%
DeepSWE: 65.3% vs. 49.0%
AutomationBench: 30.4% vs. 17.0%
WebDev Arena: 1588 vs. 1538 Elo
The model is designed to plan more carefully, handle tool calls more reliably and require fewer retries. It will also power Gemini Spark, Google's 24/7 personal agent.
The benchmarks are Google's own launch claims, but the combination is aggressive: a stronger model, shipped after just three weeks, at $0.75 per million input tokens and $3.75 per million output tokens through year-end.
Not bad, google!