Gemini 3.8 Flash is out.
beats Claude Opus 5 on some really important benchmarks.
• 54.9% on HLE-Verified vs 54.4% for Claude Opus 5
• Terminal-Bench 2.1, that test tests whether an agent can actually operate a terminal and successfully finish difficult coding, security, ML, data-science, and systems tasks. 89.4% is essentially a very high task-completion rate under that evaluation setup.
• Harvey's Legal Agent Benchmark, Gemini 3.8 Flash scores 10.0% vs Opus 5's 6.7%. The number looks low because this uses an extremely strict all-pass rule: a legal workflow gets credit only when every required criterion passes, including facts, conclusions, citations, structure, and analysis, across complex file-based legal work.
Google has not raised the price per token. On harder tasks, Gemini 3.8 Flash may think longer and make more tool calls, so it uses more tokens and the total cost of completing that task can rise even though the token price stays the same.