Claude Sonnet 5 upgrades are not uniform across every skill. e.g. its weaker than Sonnet 4.6 on CyberGym 🤔
Here, CyberGym is testing vulnerability discovery and exploit-finding behavior, not general reasoning or normal coding.
Anthropic also explicitly said in its announcment blog that Sonnet 5 was not deliberately trained for cyber tasks, so its cyber ability likely comes from general intelligence rather than targeted optimization.
So Sonnet 5's performance on CyberGym comes from general reasoning rather than specialized exploit skill.
---
From System Card of Claude Sonnet 5