Microsoft is reporting 95.95% on CyberGym for its MDASH configuration.
CyberGym measures whether AI agents can reproduce real software vulnerabilities from code.
The next-best result shown is GPT-5.5 Cyber at 85.6%. Gemini 3.5, GPT-5.6 Sol, and Mythos 5 all sit around 83% to 84%.
MAI-Cyber-1-Flash is the AI model. MDASH is the larger agent-and-orchestration system that uses it. It coordinates more than 100 specialised agents and gives them different roles, tools, prompts and stopping rules.
If it really works, it can automate the slow, expensive work of finding and fixing hidden flaws in huge codebases.
The harness, security context, signals, and action space are kept separate from the model family.
In principle, that makes the underlying model replaceable. Microsoft can introduce a cybersecurity-specific model, combine it with another model where needed, and keep the surrounding investigation and remediation workflow intact.
One particular sentence in their official blog is particularly interesting.
"Microsoft sees more than 100 trillion security signals every day"