Microsoft adds MAI-Cyber-1-Flash to MDASH and reports a 95.95% CyberGym score

On July 28, 2026, Microsoft launched MAI-Cyber-1-Flash inside MDASH, its vulnerability identification and remediation harness. Paired with GPT-5.4, the system scored 95.95% on CyberGym, while Microsoft says it costs 50% less than its strongest existing MDASH model mix.
Why the system-level result matters
The headline score belongs to MDASH running two models, not to MAI-Cyber-1-Flash alone. Microsoft designed the smaller model to process up to 90% of tasks, reserving GPT-5.4 for the hardest 10%. This routing approach aims to reduce expensive model calls without sacrificing overall performance.
CyberGym Level 1 tests whether an agent can reproduce a known vulnerability from its description and unpatched source code. It does not assess blind vulnerability discovery or confirm that generated patches are correct, limiting what the 95.95% result proves.
Architecture, benchmarks and open questions
MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer with 137 billion total parameters, five billion active parameters and a 256,000-token context window. Microsoft says the evaluated configuration replaced 80% of MDASH’s existing models and raised its reported CyberGym result from 88.4% to 95.95%.
CyberGym’s public leaderboard still showed Microsoft’s May submission at 88.4% when checked on July 28. Microsoft has not said whether the new result was submitted. Its cost comparison also omits token usage, latency, call volume, task mix and compute allocation, preventing independent reproduction.
The model is one input, the system around it is the product.
Taesoo Kim, Microsoft’s vice president of agentic security, used that distinction when describing MDASH. Separate tests produced mixed results, including zero across all ExploitGym categories. Microsoft conducted the evaluations in a network-isolated environment and warns that generated code requires review.
Access is limited to approved MDASH customers through an Azure AI Foundry private preview. For businesses, the practical opportunity is lower-cost triage and vulnerability reproduction, but procurement decisions should depend on workload-specific testing, human validation and transparent operating-cost data.

