Booz Allen: Frontier AI models to reach cyberattack parity within six months

This digest was compiled by AI from multiple sources — links to the originals are below.
Booz Allen confirmed on Sept. 2 that Anthropic's Mythos 5 can autonomously compromise a production-grade enterprise network, scoring 80 on its new Cyber Weapon Index. The firm projects frontier and Chinese AI models will reach cyberattack capability parity within six months. The warning follows a July AI-assisted attack on Taiwanese government servers that compressed reconnaissance and exploitation into four days.
Key Facts
- Booz Allen's Cyber Weapon Index scored Anthropic's Mythos 5 at 80, compared with 49 for SpaceXAI's Grok-4.5.
- The UK AI Security Institute reported in June that Mythos completed an end-to-end attack chain in 3 of 10 capture-the-flag attempts.
- OpenAI's GPT-5.5 completed the 32-step attack chain test in 2 of 10 attempts, according to the same June report.
- A July attack on Taiwanese government servers by a Chinese-speaking threat group compressed reconnaissance and attack execution into a four-day window.
- Tenable stated the agents chose which systems to map and when to expand into new sectors without step-by-step human direction.
Cyber Weapon Index
Booz Allen released the Cyber Weapon Index on Sept. 2 to benchmark a model's ability to find and exploit vulnerabilities and execute an attack. Mythos 5 scored 80 on the index, while Grok-4.5 scored 49. Brad Medairy, president of Booz Allen's National Cyber practice, said the current standings matter less than the trajectory over the next six months. Medairy expects parity between frontier models and Chinese models within roughly six months.
Autonomous Attack Evidence
The UK AI Security Institute reported in June that Mythos completed an end-to-end attack chain in capture-the-flag contests. OpenAI's GPT-5.5 completed the 32-step attack chain test in the same evaluation. Mythos succeeded in 3 of 10 attempts, and GPT-5.5 succeeded in 2 of 10 attempts. A July attack on Taiwanese government servers by a Chinese-speaking cyber threat group demonstrated near-autonomous operations over four days. Tenable's analysis of that incident and six similar AI-powered attacks found the agents chose systems to map, techniques to use, and when to expand without step-by-step human direction.
Defense Speed Gap
Medairy said human defenders currently hold the advantage, but the balance of power will shift to offense as attackers deliver effects at speed and scale. A four-hour response time is considered reasonable or even exceptional today, but Medairy said it will not be fast enough in the future. Booz Allen's findings align with other benchmarks confirming at least one frontier model can autonomously execute an end-to-end compromise.