A new analysis of public data from Anthropic’s Project Glasswing has highlighted a significant gap between the number of vulnerability findings generated by its Claude frontier model and those that ultimately prove to be real, serious, and worth fixing.
The distinction matters because it suggests that the bottleneck in vulnerability research may increasingly lie in validating new flaws and coordinating their remediation rather than in discovering them.
Barely 10% Have Made It to Disclosure Stage
Patrick Garrity, a security researcher at VulnCheck, recently analyzed Anthropic’s Vulnerability Disclosure Ledger, which is a public record tracking Project Glasswing-related findings as they move through the vulnerability disclosure and remediation process. The analysis showed that Anthropic’s Claude Mythos generated a total of 26,153 vulnerability findings across numerous software projects since Project Glasswing’s launch in April 2026.
Related:US Government Accuses Chinese AI Firms of Distilling Frontier Models
However, only 2,736 of those findings, or slightly more than 10%, had made it into the disclosure ledger, meaning they have either been disclosed to the appropriate software maintainer or are in the process of being disclosed. Less than 0.8% of flaws, a mere 202, are currently patched, and 245 were withdrawn. Another 191 vulnerabilities were in the pre-disclosure stage and had not been reported to their maintainers yet.
The remaining, nearly 90% of Claude Mythos-generated findings, had not made it to the ledger yet, suggesting human validation and coordination has become a bottleneck in determining which AI-generated findings warrant disclosure and remediation, says Garrity.
The results are "not a surprise for those of us closer to understanding how coordinated vulnerability disclosure works," Garrity tells Dark Reading. But it "is much different than the narrative frontier model providers have positioned," which has largely focused on AI's ability to dramatically accelerate vulnerability discovery.
"It seems like they are learning this through trial and error," he says.
True Positives and Severity Assessments
... continue reading