Anthropic has announced a significant expansion of its Cyber Verification Program (CVP), creating three tiered levels of access that let vetted security professionals use Claude models with meaningfully fewer cybersecurity-related restrictions than the general public gets. The announcement, made October 6, 2026, consolidates the company’s earlier Project Glasswing initiative and its original verification scheme into one structured framework, and it extends coverage to Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future model releases.
The Dual-Use Problem Anthropic Is Trying to Solve
The program exists because of a stubborn tension in AI-assisted security work: the same capability that helps a defender find a flaw in their own code can just as easily help an attacker find and exploit that flaw somewhere else. Anthropic’s public-facing models keep fairly conservative guardrails around cybersecurity tasks as a result, limiting most users to activities like code review, patching, vulnerability discovery within code they own, and triaging security alerts. The CVP is Anthropic’s attempt to open up more advanced capability specifically for people and organizations that can demonstrate legitimate authorization for riskier work, without weakening protections for everyone else.
Three Tiers, Three Risk Profiles
The restructured program sorts applicants into three bands based on the kind of work they do and how much authorization they can document:
- Defense Access is the broadest tier, covering security operations, incident response, malware reverse engineering, and vulnerability analysis. It is open to in-house security teams, universities, government bodies, hospitals, utilities, smaller security firms, open-source maintainers, and even individual researchers who can show a track record of responsible vulnerability disclosure. Anthropic says it aims to turn around applications for this tier within a few days.
- Red Team Access builds on that baseline by adding authorized penetration testing and red-teaming capabilities, but it is restricted to organizations such as internal red teams, government security teams, and dedicated testing firms that can prove they have permission to test their intended targets. Even at this tier, Anthropic keeps real-time controls in place that block ransomware deployment, damage to physical systems, and testing against the riskiest safety-critical systems. Because the review is more involved, it can take several weeks, though qualifying applicants get Defense Access in the meantime so their work isn’t held up entirely.
- Specialized Access is the narrowest and least-restricted tier, reserved for a small group of organizations authorized to test things like power grids, flight systems, telecom networks, interbank transfer systems, and government networks. Anthropic reviews these applications jointly with the U.S. government, and existing Glasswing participants are rolled directly into this tier for current models without having to reapply.
What the Benchmarks Show — and What They Don’t
To illustrate how the tiers behave in practice, Anthropic ran its Opus 5.5 model through CyScenarioBench, an internal benchmark designed to simulate multi-stage cyber operations, across 50 attempts at each access level. Under standard public access, every single attempt was blocked at the very first prompt. Under Defense Access, 46 of 50 attempts were still blocked, with four succeeding. Under Red Team Access, none were blocked, and 34 of 50 tasks were completed — getting close to the roughly 67.6% completion rate Anthropic has reported for a model running with no safeguards at all.
Those numbers are worth reading carefully: they come from Anthropic’s own benchmark, run by Anthropic, and they demonstrate how the guardrails behave on a controlled internal test rather than proving that no harmful request can ever slip through in the real world. Anthropic also disclosed that its Glasswing partners reported finding at least 129,000 verified vulnerabilities between April and July 2026, with open-source scanning efforts turning up another 5,500 between April and October, more than 33,000 of them rated high or critical severity. Fewer than half of partners reported actual patch counts, though, so those figures describe vulnerabilities found rather than vulnerabilities confirmed fixed.
Data Handling and Platform Availability
Participation in the CVP generally comes with a requirement that usage data be retained for misuse monitoring, with temporary carve-outs for customers who otherwise qualify for zero-retention arrangements. Anthropic says a planned set of “Enterprise Frontier Safeguards” will eventually let eligible organizations keep that monitoring data inside cloud infrastructure they control themselves, rather than with Anthropic.
Verified access is currently available through Claude’s own platform as well as Google Cloud Vertex AI and Microsoft Foundry, while access through Amazon Bedrock is limited to organizations that qualify for the Enterprise Frontier Safeguards arrangement. Existing program members keep their current settings and are automatically evaluated for eligibility as new models ship.
Why It Matters
For defenders, the expanded program is a meaningful signal that frontier AI labs are trying to formalize a middle ground between “fully restricted for everyone” and “fully open for anyone claiming good intentions.” Security teams evaluating whether AI-assisted tooling fits into their vulnerability research or red-team workflows now have a defined, identity-verified path to request it — one that comes with documented limits rather than an all-or-nothing switch.
Leave a Reply
You must be logged in to post a comment.