Skip to content
All insights
AI7 min read

OpenAI, Google and Anthropic gate their most capable cyber security models

Three frontier labs announced cyber security models in 48 hours. The capability is real, but the strongest versions sit behind vetted access programs that most organisations will not qualify for.

Ada

Ada

Writer

OpenAI, Google and Anthropic gate their most capable cyber security models

Three frontier AI labs published cyber security announcements within roughly 48 hours at the start of September. Read together, they say something more useful than any one of them says alone.

What was announced

On 1 September, OpenAI stated that its forthcoming model, Astra, meets the Critical cyber security capability threshold under its Preparedness Framework. That threshold is met when a model can identify and develop working zero-day exploits across many hardened real-world systems without human intervention, or devise and execute end-to-end attack strategies against hardened targets given only a high-level goal. Astra is the first model OpenAI has designated at this level.

The supporting evidence is specific. In expert-led assessments, Astra found previously unknown vulnerabilities in a hardened browser and built them into a working chain that escaped the sandbox and executed commands on the host when the browser opened an HTML file. It also chained multiple flaws in a hardened operating system into a local privilege escalation from an unprivileged user to root. On an internal benchmark of 20 recently disclosed high-severity V8 vulnerabilities, it discovered and used two zero-days as part of an exploit chain, which OpenAI says it is disclosing to the maintainers.

OpenAI notes that those results reflect the model running with Daybreak Blue access rather than its default production configuration, and that it delayed parts of Astra's development and release for several weeks while it strengthened protections.

On 2 September, Google launched the Fairwind Program, described as limited access for governments and trusted partners. It pairs Gemini 3.8 Flash Cyber with Google's CodeMender harness to find, verify and fix vulnerabilities at agentic scale, with Google claiming deployment-ready patches in minutes rather than weeks of manual work, generated inside the organisation's own cloud environment. Google says it has more than 650 participating partners globally.

Also in early September, Anthropic released Claude Fable 5.1 and Claude Mythos 5.1. They are the same underlying model with different levels of safeguards. Fable 5.1 is generally available and can now be used to identify software vulnerabilities, though not to develop exploits for them. Mythos 5.1 is available only through trusted access programs, and Anthropic says it demonstrates the strongest cyber capabilities of any model the company has released.

The common thread is the door, not the model

All three labs shipped their most capable cyber security work behind a gate, and the gates are narrow.

Google's staging is explicit: governments and national cyber authorities, critical infrastructure operators across healthcare, telecommunications, energy and financial networks, and core technology platforms. Participating organisations agree to operational conditions, including restricting use to staff inside their own cyber security, incident response or penetration testing teams, and deploying protections such as multi-factor authentication.

OpenAI's advanced cyber security capabilities go first to a small group of alpha testers, with wider access following through Daybreak Blue.

Anthropic routes Mythos 5.1 through a Cyber Verification Program, and states that the model is currently available only to a set of US organisations, with expansion to broader domestic and international partners being coordinated with the US government.

Google frames the reasoning plainly. Early access gives trusted defenders an adaptation window to harden their systems before attackers can exploit new capabilities.

What this means for organisations outside those lists

None of the three announcements describes availability for Australian organisations specifically, and the access criteria published so far point at governments, critical infrastructure and large cloud and security partners. Most Australian businesses, including a good portion of the mid-market, will not be on those lists in the near term.

That produces an uncomfortable asymmetry, and it is worth stating without drama. The capability described in OpenAI's assessment is a capability that exists now. The defensive tooling built on top of it is being distributed slowly and selectively, by design. Organisations outside the approved lists do not get the adaptation window. They get the same threat environment on a longer timeline.

The practical response is not to wait for access. It is to assume that the interval between a vulnerability becoming public and it being exploited will keep compressing, and to work on the things that determine whether that compression hurts: knowing what you actually run, patching quickly and verifiably, and being able to detect activity you did not authorise.

The friction runs both ways

There is a second-order effect worth planning for if your teams already use these models.

OpenAI warns that Astra's safeguards may flag legitimate activity as cyber misuse, and that this can slow, pause or stop work, including defensive cyber security work and tasks that do not obviously look security-related. In ChatGPT or Codex a user may be asked to review the action before continuing. Through the API, the task simply stops.

Anthropic has moved in the other direction on precision, saying its updated cyber safeguards produce around 60 per cent fewer interventions per session in Claude Code than the previous generation. It still redirects several categories of dual-use work to its Opus models, specifically penetration testing, exploit generation and binary-based vulnerability scanning.

For a security team building AI into its workflow, that is an operational detail with real consequences. A pipeline that silently halts partway through a scan is a pipeline that needs monitoring, fallbacks and an understanding of which tasks will be refused before they are relied upon.

The bottom line

The headline is not that AI can now find vulnerabilities. That has been true for a while. The headline is that the organisations building these systems now describe them as capable enough to require gated distribution, and are being fairly transparent about it.

For everyone on the outside of those gates, the sensible reading is that the attacker timeline is shortening while the defender timeline is being rationed. Fundamentals get more valuable in that environment, not less.

Book your free security consultation

A no-obligation conversation with people who actually understand security. We'll review where you stand and show you the fastest way to close your biggest gaps.