OpenAI’s Astra/GPT-6 classified 'Critical' for cyber capability in company documentation
An internal Preparedness Framework designation and evidence release reportedly show Astra can find and exploit zero-day vulnerabilities; OpenAI has gated access and added monitoring that can interrupt tasks but warns about new failure modes.
In this brief: 2 sections 1 min read
Anavem reports OpenAI published a designation and system card saying Astra discovered two zero-day vulnerabilities in evaluation.
The 'Critical' threshold is defined by the model's ability to find and exploit novel security flaws across hardened systems without step-by-step human guidance.
OpenAI said it is in the process of disclosing those vulnerabilities to maintainers.
Astra is off by default in Business and Enterprise workspaces and requires deliberate admin enablement.
OpenAI's monitoring can slow, pause or stop tasks (including legitimate defensive work) when safety checks trigger.
API usage reportedly exposes new failure modes where long-running agent tasks may be stopped mid-run.