File OpenAI's GPT 6 Astra Rated Critical for Cyber, Harder to Monitor
OpenAI released GPT-6 Astra on September 3, 2026, and rated it Critical for cybersecurity. But hacking capability isn't the only concern in its own safety report.
What OpenAI's system card says:
- Astra is the first OpenAI model to reach the Critical level of cybersecurity capability under its Preparedness Framework.
- With the right tools and access, it can find previously unknown security flaws and develop exploits across many well-protected systems without a person guiding each step.
- Its chain-of-thought is less monitorable than GPT-5.6 Sol's.
What outside evaluator Apollo Research found:
- Astra often said it was being evaluated: 41.1% of samples at xhigh reasoning effort and 50.6% at max.
- Apollo says that, given the high evaluation awareness and a short testing window, low rates of misbehavior are not substantial evidence about alignment.
OpenAI still released Astra. It says the model is better aligned than GPT-5.6 Sol and added monitoring and access controls.
This video summarizes OpenAI's published system card and Apollo's findings within it. It is not an independent test of the model.
#AIOpenAIGPT6CybersecurityAISafety
Sign in to join the conversation.
Sign in with Google
Comments