OpenAI has released the system card for its new model, GPT-6 Astra, showing significant reductions in factual errors compared to its predecessor, GPT-5.6 Sol. The model achieved a 99.99 percent defense rate against direct prompt injections using OpenAI’s automated ‘GPT-Red’ training method, alongside strong refusal rates for single-turn jailbreak attempts. However, safety evaluations reveal that the model remains vulnerable to complex, real-world deployment risks.
Under multi-turn conversational attacks, Astra’s defense rate fell to approximately 67 percent. Furthermore, external evaluation by security firm Gray Swan found that Astra failed at least once in 8.5 percent of indirect prompt injection scenarios embedded within documents. While this marks an improvement over GPT-5.6 Sol’s 27 percent failure rate, it trails competing frontier models like Claude Opus 5, which recorded a 4.8 percent failure rate in similar tests.
These findings highlight ongoing security challenges for enterprise autonomous AI agents executing code and handling sensitive documents at scale. OpenAI noted that the system card evaluation metrics reflect bare model performance without additional production-level safety classifiers active.
Why it matters
Autonomous agent deployments remain vulnerable to indirect prompt injections, requiring secondary safety layers beyond raw model capabilities.
Multi-turn conversational jailbreaks still succeed roughly one-third of the time, posing persistent risk for client-facing AI applications.
Model evaluations show Claude Opus 5 maintaining a security edge over OpenAI models in handling untrusted document inputs.
Source: the-decoder.com



