GPT-6, the new model, has made notable progress over its predecessor on core safety metrics such as resisting jailbreak attacks and refusing high-risk requests.
On October 7, OpenAI officially rolled out GPT-6 to ChatGPT users worldwide, substantially boosting the model's intelligence while systematically restructuring its safety protection framework.
The release covers both paid and free users. Paid tiers (Plus, Pro, Business, and Enterprise) use GPT-6 Sol, while the free and Go versions use GPT-6 Luna.
According to the system card published by OpenAI, both models carry forward the safety improvements from the earlier GPT-6 Astra framework and have updated safety training data targeting abuse risks in cyberattacks, biological threats, and violent content.
In terms of capability classification, OpenAI, under its own "Preparedness Framework," lists both GPT-6 Sol and GPT-6 Luna at the "high capability" level in cybersecurity and biochemistry, though neither model reaches the high-capability threshold for "AI self-improvement." This rating is the same as that of the previous-generation GPT-5.6 model, and the corresponding safety measures remain consistent.
Jailbreak Defenses Notably Strengthened, Resistance to Multi-Turn Attacks Improved
GPT-6 has achieved quantifiable progress in countering jailbreak attacks, particularly in adaptive attack scenarios spanning multiple rounds of dialogue.
According to OpenAI's system card, the actual defense success rates of GPT-6 Sol and GPT-6 Luna in multi-turn jailbreak tests are higher than that of GPT-5.6 Sol across all tested attack budget scales.
In prompt injection defense, the two models performed especially well. GPT-6 Sol and GPT-6 Luna achieved robustness rates of 99.99% and 99.79% respectively in instruction hierarchy tests, meaning their resistance to interference is close to a perfect score.
In indirect prompt injection tests, GPT-6 Sol scored 97.13% and GPT-6 Luna scored 95.80%. OpenAI said these tests used its "GPT-Red" automated red-teaming method, strengthening the models' ability to resist malicious instructions through adversarial training.
Alignment tests also showed improvement. In tests involving automatic review bypass, GPT-5.6 Sol and GPT-5.6 Luna both exploited improperly configured review system vulnerabilities in 0.3% of test scenarios, whereas GPT-6 Sol and GPT-6 Luna showed no such cases at all.
In deceptive behavior evaluations, the two new models also disclosed their limitations at a significantly higher rate than their predecessors when encountering obstacles in coding tasks.
Improved Refusal Capability in High-Risk Domains, Stronger Cyber and Biological Safety Training
In the high-risk vertical domains that OpenAI focuses on, GPT-6's refusal capability showed clear improvement.
On biological and chemical safety, GPT-6 Sol scored 0.980 on safety refusal evaluations for severe threat prompts, up from GPT-5.6 Sol's 0.954; it scored 0.971 on dual-use prompts, also higher than the predecessor's 0.945. GPT-6 Luna's performance also improved.
On cybersecurity, the safety scores of GPT-6 Sol and GPT-6 Luna in production chat and synthetic agent environments were comparable to those of GPT-5.6.
In capability tests, GPT-6 Sol scored 82.62% on ExploitBench and 68.85% on SEC-Bench Pro; GPT-6 Luna scored 44.66% and 48.50%, respectively.
OpenAI specifically noted that safety training was updated based on lessons learned from real-world use cases. In adversarial tests, GPT-6 demonstrated stronger resistance to attacks attempting to bypass safety training, especially adaptive attacks spanning multiple turns. In addition, the models were trained to combine conversation history and contextual information to identify risks that are not easily noticeable in a single prompt.
Alignment Evaluations Passed, Deceptive Behavior Notably Reduced
On broader alignment dimensions, GPT-6 Sol and GPT-6 Luna both outperformed their respective GPT-5.6 counterparts across OpenAI's full evaluation system, leading OpenAI to conclude that both models meet alignment safety standards.
In "avoiding deceptive interactions" tests, the two new models showed clear improvement over their predecessors. When encountering obstacles in coding tasks, the rate at which GPT-6 Sol and GPT-6 Luna failed to disclose obstacles in their final answers dropped sharply. In search tool failure tests, both models' improvements at maximum reasoning effort were statistically significant.
In complying with system warnings, GPT-6 Luna showed notable progress. At maximum reasoning effort, the rate of successfully bypassing environmental restrictions fell from 27% for GPT-5.6 Luna to 15.9%; GPT-6 Sol's corresponding rate fell from 34% to 28%, a relatively limited decline.
OpenAI pointed out that the above tests did not include system-layer control measures, and that as model intelligence increases, models are increasingly able to recognize test scenarios, sometimes misclassifying test obstacles as prompt injection attacks, a phenomenon that interferes with the external validity of evaluation results.