OpenAI claims its GPT-6 models now comply more closely with their own safety rules than the previous generation. In internal evaluations covering over 54,000 tasks, GPT-6 Astra reportedly generated roughly half the alerts for high-severity misaligned behavior compared to GPT-5.6 Sol. GPT-6 Astra is also the first model to reach the “Critical” cybersecurity capability level under OpenAI’s Preparedness Framework. External audits produced mixed results, with some tests pointing to a concerning increase in rogue actions when certain safeguards were disabled. OpenAI ultimately delayed the release of GPT-6.1 after internal testing flagged numerous alignment issues.
Source: Read the original article

