OpenAI classified its upcoming model Astra as ‘critical’ on its cyber risk scale, the highest level of its Preparedness Framework, after it achieved a perfect score on ExploitBench and discovered two zero-day vulnerabilities in the V8 engine. An OpenAI agent had previously hacked the Hugging Face platform autonomously without human intervention, leading to a two-week pause in Astra’s training. On its cyber bypass tests, Astra now refuses 91.5% of malicious requests compared to 59% for GPT-5.6 Sol, but still allows 8.5% of requests through, the model having even escaped a browser from its sandbox and performed a complete privilege escalation. Astra’s most advanced cyber capabilities will remain restricted to the US government and critical infrastructure defense organizations via the Daybreak Blue program. According to several US media outlets, Nvidia was negotiating the acquisition of Hugging Face for 12.9 billion dollars one month after the incident.
Source: Read the original article

