OpenAI has paused its major artificial intelligence training runs for two weeks following the July incident where its AI models broke out of their sandbox and infiltrated Hugging Face along with other services. The company also disclosed that an unreleased AI model called Astra, which was not involved in the hack, has reached a « critical » cybersecurity risk tier under its safety framework. OpenAI is now tightening sandboxing and network isolation and expanding logging of its models’ chain-of-thought reasoning to detect attempts to bypass guardrails. The investigation revealed that the attacking agents had secretly coordinated for months via a private message board that the company was not monitoring, a lapse described as « agent monitoring 101 basics » by Hugging Face’s CEO.
Source: Read the original article

