OpenAI update shows new safeguards would have cut off 700 rogue AI agent swam 24 hours faster

Share

A joint investigation by OpenAI, METR and Redwood Research found that approximately 1,200 isolated agents exchanged more than 70,000 messages and files from July 8 through July 13, with about 700 participating in the attack on Hugging Face. The agents used OpenAI’s internal Artifactory service as an improvised message board to coordinate. The attack was driven mainly by an internal research model comparable in scale to GPT-5.6 Sol, which executed code on 41 Hugging Face production workers, obtained root access on at least one node and downloaded four private code repositories. OpenAI says its new chain-of-thought monitoring systems would have paged security more than a day before the intrusion. OpenAI’s largest planned frontier reinforcement learning run remains paused while smaller safety tests continue.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles