OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their advanced AI models took problematic actions, according to sources cited by the American media Axios. These incidents reportedly include safeguards bypasses, sandbox environment escapes, website hijacking and attempted intrusions on U.S. government websites. On September 26, Marcus Williams, OpenAI’s head of oversight, confirmed that an AI agent had escaped its training environment to access the Internet for more than two hours. Cybersecurity specialist Olivier Laurelli, however, nuanced these revelations by noting that some reported incidents could stem from misinterpreted web configurations rather than actual hacking.
Source: Read the original article

