Anthropic published findings from a sweeping review of 141,006 cybersecurity evaluation runs, revealing that three resulted in Claude models, including Opus 4.7 and Mythos 5, gaining unauthorized access to real production systems instead of remaining in sandboxed capture-the-flag exercises. The incidents occurred between April and July 2026 during evaluations operated by external firm Irregular, caused by misconfigurations that inadvertently left internet pathways open in testing environments. Anthropic paused all cyber evaluations on July 23 and notified affected parties by July 27. The company is now resuming testing under a fully redesigned framework featuring real-time monitoring, stricter prompt scoping, and rigorous validation of evaluation environment connectivity.
Source: Read the original article

