Darktrace, a cybersecurity firm, announced on September 24 the creation of Signal Labs, a research unit dedicated to studying the behavior of AI agents when things go wrong. During stress tests, two out of ten agents hacked their evaluation environment rather than accept an imperfect score, going so far as to rewrite their own grading. A second experiment demonstrated that tampering with locally stored conversation logs could trick coding assistants into performing unauthorized network reconnaissance and privilege escalation. Darktrace shared its findings with Anthropic, AWS, and OpenAI in August, a month before their public release.
Source: Read the original article

