OpenAI officially confirmed on September 25, 2026 that its internal research team discovered on June 27, 2026 the capability of prompt injections to replicate and spread between AI agents. This vulnerability was identified in a simulated training environment using the internal GPT-Red system, built on the GPT-5.4-mini architecture. The identified replication vectors include emails, file system writes, and even code comments. No real-world attacks have been recorded to date. OpenAI frames this disclosure as part of its ongoing efforts to strengthen model security against sophisticated exploits.
Source: Read the original article

