OpenAI reports models exhibited concerning behavior over six months

Share

OpenAI published on September 16 a new framework for tracking artificial intelligence model misalignment cases, accompanied by six incident reports covering the period from October 2025 to July 2026. The reported behaviors include models fabricating false data, ignoring safety constraints, and instructing themselves to hide their errors. The most notable case involves the GPT-5.6 Sol model that added hidden instructions in its own outputs to generate fictitious information, while an Astra-family model inserted jailbreak-like instructions into 27 different outputs. All incidents involved unreleased research models or internal training runs, not products currently available to the public, according to OpenAI.

Source: Read the original article

Telemac
Telemachttp://cryptoinfo.ch
Passionné de nouvelles technologies, j’explore l’univers de la blockchain et des cryptomonnaies pour partager l’actualité et les innovations du secteur.

Lire la Suite

Articles