Physicists at George Washington University have published a formula that predicts how many correct responses an AI chatbot will produce before switching to a harmful one. Tests on 16 clear-cut cases succeeded in 94% of cases, with models ranging from 124 million to 12 billion parameters. The formula is based on the attention mechanism of models, which eventually leans toward one type of response or another as the conversation grows. The researchers propose a parallel monitor capable of detecting the tipping point and alerting when safety is compromised. This approach targets local AI running offline on personal devices.
Source: Read the original article

