IBM has developed BenchDrift, an open-source tool that measures how sensitive language models are to question rephrasing in benchmarks. A study conducted by IBM Research tested eight models across three widely used benchmarks, GSM8K, MMLU and MATH-Hard, revealing that minor wording changes can flip an answer from correct to incorrect. The most powerful models proved more vulnerable to rephrasing than weaker models, losing significant accuracy even when the meaning of the questions remained identical. The tool is available on GitHub with its complete source code, Jupyter notebooks and configuration files. This research is part of IBM’s broader effort to improve transparency and reliability in artificial intelligence assessments.
Source: Read the original article

