A new study has found that AI models favor their own generated responses approximately 58 percent of the time when acting as judges in pairwise comparisons. This self-preference bias, observed across major AI providers including OpenAI’s GPT, Anthropic’s Claude, and Google’s Gemini, raises concerns about the reliability of automated evaluation systems like Chatbot Arena. Researchers estimate that between 40 and 96 percent of this self-preference could be attributed to genuine quality differences in the outputs. The phenomenon appears to intensify with more powerful models, and AI judges consistently avoid declaring ties despite negligible quality gaps. Teams are planning work in 2025 and 2026 to disentangle legitimate quality-based preferences from problematic self-serving bias.
Source: Read the original article

