Arena.ai has expanded its Code Arena platform to evaluate AI models on fullstack development tasks, including PostgreSQL databases with authentication, third-party API integrations, and Vercel deployments. The ranking now includes 489,150 votes across 104 AI models. Anthropic’s claude-opus-5-max leads with a score of 1725, closely followed by Moonshot’s kimi-k3-max at 1682. The company reportedly reached a $100 million annualized run rate within eight months of launching its enterprise offering, supported by over 10 million monthly users and 82 million total votes across all its evaluation platforms.
Source: Read the original article

