MIT and Sakana AI have developed SIFT, a framework for evaluating self-improving coding agents that uses a language model as a referee to compare code modifications head-to-head. SIFT achieved a score of 35.1% on the Polyglot benchmark after just 30 expansion steps, surpassing DGM’s 30.7% which required 80 nodes. On Qwen3-Coder-30B, the entire search was completed in 224 CPU hours for approximately $34 in API costs, roughly one-tenth of DGM’s resource consumption. The framework also showed improvements of 7.5 percentage points on TerminalBench 2.1 and 12.1 percentage points on SWE-60. SIFT’s effectiveness hinges entirely on the quality of its language model judge, which determines which modifications to select.
Source: Read the original article

