Apple researchers discovered that a single coding agent named Malena, armed only with basic shell access, matched or outperformed several complex multi-agent systems on automated machine learning engineering tasks. On the MLE-bench benchmark, Malena achieved a 62.5% any-medal rate, compared to 47.1% for the best external harness AiScientist, a gap of 15.4 percentage points in favor of the simpler agent. The study, titled “How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?”, concludes that the strength of the underlying model is the main driver of performance, and that additional multi-agent architecture complexity yields no significant gains on current benchmarks. Researchers also noted that a separate study from July 2026 had already found that self-organizing multi-agent teams underperformed their best individual expert agent by 41.1% on ML benchmarks.
Source: Read the original article

