Apple Study Reveals Single Agent Outperforms Multi-Agent Systems in ML Tasks
A new study from Apple Machine Learning Research has challenged the effectiveness of multi-agent systems in machine learning engineering. Researchers found that a single, well-prompted coding agent, named Malena, demonstrated performance on par with or superior to several complex multi-agent systems.
The study, titled “How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?” was submitted to arXiv on September 30, 2026, and highlights the potential inefficiencies of elaborate agent systems. Malena operates with basic shell access and can read, write files, and execute bash commands.
In head-to-head comparisons, Malena achieved a 62.5% success rate on the MLE-bench benchmark, significantly outperforming the best external harness, AiScientist, which managed a 47.1% success rate. This 15.4 percentage point gap suggests that simpler systems may be more effective than previously thought.
The study also pointed out that adding complexity through multi-agent coordination did not yield significant performance improvements. The researchers emphasized that the underlying model's strength was the primary factor in performance, rather than the complexity of the harness surrounding it.
This research aligns with previous findings that questioned the efficacy of multi-agent systems, suggesting a need for a reevaluation of how AI agents are structured and deployed in practical applications.
FAQ
What is the main finding of the Apple study on machine learning agents?
The study found that a single, well-prompted coding agent named Malena outperformed several complex multi-agent systems in machine learning tasks, suggesting that simpler systems may be more effective.
What benchmark did Malena achieve a success rate on?
Malena achieved a 62.5% success rate on the MLE-bench benchmark, significantly outperforming the best external harness, AiScientist, which had a 47.1% success rate.
What capabilities does the agent Malena possess?
Malena operates with basic shell access, allowing it to read and write files and execute bash commands.
What does the study suggest about the complexity of multi-agent systems?
The study suggests that adding complexity through multi-agent coordination did not yield significant performance improvements, indicating that the strength of the underlying model is more crucial than the complexity of the system.
When was the study submitted to arXiv?
The study was submitted to arXiv on September 30, 2026.
Comments
Comments are moderated before publish.
No comments yet — be the first.