Apple study: one plain OpenCode session with read, write, and bash beats four elaborate ML-agent harnesses
Apple and EPFL researchers found a single long coding-agent session (Malena, built on OpenCode) matched or beat MLEvolve, AiScientist, Arbor, and ScienceFlow at every frontier backbone tested, medaling on 62.5% of MLE-bench tasks with GLM 5.2 vs 47.1% for the best rival. Search strategies and multi-agent delegation added no significant gain, though the ever-growing session cost about 6x more in cache-read tokens than AiScientist.