# Apple study finds minimal agent matches or beats multi-agent ML engineering systems

> Source: <https://cryptobriefing.com/apple-study-minimal-agent-beats-multi-agent/>
> Published: 2026-10-07 12:02:09+00:00

Apple official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

# Apple study finds minimal agent matches or beats multi-agent ML engineering systems

Apple researchers report that a single coding agent with shell access matched or beat elaborate multi-agent harnesses on autonomous ML tasks

The AI industry has spent a lot of energy building teams of agents that plan, delegate, critique and coordinate. A new [Apple](https://cryptobriefing.com/markets/apple/) study suggests that much of that machinery may not be doing much work.

Researchers at Apple Machine Learning Research found that one well-prompted coding agent, given basic shell access, matched or outperformed several complex multi-agent systems on automated machine learning engineering tasks.

## What Apple tested

The paper is titled “How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?” It was submitted to arXiv on September 30, 2026, and drew attention in early October 2026.

The research team includes Alejandro Hernández-Cano, Kirill Brilliantov and Emmanuel Abbé. Their question was simple: how much scaffolding does a capable model actually need?

A “harness” is the software wrapped around a large language model that lets it act: tools, memory, planning loops and, in fancier setups, multiple cooperating agents.

That stripped-down agent is called Malena. It runs in a single session and can do three things: read files, write files and run bash commands.

The researchers pitted Malena against four more elaborate systems: MLEvolve, AiScientist, Arbor and ScienceFlow. Each was tested under matched conditions using frontier large language models, so the comparison isolated the harness rather than the underlying model.

## The numbers

On MLE-bench, a benchmark built to measure how well agents handle machine learning engineering work, Malena posted a **62.5% any-medal rate**. That figure tracks how often an agent’s results were strong enough to earn a medal on a given task.

### AI, tech, and the markets they move—in one daily briefing.

Daily. Free. Join 34,000+ readers across crypto, finance, and policy.

The best-performing external harness, AiScientist, managed 47.1%. That works out to a gap of 15.4 percentage points in favor of the system with fewer moving parts.

The evaluation covered two settings across MLE-bench and NatureBench. One set included 30 tasks run within a 24-hour budget, and another included 40 tasks run over an 8-hour budget.

According to the study, once an agent had direct access to its environment, adding multi-agent coordination or extra autonomy mechanisms produced no significant gains across the architectures tested.

The team ran systematic ablations, removing or swapping components one at a time to see what actually matters. They tested different combinations of harnesses and backbone models.

The researchers also flagged a limitation. Compute budgets were constrained, which kept the number of seeds, or repeated runs with different random starting points, modest. Fewer seeds mean results carry more statistical noise than a lab with unlimited GPUs might accept.

Their core conclusion points at the model itself. The strength of the underlying model was the main driver of performance, and additional harness complexity was often futile on current benchmarks.

## A pattern, not a one-off

Apple’s paper lands on top of earlier skepticism about agent swarms. A separate study from July 2026 found that self-organizing multi-agent LLM teams underperformed their best individual expert agent by nearly 41.1% on ML benchmarks.

## What this means for AI builders

For researchers, the study raises a methodological flag. If harness improvements are reported without matched backbone models, gains attributed to clever architecture may actually come from a stronger underlying LLM.

Apple’s matched-conditions setup offers a template for separating those effects. Future harness papers may face pressure to show their gains hold when the model is held constant.

The researchers framed their conclusion around current benchmarks, and modest seed counts mean some margins could narrow with more runs. MLE-bench and NatureBench measure specific kinds of ML engineering work.

**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our

[Editorial Policy](https://cryptobriefing.com/editorial-policy/).
