How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Published in arXiv preprint, 2026

We compare elaborate autonomous machine learning engineering harnesses with a minimal coding-agent baseline under the same time budget and frontier language model. Across systematic ablation studies, the more elaborate harnesses provide no advantage, suggesting that the model backbone is the primary driver of performance on current benchmarks.

Recommended citation: Kirill Brilliantov, Alejandro Hernández-Cano, and Emmanuel Abbé. "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?" arXiv preprint arXiv:2609.40303, 2026.
Download Paper