Papers
arxiv:2609.35855

Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution

Published on Sep 25
· Submitted by
Yubin Lyu
on Oct 9
Authors:
,
,
,
,

Abstract

Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across preceding attempts. The procedure limits each refinement chain to a fixed depth and applies Pareto-filtered Top-N selection to bound the candidate pool. Across AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, Mara Chain delivers greater task-performance gains with fewer rollouts. It outperforms GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. It improves the pass rate by 20.2 and 22.5 percentage points over AHE and Meta-Harness on TerminalBench 2.1, respectively, and improves MuSiQue test nDCG@10 and Recall@10 by 0.104 and 0.131 over a hand-written retrieval pipeline.

Community

Paper author Paper submitter

Excited to share Mara Chain — rethinking failure as a stepping stone for AI system auto-evolution! 🧬

Standard propose–evaluate–select optimization discards every candidate that fails the acceptance bar — so later proposals keep revisiting the same failure modes. Mara Chain instead retains each rejected candidate together with its rollout traces, structured analyses, and residual failures, and iteratively refines it on the same minibatch. Failure becomes a stepping stone, not a dead end. Chain depth is fixed and the pool is bounded by Pareto-filtered Top-N selection.

📊 Highlights (same model, same budget):
• AppWorld (skill optimization): up to +20.5% relative over the best prior optimizer; reaches the target score with 65.5% fewer rollouts than GEPA (3,280 vs 9,514 rollouts; 11.9h vs 35.8h); best final score on all four metrics — 89.9/83.9/76.7/59.0 vs 48.2/35.7/33.3/19.4 unoptimized
• TerminalBench 2.1 (harness optimization): 71.9% pass rate, +20.2pp over AHE and +22.5pp over Meta-Harness — with fewer output tokens than the unoptimized baseline
• MuSiQue (retrieval-pipeline optimization): test nDCG@10 +0.104 / Recall@10 +0.131 over a hand-written pipeline, on a test set disjoint from the tuning split — real generalization, not memorization
• Verified on GLM-5 / DeepSeek / Qwen backbones (avg +40.3/+32.7/+23.4pp): the system improves, not a specific model

One framework optimizes skills, agent harnesses, and retrieval pipelines — no weight updates needed, fully white-box.

🔗 Code: https://github.com/ant-research/AntOmniEvo (quickstart in README)
📄 Paper: https://arxiv.org/abs/2609.35855

Bonus 🦞: it ships with a "lobster gym" visualizer — every candidate is a lobster leveling up, weak ones train in the basement, strong ones enter the golden hall. Full lineage tree, every fix traceable.

Happy to answer any questions — feedback and stars welcome!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35855
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35855 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.35855 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35855 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.