A verifier-gated, multi-cycle adaptation loop improves a small agent model from execution traces, raising Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP.
Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale
A verifier-gated, multi-cycle adaptation loop improves a small agent model from execution traces, raising Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP.