CurateEvo evolves executable data-curation code using failed agent trajectories, improving post-training performance by 3.2 and 2.7 points over baselines on labeled and wild data respectively.
A gent T uning: Enabling Generalized Agent Abilities for LLM s
5 Pith papers cite this work, alongside 26 external citations. Polarity classification is still indexing.
years
2026 5representative citing papers
BEACON uses milestone partitioning, temporal reward shaping, and dual-scale advantage estimation to nearly double success rates on long-horizon ALFWorld tasks while raising effective sample use from 23.7% to 82%.
ATG maintains explicit DAGs of subtasks to enable dependency tracking, parallel execution, and localized repair in LLM agents, outperforming baselines on three benchmarks with 7B-8B models.
StepGuard framework with DDPO and CANR claims SOTA navigation and answer accuracy on web benchmarks by switching policies and triggering reflection on low-confidence steps.
citing papers explorer
-
CurateEvo: Data-Curation Evolving for Agentic Post-Training
CurateEvo evolves executable data-curation code using failed agent trajectories, improving post-training performance by 3.2 and 2.7 points over baselines on labeled and wild data respectively.
-
Milestone-Guided Policy Learning for Long-Horizon Language Agents
BEACON uses milestone partitioning, temporal reward shaping, and dual-scale advantage estimation to nearly double success rates on long-horizon ALFWorld tasks while raising effective sample use from 23.7% to 82%.
-
Atomic Task Graph: A Unified Framework for Agentic Planning and Execution
ATG maintains explicit DAGs of subtasks to enable dependency tracking, parallel execution, and localized repair in LLM agents, outperforming baselines on three benchmarks with 7B-8B models.
-
StepGuard: Guarding Web Navigation via Single-Step Calibration
StepGuard framework with DDPO and CANR claims SOTA navigation and answer accuracy on web benchmarks by switching policies and triggering reflection on low-confidence steps.
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents