Pith. sign in

REVIEW 3 major objections 4 minor 69 references

PATS reframes skills as a disposable training scaffold: it expands, revises, and compresses context during RL, then discards it, matching state-of-the-art performance on agentic benchmarks while using 25–50% fewer deployment tokens.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 07:28 UTC pith:4XDUZTCF

load-bearing objection PATS is a solid, well-ablated paper with a real conceptual shift, but headline claims need qualification and the refiner's semantic reliability is an unmeasured dependency. the 3 major comments →

arxiv 2607.21419 v2 pith:4XDUZTCF submitted 2026-07-23 cs.AI

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

classification cs.AI
keywords reinforcement learningLLM agentstraining scaffoldskill internalizationcurriculum learningGRPOtoken efficiencyagentic RL
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that the right role for skills in agentic reinforcement learning is not to serve as inference-time assets, but to act as a temporary, policy-aware training scaffold. PATS converts each rollout group into an evidence card, tracks per-task-type success rates, and chooses whether to expand, semantically revise, compress, or prune a natural-language context bank before the next rollout batch. The policy is optimized only by environmental rewards, and the scaffold is removed entirely at deployment. If the central claim holds, an agent can end up skill-free and still match skill-retaining state-of-the-art methods, with 25–50% lower deployment token cost, because the explicit guidance has been parameterized into the model rather than retrieved at test time.

Core claim

On its own terms, PATS's central discovery is that training-effective support is not the same as inference-effective support. A fixed, high-quality skill library can raise immediate success while shrinking the success–failure variation inside rollout groups, which is exactly the contrast GRPO needs; PATS instead lets a controller schedule support so that weak policies receive concrete guidance, intermediate policies receive semantic revisions that track residual failures, and strong policies receive compression. The paper reports that across ALFWorld, WebShop, and seven search-augmented QA benchmarks, this schedule yields performance within about ±3% of state-of-the-art skill-based methods,

What carries the argument

The central mechanism is the policy-aware scaffold controller. It maintains a three-layer bank of natural-language entries (general principles, task-type procedures, mistake corrections), and for each task type it tracks an exponential moving average of success rate plus a global capacity pressure. These two signals select one of four edit modes—EXPAND, REVISE, COMPRESS, or FORCED_PRUNE—each with write budgets enforced by a deterministic validator. Rollout groups are condensed into evidence cards by a deterministic builder; a frozen 7B language-model refiner proposes schema-valid edits; and atomic commits make those edits visible only to the next rollout snapshot, so no batch consumes its ow

Load-bearing premise

The load-bearing premise is that the frozen 7B language model, guided only by evidence cards and mode instructions, proposes semantically valid and training-useful scaffold edits; schema checks catch format and duplication errors but not misleading content, so a bad proposal would inject harmful context into the next rollout.

What would settle it

Run PATS with the same controller and budgets but replace the frozen refiner with a random-but-schema-valid proposer (or an adversarial one). If the gains persist, the scheduler alone carries the method; if they collapse, semantic edit quality is the load-bearing component. A complementary audit would have humans label refiner proposals for semantic correctness and test whether invalid proposals correlate with next-iteration performance drops.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Deployment becomes skill-free by construction: the final policy is evaluated with the experience block removed, so any performance gain must be carried in the parameters, not in a retrieved library.
  • Token cost at inference drops 25–50% relative to skill-retaining baselines, because the scaffold contributes no prompt tokens once training ends.
  • The expand→revise→compress schedule gives a principled answer to when to add versus withdraw guidance: add when success is low, revise when failures shift, compress when competence is reliably high.
  • The offline-replay analysis implies that a static skill library can be a worse training curriculum even when it is a better inference aid, which reframes skill curation in agentic RL around rollout contrast rather than immediate utility.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: the same controller could schedule not just textual scaffolds but any auxiliary context—plans, rubrics, retrieved documents—by treating rollout success-rate EMA as a general curriculum signal; this is testable by swapping the scaffold content while keeping the scheduler fixed.
  • Inference: since only schema validity is checked deterministically, a natural stress test is to compare PATS against a random-but-valid refiner; the performance gap would isolate whether the gains come from scheduling or from the frozen refiner's semantic quality.
  • Inference: the likelihood-gap number is measured on a narrow two-gamefile diagnostic, so the broad claim of 'skill internalization' should be read as a mechanism hypothesis until replicated on a wider distribution—the paper itself flags this boundary.
  • Inference: if the intermediate-success-rate argument generalizes, then deliberately sampling tasks at ~50% success, or adapting scaffold support to hold groups near that contrast peak, could amplify GRPO signal beyond what PATS currently does; the paper explicitly does not target 50% success, so this is an open extension.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes PATS, a policy-aware training scaffold for agentic RL. At each RL iteration, rollout groups are converted into evidence cards; a competence controller chooses among EXPAND, REVISE, COMPRESS, and FORCED_PRUNE based on a task-type success-rate EMA and scaffold pressure; a frozen Qwen2.5-7B refiner proposes schema-constrained edits that are deterministically validated and atomically committed; the resulting scaffold is injected into rollout prompts during training and removed at deployment. The paper claims that PATS achieves performance competitive with SOTA skill-based baselines while using 25%–50% fewer deployment tokens, based on experiments on ALFWorld, WebShop, and seven search-augmented QA benchmarks with Qwen2.5-1.5B/7B, full test sets, three seeds, and component ablations.

Significance. If the empirical claims hold, the paper makes a useful conceptual contribution: reframing skills as a disposable, policy-conditioned training scaffold rather than as persistent test-time artifacts. The strengths are real: evaluations use fixed full test sets; results are reported as three-seed means with standard deviations; ablations isolate components (online adaptation, REVISE, group evidence, interface initialization); the appendices provide detailed schemas, budgets, operation logs, and audit trails; and the code is publicly linked. The paper also states its own limitations clearly, notably that the Stage-0 internalization study is a controlled diagnostic rather than a general skill-acquisition claim. However, two issues currently sit at the load-bearing points of the central claim: the semantic reliability of the refiner's edits is not measured, and the headline comparison to GRPO conflates the scaffold-conditioned warmup with the online adaptation mechanism. Both are fixable with additional analysis, so the appropriate disposition is major revision rather than rejection.

major comments (3)
  1. [§3.3, Appendix B.1, Appendices B.8–B.10] The central causal claim is that policy-aware adaptation improves training by injecting correct, useful guidance. The validator described in §3.3 and Appendix B.1 checks only schema validity, evidence-card count, duplication, budgets, and capacity; it does not verify that a proposed principle is entailed by the evidence, is actionable, or is not misleading. Because the scaffold is inserted into every rollout prompt (Eq. 7), a bad refiner proposal changes the training distribution. The paper's own audit (Appendices B.8–B.10) shows one correct lineage and one commit, and explicitly disclaims causal attribution; it provides no error rate. I request a quantitative audit of refiner outputs across seeds and task types—e.g., human or LLM-judge annotations of semantic validity, usefulness, and harmfulness—plus a report of the fraction of commits that were later reverted or contradicted by subseq
  2. [§4.1, §4.2, Tables 1 and 3] The headline comparison to GRPO conflates two effects. In Table 3, 'w/o training scaffold' (same interface-initialized checkpoint, no online scaffold) reaches 73.57±0.89 SR on 1.5B ALFWorld, while the GRPO baseline in Table 1 is 67.86±3.50; thus the scaffold-conditioned SFT warmup alone accounts for +5.7 SR, and the online scaffold contributes the remaining +7.1 of the reported 12.9-point end-to-end gain. Since the warmup is trained on public SkillRL-format data (Appendix B.1), the Table 1 comparison attributes to PATS an improvement that is partly a data/SFT effect. Also, Table 3's 'w/o interface initialization' row (47.86±3.29) shows that without the warmup PATS is far below GRPO. I request that the paper present the decomposition 'GRPO vs. warmup-only vs. warmup + online scaffold' as the primary reading, and that all abstract/conclusion claims specify whether they refer to the full PA
  3. [Abstract, §4.1, Table 1, Appendix B.1] The '25%–50% fewer tokens' claim is stated in the Abstract and Conclusion as if it were a global resource reduction, but the metric in Table 1 is explicitly deployment interaction tokens only. Training-time service costs are substantial: the audited 1.5B run logs 831 refiner requests, 44.1M input characters, and ~180 GPU-hours (Appendix B.1). If the claim is about total compute or total token consumption, the current statement is misleading; if it is about deployment tokens, the Abstract should say 'deployment tokens' and the paper should provide the full training-time token/cost breakdown so readers can assess the end-to-end efficiency trade-off.
minor comments (4)
  1. [Appendix C and §3.3] The Bernoulli calculation is correct but does not determine the controller thresholds; the appendix itself says the thresholds remain empirical. Please state explicitly in the main text that the intermediate-competence argument is motivational and not a parameter-free derivation, to avoid an unintended implication.
  2. [Table 2] The 'Avg.' row appears to be a macro average across seven datasets of very different sizes (NQ 3,610 examples vs. Bamboogle 125). Please define whether this is macro or micro, and consider reporting the micro average as well, since the conclusion depends on the pooled number.
  3. [Figure 1] The figure would be easier to interpret with explicit labels on both y-axes and a statement of whether the PATS curve is evaluated with or without the scaffold at validation. Currently the caption leaves the evaluation condition implicit.
  4. [§4.3] Minor wording/typo: 'Analyze behavioral shifts' should be 'Analyzing behavioral shifts'; also, the fact that the NLL set covers only two gamefiles should be stated in the main text near the claim, not only in Appendix A.

Circularity Check

1 steps flagged

Minor self-confirmatory framing around Figure 1; central empirical claims are independent.

specific steps
  1. self definitional [Section 1 (Introduction), discussion of Figure 1; cf. Eq. (14) controller modes]
    "Figure 1 shows the intended training scaffolding pattern: skills expand when the weak policy needs concrete guidance and contract as performance improves. This confirms that the performance gains stem from adaptive training support rather than reliance on well-crafted or continuously evolving skills."

    The observed expand-then-contract pattern is generated by the paper's own controller (Eq. 14): EXPAND is selected when s̄τ,k < ρr and COMPRESS when s̄τ,k ≥ ρc, so the curve is an implementation of the rule, not independent evidence that adaptation causes the gains. The causal claim therefore does not follow from Figure 1 alone. However, the central performance claim is supported separately by ablations (Tables 3 and 8: frozen, EXPAND-only, and w/o REVISE variants all degrade), so this confirmatory sentence is rhetorical rather than load-bearing.

full rationale

Apart from the Figure 1 framing, PATS's derivation is not circular. The central claim is an empirical training-efficiency comparison (Tables 1 and 2), not a quantity fitted from its own output. The only formal derivation, Appendix C's Bernoulli calculation (Eqs. 20-23), is a standard probability identity used to motivate the REVISE regime and is explicitly disclaimed as a convergence proof or a target; it does not back-fit the results. The frozen qwen2.5-7b refiner and public SkillRL warmup data are external ingredients, not fitted predictions, and the unmeasured semantic reliability of refiner edits is a robustness/correctness limitation rather than circularity. Self-citations (e.g., Skill1) appear only in related work and are not load-bearing. The paper's own limitation statements (Appendices A.3, B.8, C.2) appropriately constrain interpretation. Token reductions against skill-retaining baselines follow in part from the scaffold-free deployment definition (Eq. 2), but the paper also reports token reductions versus scaffold-free GRPO and, crucially, competitive performance after scaffold removal, so the headline claim does not reduce to its definition.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 0 invented entities

The paper introduces no new physical or ontological entities. Free parameters are the controller thresholds, EMA coefficient, capacity limits, and edit budgets, all hand-chosen. The key domain assumptions are the Bernoulli model for mixed groups, the efficacy of SFT warmup, and the reliability of the refiner LLM—the last is the most fragile but is not independently validated in the text.

free parameters (4)
  • Controller thresholds ρ_r, ρ_c = ρ_r=0.3, ρ_c=0.85 (ALFWorld/WebShop 1.5B); ρ_r=0.2, ρ_c=0.4 (QA 7B)
    Hand-chosen per task family; determine when EXPAND/REVISE/COMPRESS fire and directly control the scaffold dynamics.
  • EMA coefficient α = 0.1
    Chosen without sensitivity analysis; affects responsiveness of success-rate estimate.
  • Scaffold capacity (Nmax, Lmax) = 30 entries / 2,000 tokens
    Hand-set; FORCED_PRUNE triggers when these are hit.
  • Retrieval and edit budgets = top 4 task / 3 general / 3 mistake; EXPAND permits 2+1+1 proposals
    Hand-set operation caps enforced in refiner validation.
axioms (4)
  • domain assumption Binary-reward approximation: group success probability p and Bernoulli model for mixed groups
    Appendix C motivates REVISE via Pr(mixed) maximized at p=0.5; paper acknowledges real environments violate this (reward density, stochasticity).
  • domain assumption Policy can internalize textual scaffold guidance after SFT warmup
    Scaffold-conditioned SFT (Eq. 15) assumes the model learns to follow principles/mistakes; ablation shows removing it collapses performance (Table 3, -32.85).
  • domain assumption Frozen refiner produces semantically valid and useful edits
    The entire adaptation path depends on Qwen2.5-7B refiner proposals; only schema validity is checked deterministically, not semantic usefulness.
  • standard math GRPO group-relative advantages are the right learning signal
    Standard PPO-style clipped objective used as-is; not a new contribution.

pith-pipeline@v1.3.0-alltime-deepseek · 161 in / 10391 out tokens · 131170 ms · 2026-08-01T07:28:35.372666+00:00 · methodology

0 comments
read the original abstract

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. However, they remain centered on the skills themselves rather than being designed as adaptive training-time support for the evolving policy. To address this, we propose a policy-centric training paradigm that reframes skills as a dynamic training scaffold. Our framework, PATS, converts rollout groups from the latest policy into evidence cards and uses task-specific evaluation to adjust the context used in subsequent rollouts. Concrete guidance helps weak policies to complete challenging tasks. As policy improves, redundant context is revised or removed to reduce reliance on explicit guidance while preserving useful rollout variation. The policy is optimized with environmental rewards using standard RLVR, and the training scaffold is discarded at deployment. Across ALFWorld, WebShop, and seven search-augmented QA benchmarks, PATS achieves performance competitive with SOTA baselines while using 25%-50% fewer tokens.

Figures

Figures reproduced from arXiv: 2607.21419 by Peng Chen, Qitai Tan, Yang Li, Yipeng Shi, Yue Wang, Zhengzhou Zhu, Zhipeng Ma.

Figure 1
Figure 1. Figure 1: Training dynamics on ALFWorld under the shared [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of different training methods. Left: Skill-expansion methods retain a growing set of experiences, whereas [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of Pats. Skills are injected into rollout requests as context, and the resulting trajectories are used to update the policy via GRPO. Policy updates and skill updates are executed in parallel. The updated skills take effect in the subsequent iteration and are discarded during deployment. 3 Methodology 3.1 Overview We study interactive tasks drawn from distribution D. Each task q belongs to a task … view at source ↗
Figure 4
Figure 4. Figure 4: NLL Gap between two variants. The narrowing [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Controller-mode evolution on 1.5B ALFWorld. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Interface initialization precedes Stage 1. Step-zero validation differs only modestly, but the initialized policy learns [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Policy-facing experience block rendered before the current interaction history. [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Schema of the deterministic group-level evidence card. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Minimal refiner instruction used before the full scaffold review template. [PITH_FULL_IMAGE:figures/full_fig_p012_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Full policy injection template used in the supported training condition. [PITH_FULL_IMAGE:figures/full_fig_p013_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Scaffold review message template supplied to the refiner. [PITH_FULL_IMAGE:figures/full_fig_p013_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Mode-specific scaffold review instructions. [PITH_FULL_IMAGE:figures/full_fig_p014_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Allowed scaffold edit tool calls returned by the refiner. [PITH_FULL_IMAGE:figures/full_fig_p014_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Formal macro-review system message used by the scaffold refiner. [PITH_FULL_IMAGE:figures/full_fig_p015_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Fixed-policy ALFWorld replay cases. Green checks denote successful trajectories and red crosses denote failures. [PITH_FULL_IMAGE:figures/full_fig_p018_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 3 canonical work pages · 2 internal anchors

  1. [1]

    arXiv preprint arXiv:1707.06347 , year =

    Proximal Policy Optimization Algorithms , author =. arXiv preprint arXiv:1707.06347 , year =. doi:10.48550/arXiv.1707.06347 , url =

  2. [2]

    arXiv preprint arXiv:1707.06347 , year=

    Proximal policy optimization algorithms , author=. arXiv preprint arXiv:1707.06347 , year=

  3. [3]

    arXiv preprint arXiv:1802.07569 , year =

    Continual Lifelong Learning with Neural Networks: A Review , author =. arXiv preprint arXiv:1802.07569 , year =. doi:10.48550/arXiv.1802.07569 , url =

  4. [4]

    arXiv preprint arXiv:2010.03768 , year =

    Shridhar, Mohit and Yuan, Xingdi and C. arXiv preprint arXiv:2010.03768 , year =. doi:10.48550/arXiv.2010.03768 , url =

  5. [5]

    2022 , doi =

    Yao, Shunyu and Chen, Howard and Yang, John and Narasimhan, Karthik , journal =. 2022 , doi =

  6. [6]

    Proceedings of the 26th Annual International Conference on Machine Learning , pages =

    Curriculum Learning , author =. Proceedings of the 26th Annual International Conference on Machine Learning , pages =. 2009 , doi =

  7. [7]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

    When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories , author =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2023 , doi =

  8. [8]

    Transactions of the Association for Computational Linguistics , volume =

    Natural Questions: A Benchmark for Question Answering Research , author =. Transactions of the Association for Computational Linguistics , volume =. 2019 , doi =

  9. [9]

    and Zettlemoyer, Luke , booktitle =

    Joshi, Mandar and Choi, Eunsol and Weld, Daniel S. and Zettlemoyer, Luke , booktitle =. 2017 , doi =

  10. [10]

    and Salakhutdinov, Ruslan and Manning, Christopher D

    Yang, Zhilin and Qi, Peng and Zhang, Saizheng and Bengio, Yoshua and Cohen, William W. and Salakhutdinov, Ruslan and Manning, Christopher D. , booktitle =. 2018 , doi =

  11. [11]

    Constructing A Multi-hop

    Ho, Xanh and Nguyen, Anh-Khoa Duong and Sugawara, Saku and Aizawa, Akiko , booktitle =. Constructing A Multi-hop. 2020 , doi =

  12. [12]

    2022 , doi =

    Trivedi, Harsh and Balasubramanian, Niranjan and Khot, Tushar and Sabharwal, Ashish , journal =. 2022 , doi =

  13. [13]

    arXiv preprint arXiv:2210.03350 , year =

    Measuring and Narrowing the Compositionality Gap in Language Models , author =. arXiv preprint arXiv:2210.03350 , year =. doi:10.48550/arXiv.2210.03350 , url =

  14. [14]

    Advances in Neural Information Processing Systems , volume =

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author =. Advances in Neural Information Processing Systems , volume =. 2022 , url =

  15. [15]

    Retrieval-Augmented Generation for Knowledge-Intensive

    Lewis, Patrick and Perez, Ethan and Piktus, Aleksandra and Petroni, Fabio and Karpukhin, Vladimir and Goyal, Naman and K. Retrieval-Augmented Generation for Knowledge-Intensive. Advances in Neural Information Processing Systems , volume =. 2020 , url =

  16. [16]

    arXiv preprint arXiv:2212.03533 , year =

    Text Embeddings by Weakly-Supervised Contrastive Pre-training , author =. arXiv preprint arXiv:2212.03533 , year =. doi:10.48550/arXiv.2212.03533 , url =

  17. [17]

    Billion-Scale Similarity Search with

    Johnson, Jeff and Douze, Matthijs and J. Billion-Scale Similarity Search with. arXiv preprint arXiv:1702.08734 , year =. doi:10.48550/arXiv.1702.08734 , url =

  18. [18]

    Transactions of the Association for Computational Linguistics , volume =

    Lost in the Middle: How Language Models Use Long Contexts , author =. Transactions of the Association for Computational Linguistics , volume =. 2024 , doi =

  19. [19]

    2022 , doi =

    Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan , journal =. 2022 , doi =

  20. [20]

    arXiv preprint arXiv:2303.11366 , year =

    Reflexion: Language Agents with Verbal Reinforcement Learning , author =. arXiv preprint arXiv:2303.11366 , year =. doi:10.48550/arXiv.2303.11366 , url =

  21. [21]

    2023 , doi =

    Li, Guohao and Hammoud, Hasan Abed Al Kader and Itani, Hani and Khizbullin, Dmitrii and Ghanem, Bernard , journal =. 2023 , doi =

  22. [22]

    2023 , doi =

    Zhong, Wanjun and Guo, Lianghong and Gao, Qiqi and Ye, He and Wang, Yanlin , journal =. 2023 , doi =

  23. [23]

    arXiv preprint arXiv:2305.16291 , year =

    Voyager: An Open-Ended Embodied Agent with Large Language Models , author =. arXiv preprint arXiv:2305.16291 , year =. doi:10.48550/arXiv.2305.16291 , url =

  24. [24]

    2023 , doi =

    Zhao, Andrew and Huang, Daniel and Xu, Quentin and Lin, Matthieu and Liu, Yong-Jin and Huang, Gao , journal =. 2023 , doi =

  25. [25]

    and Burger, Doug and Wang, Chi , journal =

    Wu, Qingyun and Bansal, Gagan and Zhang, Jieyu and Wu, Yiran and Li, Beibin and Zhu, Erkang and Jiang, Li and Zhang, Xiaoyun and Zhang, Shaokun and Liu, Jiale and Awadallah, Ahmed Hassan and White, Ryen W. and Burger, Doug and Wang, Chi , journal =. 2023 , doi =

  26. [26]

    Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, Y. K. and Wu, Y. and Guo, Daya , journal =. 2024 , doi =

  27. [27]

    Advances in Neural Information Processing Systems , volume =

    Training Language Models to Follow Instructions with Human Feedback , author =. Advances in Neural Information Processing Systems , volume =. 2022 , url =

  28. [28]

    Equipping Agents for the Real World with Agent Skills , year =

  29. [29]

    Back to Basics: Revisiting

    Ahmadian, Arash and Cremer, Chris and Gall. Back to Basics: Revisiting. arXiv preprint arXiv:2402.14740 , year =. doi:10.48550/arXiv.2402.14740 , url =

  30. [30]

    arXiv preprint arXiv:2409.07429 , year =

    Agent Workflow Memory , author =. arXiv preprint arXiv:2409.07429 , year =. doi:10.48550/arXiv.2409.07429 , url =

  31. [31]

    2025 , doi =

    Chhikara, Prateek and Khant, Dev and Aryan, Saket and Singh, Taranjeet and Yadav, Deshraj , journal =. 2025 , doi =

  32. [32]

    2025 , doi =

    Wang, Zihan and Wang, Kangrui and Wang, Qineng and Zhang, Pingyue and Li, Linjie and Yang, Zhengyuan and Jin, Xing and Yu, Kefan and Nguyen, Minh Nhat and Liu, Licheng and Gottlieb, Eli and Lu, Yiping and Cho, Kyunghyun and Wu, Jiajun and Fei-Fei, Li and Wang, Lijuan and Choi, Yejin and Li, Manling , journal =. 2025 , doi =

  33. [33]

    Group-in-Group Policy Optimization for

    Feng, Lang and Xue, Zhenghai and Liu, Tingcong and An, Bo , journal =. Group-in-Group Policy Optimization for. 2025 , doi =

  34. [34]

    2025 , doi =

    Jin, Bowen and Zeng, Hansi and Yue, Zhenrui and Yoon, Jinsung and Arik, Sercan and Wang, Dong and Zamani, Hamed and Han, Jiawei , journal =. 2025 , doi =

  35. [35]

    2025 , doi =

    Li, Xiaoxi and Dong, Guanting and Jin, Jiajie and Zhang, Yuyao and Zhou, Yujia and Zhu, Yutao and Zhang, Peitian and Dou, Zhicheng , journal =. 2025 , doi =

  36. [36]

    2025 , doi =

    Sun, Hao and Qiao, Zile and Guo, Jiayan and Fan, Xuanbo and Hou, Yingyan and Jiang, Yong and Xie, Pengjun and Zhang, Yan and Huang, Fei and Zhou, Jingren , journal =. 2025 , doi =

  37. [37]

    2025 , doi =

    Wang, Ziliang and Zheng, Xuhui and An, Kang and Ouyang, Cijun and Cai, Jialu and Wang, Yuhang and Wu, Yichao , journal =. 2025 , doi =

  38. [38]

    2025 , doi =

    Zhang, Guibin and Fu, Muxin and Wan, Guancheng and Yu, Miao and Wang, Kun and Yan, Shuicheng , journal =. 2025 , doi =

  39. [39]

    Tang, Xiangru and Qin, Tianrui and Peng, Tianhao and Zhou, Ziyang and Shao, Daniel and Du, Tingting and Wei, Xinming and Xia, Peng and Wu, Fang and Zhu, He and Zhang, Ge and Liu, Jiaheng and Wang, Xingyao and Hong, Sirui and Wu, Chenglin and Cheng, Hao and Wang, Chi and Zhou, Wangchunshu , journal =. Agent. 2025 , doi =

  40. [40]

    2025 , doi =

    Wang, Yu and Chen, Xia , journal =. 2025 , doi =

  41. [41]

    arXiv preprint arXiv:2507.21046 , year =

    A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence , author =. arXiv preprint arXiv:2507.21046 , year =. doi:10.48550/arXiv.2507.21046 , url =

  42. [42]

    2025 , doi =

    Fang, Runnan and Liang, Yuan and Wang, Xiaobin and Wu, Jialong and Qiao, Shuofei and Xie, Pengjun and Huang, Fei and Chen, Huajun and Zhang, Ningyu , journal =. 2025 , doi =

  43. [43]

    and Daruki, Samira and Tang, Xiangru and Tirumalashetty, Vishy and Lee, George and Rofouei, Mahsan and Lin, Hangfei and Han, Jiawei and Lee, Chen-Yu and Pfister, Tomas , journal =

    Ouyang, Siru and Yan, Jun and Hsu, I-Hung and Chen, Yanfei and Jiang, Ke and Wang, Zifeng and Han, Rujun and Le, Long T. and Daruki, Samira and Tang, Xiangru and Tirumalashetty, Vishy and Lee, George and Rofouei, Mahsan and Lin, Hangfei and Han, Jiawei and Lee, Chen-Yu and Pfister, Tomas , journal =. 2025 , doi =

  44. [44]

    2025 , doi =

    Wang, Yu and Takanobu, Ryuichi and Liang, Zhiqi and Mao, Yuzhen and Hu, Yuanzhe and McAuley, Julian and Wu, Xiaojian , journal =. 2025 , doi =

  45. [45]

    2025 , doi =

    Wu, Rong and Wang, Xiaoman and Mei, Jianbiao and Cai, Pinlong and Fu, Daocheng and Yang, Cheng and Wen, Licheng and Yang, Xuemeng and Shen, Yufan and Wang, Yuxin and Shi, Botian , journal =. 2025 , doi =

  46. [46]

    2025 , doi =

    Xia, Peng and Zeng, Kaide and Liu, Jiaqi and Qin, Can and Wu, Fang and Zhou, Yiyang and Xiong, Caiming and Yao, Huaxiu , journal =. 2025 , doi =

  47. [47]

    2025 , doi =

    Liu, Jiaqi and Xiong, Kaiwen and Xia, Peng and Zhou, Yiyang and Ji, Haonian and Feng, Lu and Han, Siwei and Ding, Mingyu and Yao, Huaxiu , journal =. 2025 , doi =

  48. [48]

    and Wang, Chi and Chen, Shuo and Pereira, Fernando and Kang, Wang-Cheng and Cheng, Derek Zhiyuan , journal =

    Wei, Tianxin and Sachdeva, Noveen and Coleman, Benjamin and He, Zhankui and Bei, Yuanchen and Ning, Xuying and Ai, Mengting and Li, Yunzhe and He, Jingrui and Chi, Ed H. and Wang, Chi and Chen, Shuo and Pereira, Fernando and Kang, Wang-Cheng and Cheng, Derek Zhiyuan , journal =. 2025 , doi =

  49. [49]

    2025 , doi =

    Zhang, Guibin and Ren, Haotian and Zhan, Chong and Zhou, Zhenhong and Wang, Junhao and Zhu, He and Zhou, Wangchunshu and Yan, Shuicheng , journal =. 2025 , doi =

  50. [50]

    Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General

    Zhou, Yang and Li, Sunzhu and Liu, Shunyu and Fang, Wenkai and Zhang, Kongcheng and Zhao, Jiale and Yang, Jingwen and Zhou, Yihe and Lv, Jianwei and Zheng, Tongya and Lu, Hengtong and Chen, Wei and Xie, Yan and Song, Mingli , journal =. Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General. 2025 , doi =

  51. [51]

    2025 , doi =

    Jiang, Guochao and Feng, Wenfeng and Quan, Guofeng and Hao, Chuzhan and Zhang, Yuewei and Liu, Guohua and Wang, Hao , journal =. 2025 , doi =

  52. [52]

    2026 , doi =

    Xia, Peng and Chen, Jianwen and Wang, Hanyang and Liu, Jiaqi and Zeng, Kaide and Wang, Yu and Han, Siwei and Zhou, Yiyang and Zhao, Xujiang and Chen, Haifeng and Zheng, Zeyu and Xie, Cihang and Yao, Huaxiu , journal =. 2026 , doi =

  53. [53]

    2026 , doi =

    Liu, Jiaqi and Su, Yaofeng and Xia, Peng and Han, Siwei and Zheng, Zeyu and Xie, Cihang and Ding, Mingyu and Yao, Huaxiu , journal =. 2026 , doi =

  54. [54]

    2026 , doi =

    Zhang, Shengtao and Wang, Jiaqian and Zhou, Ruiwen and Liao, Junwei and Feng, Yuchen and Li, Zhuo and Zheng, Yujie and Zhang, Weinan and Wen, Ying and Li, Zhiyu and Xiong, Feiyu and Qi, Yutao and Tang, Bo and Wen, Muning , journal =. 2026 , doi =

  55. [55]

    arXiv preprint arXiv:2601.12538 , year =

    Agentic Reasoning for Large Language Models , author =. arXiv preprint arXiv:2601.12538 , year =. doi:10.48550/arXiv.2601.12538 , url =

  56. [56]

    2026 , doi =

    Gao, Haowen and Zhang, Zhenyu and Pang, Liang and Guo, Fangda and Dou, Hongjian and Lv, Guannan and Liu, Shaoguo and Gao, Tingting and Shen, Huawei and Cheng, Xueqi , journal =. 2026 , doi =

  57. [57]

    2026 , doi =

    Lu, Zhengxi and Yao, Zhiyuan and Wu, Jinyang and Han, Chengcheng and Gu, Qi and Cai, Xunliang and Lu, Weiming and Xiao, Jun and Zhuang, Yueting and Shen, Yongliang , journal =. 2026 , doi =

  58. [58]

    2026 , doi =

    Lin, Hongxiang and Kuai, Zhirui and Xue, Erpeng and Wang, Lei , journal =. 2026 , doi =

  59. [59]

    2026 , doi =

    Ouyang, Siru and Yan, Jun and Chen, Yanfei and Han, Rujun and Wang, Zifeng and Mishra, Bhavana Dalvi and Meng, Rui and Li, Chun-Liang and Jiao, Yizhu and Zha, Kaiwen and Shen, Maohao and Tirumalashetty, Vishy and Lee, George and Han, Jiawei and Pfister, Tomas and Lee, Chen-Yu , journal =. 2026 , doi =

  60. [60]

    2026 , doi =

    Shi, Yaorui and Chen, Yuxin and Lu, Zhengxi and Miao, Yuchun and Liu, Shugui and Gu, Qi and Cai, Xunliang and Wang, Xiang and Zhang, An , journal =. 2026 , doi =

  61. [61]

    arXiv preprint arXiv:2605.10923 , year =

    Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning , author =. arXiv preprint arXiv:2605.10923 , year =. doi:10.48550/arXiv.2605.10923 , url =

  62. [62]

    2026 , doi =

    Pu, Hongji and Song, Xinyuan and Zhao, Liang , journal =. 2026 , doi =

  63. [63]

    doi:10.48550/arXiv.2412.15115 , url =

    arXiv preprint arXiv:2412.15115 , year =. doi:10.48550/arXiv.2412.15115 , url =

  64. [64]

    Learning Agentic Policy from Action Guidance

    Learning Agentic Policy from Action Guidance , author =. arXiv preprint arXiv:2605.12004 , year =. doi:10.48550/arXiv.2605.12004 , url =

  65. [65]

    2026 , doi =

    He, Zhongyu and Li, Yuanfan and Huang, Fei and Chen, Tianyu and Chen, Siyuan and Li, Xingyang and Yu, Meng Hsuan and Liu, Xiangrong and Wei, Leyi and Pan, Lu and Zeng, Ke and Cai, Xunliang , journal =. 2026 , doi =

  66. [66]

    Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents

    Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents , author =. arXiv preprint arXiv:2606.04815 , year =. doi:10.48550/arXiv.2606.04815 , url =

  67. [67]

    arXiv preprint arXiv:2606.15390 , year =

    Not All Skills Help: Measuring and Repairing Agent Knowledge , author =. arXiv preprint arXiv:2606.15390 , year =. doi:10.48550/arXiv.2606.15390 , url =

  68. [68]

    2026 , doi =

    Wu, Jinyang and Yang, Shuo and Lu, Zhengxi and Zhang, Fan and Shen, Yuhao and Feng, Lang and Luo, Haoran and Lian, Zheng and Zhang, Shuai and Wen, Zhengqi and Tao, Jianhua , journal =. 2026 , doi =

  69. [69]

    2026 , doi =

    Tu, Songjun and Xu, Chengdong and Zhang, Qichao and Ma, Yiwen and Zhang, Yaocheng and Li, Linjing and Li, Dong and Lan, Xiangyuan and Zhao, Dongbin , journal =. 2026 , doi =