{"total":14,"items":[{"citing_arxiv_id":"2607.06935","ref_index":121,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Mathematical methods of reinforcement learning","primary_cat":"math.OC","submitted_at":"2026-07-08T02:57:22+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":0.0,"formal_verification":"none","one_line_summary":"A survey unifying the operator-theoretic, probabilistic, and optimization-based mathematical structures underlying modern reinforcement learning algorithms.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12191","ref_index":262,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application","primary_cat":"cs.CL","submitted_at":"2026-06-10T15:15:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Trajectory-CentricOffline Evolution(§6.3) Task Synthesis (§6.3.1)e.g.,BAGEL [247], OS-Genesis [248], Insta [249], APIGen-MT [250], WebShaper [251],WebWatcher [252], WebExplorer [253], AutoPlay [254], CRMWeaver [255], WebLeaper [256],etc. Trajectory Synthesis (§6.3.2)e.g.,ToolAlpaca [257], Lingma SWE-GPT [258], Aguvis [259], FlowReasoner [260],WebSynthesis [261], AgentFold [262], ToolACE-MCP [263], ProAct [264],etc. Trajectory Refinement (§6.3.3)e.g.,Toolformer [265], ETO [266], Self-Improvement [267], GUI-Reflection [268], TiG [269],AgentFrontier [270], WebSTAR [271], SynthAgent [272], TopoCurate [273],etc. Exploration-CentricOnline Evolution(§6.4) Reasoning Structure (§6.4.1)e.g.,DeepRetrieval [274], Search-R1 [275], AutoRefine [276], SEEA-R1 [277], M3-Agent [278],Video-Thinker [279], ReSearch [280],etc."},{"citing_arxiv_id":"2606.09138","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Claw-R1: A Step-Level Data Middleware System for Agentic Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-06-08T07:35:18+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Claw-R1 provides a Gateway Server and Data Pool to manage step-level agent interaction traces as structured data assets for agentic RL training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22138","ref_index":48,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Efficient Agentic Reasoning Through Self-Regulated Simulative Planning","primary_cat":"cs.AI","submitted_at":"2026-05-21T08:11:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SR²AM achieves competitive Pass@1 accuracy on diverse tasks with 25.8-95.3% fewer reasoning tokens than much larger models by using self-regulated simulative planning trained via supervised learning and RL.","context_count":1,"top_context_role":"baseline","top_context_polarity":"baseline","context_text":"mechanismacrossdiversetasks[ 113];andself-regulation(SystemIII)thatdecideswhenandhowdeeplyto planthroughalearnedconfigurator,muchlikehumansmodulatedeliberationbasedonurgency,uncertainty, andcomplexity[ 42]. Prioreffortseachaddressespartofthisproblem,whetheritbecontrollingreasoning amount[e.g., 52,99],selectingexecutionmodeattaskonset[ 13,40],distillingrule-basedworkflows[ 48],or usingworldmodelsforobligatorysimulation[e.g., 33,20]. Nonecombinesallthreeintoaunifiedarchitecture. Inthispaper,westudywhethertheSystemI+II+IIIdecompositionyieldsbetteraccuracy-efficiencytradeoffs thanunregulatedorpartiallyregulatedalternatives,inthesettingoflanguage-basedinteractivereasoning[e.g., 54,64]. Totestthis,wedevelopSR 2AM(Self-RegulatedSimulativeReasoningAgenticLLM),whichimple-"},{"citing_arxiv_id":"2605.15224","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ICRL: Learning to Internalize Self-Critique with Reinforcement Learning","primary_cat":"cs.AI","submitted_at":"2026-05-13T08:50:05+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ICRL uses joint RL training of solver and critic with distribution-calibration re-weighting and role-wise advantage estimation to internalize critique into unassisted LLM performance, yielding 6.4-point gains on agentic tasks and 7.0 on math reasoning with Qwen3 models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.07725","ref_index":12,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SOD: Step-wise On-policy Distillation for Small Language Model Agents","primary_cat":"cs.CL","submitted_at":"2026-05-08T13:30:42+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A step-wise reweighting of on-policy distillation, based on per-step student-teacher divergence, improves tool-integrated reasoning in 0.6B and 1.7B language models.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Agentdistill: Training-free agent distillation with generalizable mcp boxes.arXiv preprint arXiv:2506.14728, 2025. [11] Yi Yao, He Zhu, Piaohong Wang, Jincheng Ren, Xinlong Yang, Qianben Chen, Xiaowan Li, Dingfeng Shi, Jiaxian Li, Qiexiang Wang, et al. O-researcher: An open ended deep research model via multi-agent distillation and agentic rl.arXiv preprint arXiv:2601.03743, 2026. [12] Weizhen Li, Jianbo Lin, Zhuosong Jiang, Jingyi Cao, Xinpeng Liu, Jiayu Zhang, Zhenqiang Huang, Qianben Chen, Weichen Sun, Qiexiang Wang, et al. Chain-of-agents: End-to-end agent foundation models via multi-agent distillation and agentic rl.arXiv preprint arXiv:2508.13167, 2025. [13] Cheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan"},{"citing_arxiv_id":"2605.04496","ref_index":40,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States","primary_cat":"cs.CL","submitted_at":"2026-05-06T04:55:59+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SCOUT achieves state-of-the-art long-text understanding with up to 8x lower token use by actively foraging for sparse query-relevant information and updating a compact provenance-grounded epistemic state.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.01347","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate","primary_cat":"cs.CL","submitted_at":"2026-05-02T09:41:37+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"MAD-OPD recasts on-policy distillation teachers as a debating collective to supply better supervision, lifting agentic and code performance over single-teacher OPD across multiple model sizes.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"[30] extend OPD to leverageprivileged informationthat is visible only to a single teacher at training time. These methods all rely on a single teacher with task-agnostic divergence; MAD-OPD instead provides debate-generated privileged information and a task-adaptive divergence. Multi-Agent Systems.Multi-agent systems (MAS) coordinate specialized agents to solve complex tasks; Chain-of-Agents [29] distills an entire MAS pipeline into a single foundation model. The MAS instantiation closest to our setting is multi-agent debate (MAD), where iterative argumentation among agents yields anemergent collective intelligencethat improves reasoning beyond any individual agent's capability [6, 9, 11, 23]. Existing MAD methods, however, consume debate outputs only"},{"citing_arxiv_id":"2605.08124","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence","primary_cat":"cs.DC","submitted_at":"2026-04-29T12:24:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A vision paper outlining a two-pronged research agenda for scaling mobile agents from isolated devices to distributed intelligent systems.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.18292","ref_index":46,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence","primary_cat":"cs.AI","submitted_at":"2026-04-20T14:01:10+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Agent-World autonomously synthesizes verifiable real-world tasks and uses continuous self-evolution to train 8B and 14B agents that outperform proprietary models on 23 benchmarks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"[45] Kuan Li, Zhongwang Zhang, Huifeng Yin, Liwen Zhang, Litu Ou, Jialong Wu, Wenbiao Yin, Baixuan Li, Zhengwei Tao, Xinyu Wang, Weizhou Shen, Junkai Zhang, Dingchu Zhang, Xixi Wu, Yong Jiang, Ming Yan, Pengjun Xie, Fei Huang, and Jingren Zhou. Websailor: Navigating super-human reasoning for web agent.CoRR, abs/2507.02592, 2025. doi: 10.48550/ARXIV.2507.02592. URLhttps://doi.org/10.48550/arXiv.2507.02592. [46] Weizhen Li, Jianbo Lin, Zhuosong Jiang, Jingyi Cao, Xinpeng Liu, Jiayu Zhang, Zhenqiang Huang, Qianben Chen, Weichen Sun, Qiexiang Wang, Hongxuan Lu, Tianrui Qin, Chenghao Zhu, Yi Yao, Shuying Fan, Xiaowan Li, Tiannan Wang, Pai Liu, King Zhu, He Zhu, Dingfeng Shi, Piaohong Wang, Yeyi Guan, Xiangru Tang, Minghao Liu, Yuchen Eleanor Jiang, Jian Yang, Jiaheng Liu, Ge Zhang, and Wangchunshu Zhou."},{"citing_arxiv_id":"2604.17931","ref_index":15,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent","primary_cat":"cs.AI","submitted_at":"2026-04-20T08:11:09+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Injecting 1% targeted synthetic data into GPT-2's pre-training substantially improves performance on 8 of 9 failing BLiMP grammatical paradigms, indicating data scarcity causes formal linguistic failures.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.06170","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework","primary_cat":"cs.CL","submitted_at":"2026-04-07T17:59:58+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Paper Circle is an open-source multi-agent system that retrieves papers via offline and online sources, applies multi-criteria scoring and diversity ranking, and converts papers into typed knowledge graphs for structured analysis and question answering.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2511.11793","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling","primary_cat":"cs.CL","submitted_at":"2025-11-14T18:52:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MiroThinker shows that scaling agent-environment interactions via reinforcement learning lets a 72B open-source model reach up to 81.9% on GAIA and approach commercial performance on research benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.08827","ref_index":283,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Survey of Reinforcement Learning for Large Reasoning Models","primary_cat":"cs.CL","submitted_at":"2025-09-10T17:59:43+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"A survey compiling RL methods, challenges, data resources, and applications for enhancing reasoning in large language models and large reasoning models since DeepSeek-R1.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}