{"total":14,"items":[{"citing_arxiv_id":"2607.07989","ref_index":64,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems","primary_cat":"cs.CR","submitted_at":"2026-07-08T23:33:49+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AgentLocate localizes multi-agent LLM failures to a responsible agent and earliest decisive step via judge hypotheses, confidence-weighted multi-evaluator verification, and LoRA refinement.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06935","ref_index":118,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Mathematical methods of reinforcement learning","primary_cat":"math.OC","submitted_at":"2026-07-08T02:57:22+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":0.0,"formal_verification":"none","one_line_summary":"A survey unifying the operator-theoretic, probabilistic, and optimization-based mathematical structures underlying modern reinforcement learning algorithms.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.23537","ref_index":9,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SQLConductor: Search-to-Policy Learning for Step-wise Text-to-SQL Orchestration","primary_cat":"cs.DB","submitted_at":"2026-06-22T16:13:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SQLConductor uses Search-to-Policy Learning with MCTS, stability-weighted SFT, and curriculum RL to train a compact policy for adaptive step-wise Text-to-SQL orchestration, reporting 73.2% EX on BIRD-Dev.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12191","ref_index":279,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application","primary_cat":"cs.CL","submitted_at":"2026-06-10T15:15:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Trajectory Refinement (§6.3.3)e.g.,Toolformer [265], ETO [266], Self-Improvement [267], GUI-Reflection [268], TiG [269],AgentFrontier [270], WebSTAR [271], SynthAgent [272], TopoCurate [273],etc. Exploration-CentricOnline Evolution(§6.4) Reasoning Structure (§6.4.1)e.g.,DeepRetrieval [274], Search-R1 [275], AutoRefine [276], SEEA-R1 [277], M3-Agent [278],Video-Thinker [279], ReSearch [280],etc. Reward Shaping (§6.4.2)e.g.,Agent-R1 [281], ToolRL [8], Chain-of-Agents [245], VRAG-RL [282], GDPO [283],FlowSteer [284], Tool-N1 [285], ToolOrchestra [286],etc. AlgorithmicOptimization (§6.4.3)e.g.,RAGEN [287], ZeroSearch [288], EvolveSearch [289], GiGPO [290], ARPO [291],MobileGUI-RL [292], SPEAR [293], VAGEN [294], SeeUPO [295],etc."},{"citing_arxiv_id":"2606.08656","ref_index":35,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"From Player to Master: Enhancing Test-Time Learning of LLM Agents via Reinforcement Learning over Memory","primary_cat":"cs.CL","submitted_at":"2026-06-07T14:53:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MemoPilot trains memory updates for LLM agents via multi-turn GRPO on RPS and poker, achieving top Elo scores and outperforming baselines including DeepSeek-V3.2.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29790","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems","primary_cat":"cs.MA","submitted_at":"2026-05-28T11:40:16+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Meta-Team is a collaborative self-evolution framework that turns multi-agent execution experience into reusable improvements at agent, coordination, and team levels, outperforming baselines on six benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22505","ref_index":34,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Towards Direct Evaluation of Harness Optimizers via Priority Ranking","primary_cat":"cs.AI","submitted_at":"2026-05-21T13:55:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Priority ranking offers a low-cost direct evaluation for harness optimizers that correlates with their real multi-step optimization performance, supported by the Shor dataset of 182 scenarios.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.14483","ref_index":22,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LEMON: Learning Executable Multi-Agent Orchestration via Counterfactual Reinforcement Learning","primary_cat":"cs.AI","submitted_at":"2026-05-14T07:24:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LEMON trains an LLM orchestrator with counterfactual-augmented GRPO to produce deployable multi-agent specifications that reach state-of-the-art results on six reasoning and coding benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.08401","ref_index":18,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AIPO: Learning to Reason from Active Interaction","primary_cat":"cs.CL","submitted_at":"2026-05-08T19:06:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AIPO adds active multi-agent consultation (Verify, Knowledge, Reasoning agents) plus custom importance sampling to RLVR training so LLMs expand their reasoning boundary and then operate without the agents.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Stone, and et al. The llama 3 herd of models.CoRR, abs/2407.21783, 2024. 1, 4.1 [17] Yuqian Fu, Tinghong Chen, Jiajun Chai, Xihuai Wang, Songjun Tu, Guojun Yin, Wei Lin, Qichao Zhang, Yuanheng Zhu, and Dongbin Zhao. SRFT: A single-stage method with super- vised and reinforcement fine-tuning for reasoning.CoRR, abs/2506.19767, 2025. 1, 1, 2, 5, D.1 [18] Hongcheng Gao, Yue Liu, Yufei He, Longxu Dou, Chao Du, Zhijie Deng, Bryan Hooi, Min Lin, and Tianyu Pang. Flowreasoner: Reinforcing query-level meta-agents.CoRR, abs/2504.15257, 2025. 3.1 [19] Etash Kumar Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, Ashima Suvarna, Benjamin Feuer,"},{"citing_arxiv_id":"2604.17503","ref_index":11,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SkillGraph: Self-Evolving Multi-Agent Collaboration with Multimodal Graph Topology","primary_cat":"cs.AI","submitted_at":"2026-04-19T15:46:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SkillGraph jointly evolves agent skills and collaboration topologies in multi-agent vision-language systems using a multimodal graph transformer and a skill designer, yielding consistent performance gains on benchmarks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"adaptive topology prediction; and MASS [55] revealed that prompt and topology 4 Z. Nie et al. search are mutually reinforcing. Preceding these learnable methods, DyLAN [29] and DSPy [20] laid important groundwork through dynamic team optimization and compiled agent pipelines, respectively. More recent work trains meta-agents to generate query-conditioned workflows end-to-end via RL [8,11,14], or enables the graph itself to self-evolve through test-time feedback [16,42]. Despite this progress,acriticalgappersists:everyexistingmethodconstructstheagentgraph purely from text, leaving the visual content of the query outside the topology- prediction loop. SkillGraph closes this gap by conditioning the graph transformer jointly on question semantics and per-agent image attention, so that the inferred"},{"citing_arxiv_id":"2603.25111","ref_index":13,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SEVerA: Verified Synthesis of Self-Evolving Agents","primary_cat":"cs.LG","submitted_at":"2026-03-26T07:32:20+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"SEVerA uses Formally Guarded Generative Models and a three-stage Search-Verification-Learning process to synthesize self-evolving agents that satisfy hard formal constraints while improving task performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2510.05746","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems","primary_cat":"cs.AI","submitted_at":"2025-10-07T10:04:48+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ARM evolves specialized reasoning modules from basic CoT via tree search to serve as reusable components in multi-agent systems that generalize across models and domains without per-task re-optimization.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2508.07407","ref_index":27,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems","primary_cat":"cs.AI","submitted_at":"2025-08-10T16:07:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A comprehensive review of self-evolving AI agents that improve themselves over time, organized via a framework of inputs, agent system, environment, and optimizers, with domain-specific and safety discussions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2507.21046","ref_index":274,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence","primary_cat":"cs.AI","submitted_at":"2025-07-28T17:59:05+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"The paper delivers the first systematic review of self-evolving agents, structured around what components evolve, when adaptation occurs, and how it is implemented.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}