{"total":17,"items":[{"citing_arxiv_id":"2607.07820","ref_index":4,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment","primary_cat":"cs.CL","submitted_at":"2026-07-08T18:03:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A verifiable offline Wikipedia tool environment plus iterative scaffold-to-ReAct self-distillation lets a 9B agent reach competitive BrowseComp/GAIA/HotpotQA scores without stronger-model distillation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.20724","ref_index":6,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration","primary_cat":"cs.AI","submitted_at":"2026-06-16T23:00:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Parallel WebBench reveals GRPO training raises web agent completion to 96% but leaves a large correctness gap from context-bound loops, premature termination, and synthesis collapse.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.12191","ref_index":271,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application","primary_cat":"cs.CL","submitted_at":"2026-06-10T15:15:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":",BAGEL [247], OS-Genesis [248], Insta [249], APIGen-MT [250], WebShaper [251],WebWatcher [252], WebExplorer [253], AutoPlay [254], CRMWeaver [255], WebLeaper [256],etc. Trajectory Synthesis (§6.3.2)e.g.,ToolAlpaca [257], Lingma SWE-GPT [258], Aguvis [259], FlowReasoner [260],WebSynthesis [261], AgentFold [262], ToolACE-MCP [263], ProAct [264],etc. Trajectory Refinement (§6.3.3)e.g.,Toolformer [265], ETO [266], Self-Improvement [267], GUI-Reflection [268], TiG [269],AgentFrontier [270], WebSTAR [271], SynthAgent [272], TopoCurate [273],etc. Exploration-CentricOnline Evolution(§6.4) Reasoning Structure (§6.4.1)e.g.,DeepRetrieval [274], Search-R1 [275], AutoRefine [276], SEEA-R1 [277], M3-Agent [278],Video-Thinker [279], ReSearch [280],etc. Reward Shaping (§6.4.2)e.g.,Agent-R1 [281], ToolRL [8], Chain-of-Agents [245], VRAG-RL [282], GDPO [283],FlowSteer [284], Tool-N1 [285], ToolOrchestra [286],etc."},{"citing_arxiv_id":"2606.11337","ref_index":72,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Can AI Agents Synthesize Scientific Conclusions?","primary_cat":"cs.AI","submitted_at":"2026-06-09T18:16:04+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A new benchmark and clean-room harness show frontier AI agents reach only 0.337 factual F1 when synthesizing conclusions from scientific evidence.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.31529","ref_index":43,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence","primary_cat":"cs.CV","submitted_at":"2026-05-29T16:43:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SVI-Bench provides 35K hours of sports video with 9 tasks across four cognitive levels, revealing models drop from ~74% on action QA to 5% on agentic evidence integration.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.30824","ref_index":16,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward","primary_cat":"cs.AI","submitted_at":"2026-05-29T04:18:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DecomposeR represents research plans as typed DAGs and uses two-stage planner-then-answerer RL to improve long-form research performance by 5.1-8.0 points over baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29697","ref_index":14,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling","primary_cat":"cs.AI","submitted_at":"2026-05-28T09:57:12+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"GDCR assigns step-level rewards via distance to the answer node in a training-time ER graph and SAPO combines these with trajectory advantages for credit assignment in agentic search.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.28003","ref_index":2,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ResearchMath-14K: Scaling Research-Level Mathematics via Agents","primary_cat":"cs.CL","submitted_at":"2026-05-27T05:54:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"The authors release ResearchMath-14k, the largest dataset of research-level math problems, and demonstrate that agent-filtered reasoning trajectories from open models improve fine-tuned Qwen3 models by 9.2 points on average.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.26494","ref_index":3,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence","primary_cat":"cs.AI","submitted_at":"2026-05-26T03:16:11+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A 10-billion-parameter-activated Mixture-of-Experts model family is reported to match or approach larger frontier systems on several agentic coding, search, and office benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22138","ref_index":49,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Efficient Agentic Reasoning Through Self-Regulated Simulative Planning","primary_cat":"cs.AI","submitted_at":"2026-05-21T08:11:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SR²AM achieves competitive Pass@1 accuracy on diverse tasks with 25.8-95.3% fewer reasoning tokens than much larger models by using self-regulated simulative planning trained via supervised learning and RL.","context_count":1,"top_context_role":"baseline","top_context_polarity":"baseline","context_text":"6[ 122],GPT-OSS-120B-high[ 65],Qwen3- 8B[79],Qwen3-30B-A3B-Thinking-2507[78],andQwen3-235B-A22B-Thinking-2507[76]). • UnregulatedDeliberation: agenticLLMstrainedtoreasonandactwithunconstrainedreasoning(Tongyi- DeepResearch [97], MiroThinker-v1.5-30B [101], WebSailor-(7B/32B) [110], ASearcher-Web-(7B/QWQ- v2)[90],SimpleTIR-(7B/32B)[103],andWebExplorer-8B[49]) • Partially-RegulatedDeliberation: agenticLLMsrealizingasubsetofourproposedthree-systemdecom- position(A 2FM[13]forModeRoutingandAFM-(Web-7B/Code-7B)[48]forWorkflowDistillation) Evaluation Protocol and MetricsWe report overall Pass@𝐾[12] following [111], defined as the un- weightedaverageofPass@ 𝐾acrossall 𝑀datasets(Pass@1bydefault,Pass@3whereapplicable)."},{"citing_arxiv_id":"2605.20876","ref_index":25,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Terminal-World: Scaling Terminal-Agent Environments via Agent Skills","primary_cat":"cs.CL","submitted_at":"2026-05-20T08:14:51+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Terminal-World is a skill-based synthesis pipeline that generates 5,723 training environments and produces Terminal-World-32B which outperforms baselines on Terminal-Bench 2.0 using only 1.2% of the data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.17561","ref_index":34,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports","primary_cat":"cs.SE","submitted_at":"2026-05-17T17:45:13+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"Retrieval-augmented generation reaches 0.66 weighted F1 for invalid bug report subclassification while agentic web search reaches 68.9% Judge LLM success for no-code fix generation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.16217","ref_index":53,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Argus: Evidence Assembly for Scalable Deep Research Agents","primary_cat":"cs.CL","submitted_at":"2026-05-15T17:29:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Argus coordinates a Navigator and multiple Searchers via an evidence graph for deep research, reporting average gains of 5.5 points with one Searcher and 12.7 points with eight parallel Searchers across eight benchmarks, reaching 86.2 on BrowseComp with 64 Searchers.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.01489","ref_index":20,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning","primary_cat":"cs.AI","submitted_at":"2026-05-02T15:26:45+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SciResearcher is a new agentic data-construction framework that trains an 8B model via supervised fine-tuning and reinforcement learning to reach 19.46% on HLE-Bio/Chem-Gold and 13-15% gains on related biology and literature benchmarks.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Verified critical step optimization for llm agents, 2026. URLhttps://arxiv.org/abs/2602.03412. [19] Junteng Liu, Yunji Li, Chi Zhang, Jingyang Li, Aili Chen, Ke Ji, Weiyu Cheng, Zijia Wu, Chengyu Du, Qidi Xu, Jiayuan Song, Zhengmao Zhu, Wenhu Chen, Pengyu Zhao, and Junxian He. Webexplorer: Explore and evolve for training long-horizon web agents, 2025. URLhttps://arxiv.org/abs/2509.06501. [20] Zexi Liu, Yuzhu Cai, Xinyu Zhu, Yujie Zheng, Runkun Chen, Ying Wen, Yanfeng Wang, Weinan E, and Siheng Chen. Ml-master: Towards ai-for-ai via integration of exploration and reasoning, 2025. URL https://arxiv.org/abs/2506.16499. [21] Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist: Towards fully automated open-ended scientific discovery, 2024."},{"citing_arxiv_id":"2604.17931","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent","primary_cat":"cs.AI","submitted_at":"2026-04-20T08:11:09+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Injecting 1% targeted synthetic data into GPT-2's pre-training substantially improves performance on 8 of 9 failing BLiMP grammatical paradigms, indicating data scarcity causes formal linguistic failures.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.04949","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Learning to Retrieve from Agent Trajectories","primary_cat":"cs.IR","submitted_at":"2026-03-30T17:59:02+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Retrievers trained on agent trajectories via the LRAT framework improve evidence recall, task success, and efficiency in agentic search benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2511.11793","ref_index":21,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling","primary_cat":"cs.CL","submitted_at":"2025-11-14T18:52:07+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MiroThinker shows that scaling agent-environment interactions via reinforcement learning lets a 72B open-source model reach up to 81.9% on GAIA and approach commercial performance on research benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}