Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:22:02.778783Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 27 inbound Pith citation observations for arXiv:2501.11425.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T18:22:02.778783Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:31:12.889246Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T13:56:19.181721Z
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fc30c406-cd27-445d-a861-b0f1b83d3663 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The claude 3 model family: Opus, sonnet, haiku
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5b2e6b-ec7d-4f59-9342-ee5d6bcc9aa9 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4 (1):1–43, 2012
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8056b4fb-a666-4045-a1bb-097288ad86ce · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2bc9fb5-a5b4-4ef4-8c9d-68613b0e36d1 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Teaching large language models to self-debug
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9787c36e-e909-4162-b804-af0ca607f57a · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mindsearch: Mimicking human minds elicits deep ai searcher.arXiv preprint arXiv:2407.20183, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6210d6f9-f99d-4d3f-81e6-2725fdfa4412 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent-FLAN: Designing data and methods of effective agent tuning for large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7fcaad-0590-42bf-ad33-d4d5966e6ef4 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mind2web: Towards a generalist agent for the web
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 26f30bf4-4a45-49d9-a9fa-2141e6728f8e · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c0765e97-8f2c-47cc-b15e-79b0dccb2673 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23979a3-e773-43ec-88df-a1913136c8e2 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CRITIC: Large language models can self-correct with tool-interactive critiquing
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation da80a68a-9e6f-4df9-b7a2-8b56d5762239 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Reasoning with language model is planning with world model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ce9379a9-42a0-4f69-a7e5-203032654627 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00a9c5c-f345-46d2-90a7-f5e9191715aa · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large Language Models Cannot Self-Correct Reasoning Yet
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc36c2b5-9564-4f71-baa2-68aab2079569 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training When can LLMs actually correct their own mistakes? a critical survey of self-correction of LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0470dd-b146-4bbc-87a6-5b3811586b23 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Language models can solve computer tasks.Advances in Neural Information Processing Systems, 36, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d1bbc986-49d0-4ca2-896b-78f13456cb4c · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Bandit based monte-carlo planning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a94272-d4cb-4793-83b2-cab28001cc60 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Tree search for language model agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ea87d23-bed6-47eb-bf51-54565b8895b2 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Training Language Models to Self-Correct via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe50344-23bc-4dd6-b5b5-afd588893e4e · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e4b7a4b8-e042-4425-be3a-058db3c31911 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3775be2-59b5-40c0-89c5-ac32b53d0c83 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65759579-653f-4f03-b936-60b3d1dc65e8 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agentbench: Evaluating LLMs as agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 50791d4d-44c9-4e19-bc7f-943e920f4d32 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Self-refine: Iterative refinement with self-feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81f6ec6e-87d6-4675-9b33-216e3f464499 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba0a2d0-9a25-41a7-b62b-6bab1660447b · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Skill set optimization: Reinforcing language model behavior via transferable skills.arXiv,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1aab42fc-a660-47e3-8191-1d5f5235d877 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Is self-repair a silver bullet for code generation? InThe TwelfthInternational Conference on Learning Representations, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4ff89f03-d10d-41b3-a00f-8d650f6bfa51 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Chatgpt, 2022
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a3d8e3b5-4dfe-4198-a0c2-ab44ed9e8891 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training GPT-4 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 017dd6c6-f96a-49f6-b60e-fbf466411766 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Automatically correcting large language models: Surveying the landscape of diverse automated correction 16 strategies
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9b16a689-438d-40d6-bb91-282ac9963ae4 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large Language Models Can Self-Improve At Web Agent Tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf92a693-a6f9-40d5-bcd9-91eab11af435 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training ADaPT: As-needed decomposition and planning with language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ade1b4b-ae86-431f-a876-ce7674bababd · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfb9819-8c81-4aef-87d9-67fdbb1116d2 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Agent planning with world knowledge model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5bf36f80-3d4f-4ec9-b0a5-af1cdd27e2a3 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Direct preference optimization: Your language model is secretly a reward model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4aabfb35-7752-428e-8b4a-bf5326e186d0 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Tarr, William W
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5809b1a9-2387-48c7-95d6-9c2197943fd2 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Training Language Models with Language Feedback at Scale
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a7a265f-2706-4240-b3d2-97d936605bf7 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Direct multi-turn preference optimization for language agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0969a0b0-779e-48d3-8434-8c62e6ceed97 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 074fb72d-495b-4f9c-9f98-5387ed6e8771 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentBank: Towards generalized LLM agents via fine-tuning on 50000+ interaction trajectories
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d42ed4-d03e-4cec-97a7-a1ae631c99e8 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Trial and error: Exploration- based trajectory optimization of LLM agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14980220-a53a-42f9-8068-999085826d5a · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c7b7ecc-f1b2-475e-9124-50a1bffb1152 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training LLMs cannot find reasoning errors, but can correct them given the error location
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f57b7be-685c-4589-a5f1-d83132e06c45 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training E2CL: Exploration-based error correction learning for embodied agents
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a95e2d-0e36-4fe4-88d8-5ed75631bc11 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1191ddaf-e89b-4d2a-890d-0857e25a1068 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a7e4b1-7a16-494b-9010-50aa380defec · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Generating sequences by learning to self-correct
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 10a268ff-9cd9-4e05-b738-7e099d664208 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf3f601e-e429-4d2c-9ab2-ec4c26db51de · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c206319f-92cb-444b-9adf-3a15c3694d30 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd919ec-6f6e-450d-a938-e4d10ae54cd9 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Revealing the Barriers of Language Agents in Planning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c075135a-30b0-4840-a78f-ac5030ce8299 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Watch every step! LLM agent learning via iterative step-level process refinement
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e74debc2-423f-4817-9ebd-665e86ea25dc · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Webshop: Towards scalable real-world web interaction with grounded language agents
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc5e646f-feee-4270-b631-4d500d6a64c6 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training doi: 10.18653/v1/2024.emnlp-main.93
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c63a08-8297-46d0-a87d-5d035a6249d8 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training React: Synergizing reasoning and acting in language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cf35c683-d0ab-423a-882c-f5201de890fb · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Griffiths, Yuan Cao, and Karthik R Narasimhan
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 92a6e8a6-28df-4483-836a-1c185e033216 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training AgentTuning: Enabling generalized agent abilities for LLMs
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24abe105-cf0b-434e-9722-75430a43dce3 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 055d0052-2503-4156-b4c9-94ca80654ecb · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c684ccf-89ed-4073-8934-75df5e4e3385 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Mr-ben: A meta-reasoning benchmark for evaluating system-2 thinking in llms
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a562d886-1e95-4db8-aedd-b9fac972a046 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training TimeArena: Shaping efficient multitasking language agents in a time-aware simulation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1a7ad6e-ff5a-4b07-a3ee-63a1deb9b859 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Understanding the Dark Side of LLMs' Intrinsic Self-Correction
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb57cd0c-894e-404e-ad6c-ce3f6c9de42d · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Large language models as commonsense knowledge for large-scale task planning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c61f8360-c129-4225-bc9b-1b7963c77796 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training doi: 10.18653/v1/2024.acl-long.215
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 862c5247-0abc-48b1-a4bf-12cf69eea0eb · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Expel: Llm agents are experiential learners
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad1a84b3-d726-436e-a107-0b84beaf7b08 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c65df3-ac3a-477f-bea1-81b2af148280 · outbound
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Skill Set Optimization: Reinforcing Language Model Behavior via Transferable Skills
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6783f560-53d4-472e-a0b7-7c0b9d4f7635 · inbound
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e371134-90ec-4c96-b54d-11d522178d0e · inbound
Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 262
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ef6a90-fbf8-41b6-b362-2ba868b5ac26 · inbound
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54b30e9-74f0-491c-911b-2e4fe225f983 · inbound
Position: Agent Should Invoke External Tools ONLY When Epistemically Necessary Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 43fc11a3-37d9-4aed-a4af-3d44e94031d9 · inbound
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3e3565-2606-4351-a445-408950042b7d · inbound
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 60ac31f0-aca5-4678-9ad2-e093e0cffafa · inbound
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e784283-0055-4fc5-9ae7-06df5910fdd3 · inbound
SAND: Boosting LLM Agents with Self-Taught Action Deliberation Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11389ede-4427-4e8b-9513-85c8adb70359 · inbound
Agent Safety Alignment via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51204792-88b2-4000-8f53-bc77a7647bbc · inbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6063efe-068b-4e87-b5ff-a213379dd495 · inbound
ReQuestNet: A Foundational Learning model for Channel Estimation Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954732c8-7d21-451e-80ed-82a9d6e7e726 · inbound
Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation de9b7b6a-e106-4af2-bfbf-effc413786b6 · inbound
TEC: A Collection of Human Trial-and-error Trajectories for Problem Solving Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bcc60619-bd81-4b0f-b25f-4c9e98bba6b2 · inbound
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 16135e9c-96b2-45d1-a29a-6a40c1678038 · inbound
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 564c48f5-03ab-4bf5-bec0-99c7ad73ab0c · inbound
SEAL: Synergistic Co-Evolution of Agents and Learning Environments Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ab2e037c-6af1-4670-ad6a-2f084ca6cb25 · inbound
COMAP: Co-Evolving World Models and Agent Policies for LLM Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1636119b-fa91-4868-b79f-a072ea86b6a0 · inbound
Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a797b419-0d5b-43d1-8e7a-6a9debd70232 · inbound
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382ab218-744c-4c8e-b635-2ac60247a19e · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 92210d39-2e0c-402e-acb5-0ec6fe9095ca · inbound
BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96571b2-8592-4b12-97e6-b2a69fdcc141 · inbound
Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af4d129-b140-4e06-ad80-fd63bfe22bdc · inbound
Agents in the Wild: Where Research Meets Deployment Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b173eeee-323c-4da1-baa0-0e874dc85a6f · inbound
Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad88add-183b-4f1f-bc9b-07f6cf689383 · inbound
From Failures to Supervision: DynamicEnvPlan for Robust Long-Horizon Embodied Planning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17822924-6ec9-4ab3-b904-8366e7604baf · inbound
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20bb84e0-bdfb-42d2-a03e-4b3ccbd56b01 · inbound
MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.