Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:21:28.356097Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 17 inbound Pith citation observations for arXiv:2507.22844.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T11:21:28.356097Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:45.506154Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
40 of 40 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation e7c58f48-76bb-413a-98aa-c35e51c3c875 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72841eb2-0aa8-436d-b40a-c3721397a918 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc625d8-f4ac-4e88-9ede-8385014f4c27 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f424d337-8c80-48a4-bd9c-b09ce0e5aaff · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5674a6bd-a0c9-4192-9ca1-59e68115dd09 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cf0d7b-ee2a-4185-b330-e95e284b7f69 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a5dedfc-061a-41d8-b712-4157991f9fe6 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentRefine: Enhancing Agent Generalization through Refinement Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f457d4-de31-4c72-9787-60e7588b2d40 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a35479b-97f9-44cd-9b9e-a06361e5caf5 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc1466c-d69b-4818-89f7-d6e44b0c8948 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Metacognition: A literature review
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde1e6f9-84a8-428e-b73f-c8c6931e5a7a · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Understanding R1-Zero-Like Training: A Critical Perspective
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 817d8432-6a6f-4089-8efc-f47b2bf6237f · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents What is metacognition? Phi delta kappan, 87 0 (9): 0 696--699, 2006
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 007c1f90-e20e-4186-91c2-bda3619ad0a0 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents s1: Simple test-time scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 374bf8bd-f7b8-4394-ad92-958409523872 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Training language models to follow instructions with human feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061d84b4-3439-411b-9caf-109d8861021b · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent planning with world knowledge model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18d5201c-9847-4104-af5f-a46c1601b22d · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Toolllm: Facilitating large language models to master 16000+ real-world apis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f230ed-348e-4100-9b70-4447b6362e8a · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Direct preference optimization: Your language model is secretly a reward model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d481db7-b25b-40e1-9bfe-246c62bd71f2 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 112a0911-c508-46da-9e21-daae599e1215 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Reflexion: Language agents with verbal reinforcement learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfcea7cb-7364-4fa2-9355-f9cb34de8804 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a88c7ea-1fe7-4db2-a16c-6ba094467c5d · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29c7ebc-a2f0-4396-8f4c-2eb7f8f98d5e · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6640c1a-146f-451c-a166-5b508956f7d4 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents, 2025 a
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1167c757-ae55-4ed4-8b77-4258b14823ad · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ScienceWorld: Is your Agent Smarter than a 5th Grader?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fecd4649-0c8e-472b-b9fe-1391c4ebe95b · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c40d27-ad2e-4f13-ad10-91e3b1feebb7 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53bb5053-9869-445c-bca5-48a438f958b9 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Watch every step! llm agent learning via iterative step-level process refinement
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9d66c38-59b3-44df-90a4-7ab831707a05 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Gpt4tools: Teaching large language model to use tools via self-instruction
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ce4d8a-523d-4628-9b25-cbf8384ce39d · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents React: Synergizing reasoning and acting in language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c28fc9-fde8-4682-b1d7-5de80f019aea · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47c97e62-e927-4e3c-b70f-fb2bdc07b4b2 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Steptool: A step-grained reinforcement learning framework for tool learning in llms
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51204792-88b2-4000-8f53-bc77a7647bbc · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 068f6e9c-9c79-462c-83e9-a33b13845154 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agenttuning: Enabling generalized agent abilities for llms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9af3635a-dd82-4606-a094-4d03986b4492 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3ffdca-1885-46b0-adfc-0d47dec6e9fb · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e28e93-3f2c-446c-bb61-04b4a72ae661 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents You only look at screens: Multimodal chain-of-action agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fab776c0-00c2-49ae-91bc-fa9fe8e225d6 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Archer: training language model agents via hierarchical multi-turn rl
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffe30d0c-eb40-4678-a19d-6972570fabd9 · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents @esa (Ref
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808f4ef1-244a-4aa2-bfc5-0a214225c43a · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe29c0e-0c7f-47c8-9952-218ab84d364d · outbound
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6287142b-4835-4649-bbb1-98811edc27ab · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 269
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f88f39fd-5a20-47bc-aadf-f41e12188cba · inbound
RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e53603-8078-4053-9f3f-5712016ef43b · inbound
Differentiable Evolutionary Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc3355c7-9b12-4dc7-a994-69ceeb212dce · inbound
Agentic Reasoning for Large Language Models RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 237
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d56dd12c-0b56-454d-9b81-0981a35525cc · inbound
RoboAgent: Chaining Basic Capabilities for Embodied Task Planning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6a405de-44c1-43f2-951e-dde3dce430d5 · inbound
Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b8b3264-a6b4-4a55-b1a6-37108d034a8e · inbound
DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 888a8468-6fc4-4324-bb82-018071e57209 · inbound
Dynamic Mixture of Latent Memories for Self-Evolving Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e1cd5bb-ed16-4e47-87b1-71e7a414d50f · inbound
Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 206
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66fbc2e9-14e9-4686-ae43-08548ab38b29 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 158c06d6-9cc9-4393-841d-356cfdfb1d41 · inbound
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 447b24f9-a4f4-4440-bff1-ce73c7e0bc3b · inbound
Diagnosing Task Insensitivity in Language Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 578e9c1e-430f-4a6a-84b3-78e2a06df079 · inbound
Where Do CoT Training Gains Land in LLM based Agents? RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac00b5aa-cc46-41ed-bd85-c273d00a1419 · inbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39e0090-b9af-40e9-bb51-fdc755be7ea1 · inbound
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · inbound
TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.