Pith. sign in

Paper Citation Record · LEDGER

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 17 inbound Pith citation observations for arXiv:2507.22844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22844 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:21:28.356097Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:45.506154Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e7c58f48-76bb-413a-98aa-c35e51c3c875 · outbound

This paper cites write newline.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.792677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.792677Z digest=sha256:76e6d99b65165da5e0d1f6025caa4a29f40f382aa68b39d31bc4c8f9e5e9da46

Observation 72841eb2-0aa8-436d-b40a-c3721397a918 · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.872638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.872638Z digest=sha256:b811ed07f42393919ca6633717c80fe546099af1d734e532e5114192fec88677

Observation 9dc625d8-f4ac-4e88-9ede-8385014f4c27 · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.987699Z digest=sha256:988b2cff82702d4577edf48ebb61b6aae1caea3767de40216c87ea4faa652fd8

Observation f424d337-8c80-48a4-bd9c-b09ce0e5aaff · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.067188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.067188Z digest=sha256:ad4dcd13720d56b71c86a59de9b1602f7516c847bdc9ae429152e4f42ae987b5

Observation 5674a6bd-a0c9-4192-9ca1-59e68115dd09 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.125749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.125749Z digest=sha256:0368c9a867bb00c30ae2add1ad4f4d3c6dadfc7c03db95968d986bf56a19cc23

Observation 44cf0d7b-ee2a-4185-b330-e95e284b7f69 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.329821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.329821Z digest=sha256:8cd18878097924947d9839b55c39dbc5eca072122bf8813ffb4ac38db371a66a

Observation 6a5dedfc-061a-41d8-b712-4157991f9fe6 · outbound

This paper cites AgentRefine: Enhancing Agent Generalization through Refinement Tuning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentRefine: Enhancing Agent Generalization through Refinement Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.444678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.444678Z digest=sha256:b5f9b5f112b389862c1fd376525e6b918a23c6c3989975038cc842bf68c79d94

Observation 18f457d4-de31-4c72-9787-60e7588b2d40 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.567372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.567372Z digest=sha256:a607b922279fd0c2e80ddb9e2b327cd1ea3e8866896be7349af4a6bf54b96f76

Observation 5a35479b-97f9-44cd-9b9e-a06361e5caf5 · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.638167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.638167Z digest=sha256:9af056a6615badbbc2bf2b599286a86469bc4b61f77b1dd65fffec3d0bafad3f

Observation 0fc1466c-d69b-4818-89f7-d6e44b0c8948 · outbound

This paper cites Metacognition: A literature review.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Metacognition: A literature review

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.727837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.727837Z digest=sha256:20df6d6bb34a291d09788b78bc3b71bf72eba7a07d4eac72da0652514b1c3b61

Observation cde1e6f9-84a8-428e-b73f-c8c6931e5a7a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.854273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.854273Z digest=sha256:f5cca41d50504aa02960c298f0fe575418972e7bd690e89dc44c4a646c1312a7

Observation 817d8432-6a6f-4089-8efc-f47b2bf6237f · outbound

This paper cites What is metacognition? Phi delta kappan, 87 0 (9): 0 696--699, 2006.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents What is metacognition? Phi delta kappan, 87 0 (9): 0 696--699, 2006

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:30.281687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:25.985422Z digest=sha256:e08d206e08e96ea4ae9420e2c5248cb1fa58fc8c736f7df5d5ec1cf858ab976d

Observation 007c1f90-e20e-4186-91c2-bda3619ad0a0 · outbound

This paper cites s1: Simple test-time scaling.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.067539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.067539Z digest=sha256:419d3ea493c26dd5676f832e883425264ea967caf4a994781e015ce911c0a17e

Observation 374bf8bd-f7b8-4394-ad92-958409523872 · outbound

This paper cites Training language models to follow instructions with human feedback.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.137344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.137344Z digest=sha256:76f3118da5e8f06ce49fcb84b0dc5b42c7031a2354f7980e13214cfb702a7c80

Observation 061d84b4-3439-411b-9caf-109d8861021b · outbound

This paper cites Agent planning with world knowledge model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent planning with world knowledge model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:30.038442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:26.223002Z digest=sha256:f43130c804dc65a75adc8c5b5c74f1fbfa540e1abf990b14d77690955894cb2c

Observation 18d5201c-9847-4104-af5f-a46c1601b22d · outbound

This paper cites Toolllm: Facilitating large language models to master 16000+ real-world apis.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Toolllm: Facilitating large language models to master 16000+ real-world apis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.314995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.314995Z digest=sha256:5d192e2a1b41472a337f3563bc323adf310aca63d47561d325976349ab715f39

Observation 51f230ed-348e-4100-9b70-4447b6362e8a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Direct preference optimization: Your language model is secretly a reward model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.421243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.421243Z digest=sha256:11257489c2e54a4a850191d6f14f8c159452b15ca8b67fa1bc71adddb61476ff

Observation 8d481db7-b25b-40e1-9bfe-246c62bd71f2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.503078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.503078Z digest=sha256:b19476fba6cfc12b8125261ee5a51ddf254b84fe72d5cc64303732a8b9adc3f4

Observation 112a0911-c508-46da-9e21-daae599e1215 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Reflexion: Language agents with verbal reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.592831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.592831Z digest=sha256:dafad092a5bf455e1e27bd5bab5e8f7e9c70d38dd9725f7ef2b7123e176a36f5

Observation dfcea7cb-7364-4fa2-9355-f9cb34de8804 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.682417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.682417Z digest=sha256:3569d5e09e1880e4eb754bb14d2c9354b60a273fd872fd2ef614abd91b906468

Observation 2a88c7ea-1fe7-4db2-a16c-6ba094467c5d · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.783454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.783454Z digest=sha256:73ea9d0128ff74c6a6966cd0917310125336dee7eff7c4a90be2a49dd85b995a

Observation d29c7ebc-a2f0-4396-8f4c-2eb7f8f98d5e · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.873512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.873512Z digest=sha256:07cb75414ee0654e40d4062f6919cf42fb188d1ec17b5354f5f4c592f59c02a9

Observation d6640c1a-146f-451c-a166-5b508956f7d4 · outbound

This paper cites Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents, 2025 a.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents, 2025 a

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.966172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.966172Z digest=sha256:5a4680e9ea68c33c13a769f2aa39b49e68a8788c5886c263fa25f4eb7c58afda

Observation 1167c757-ae55-4ed4-8b77-4258b14823ad · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.035759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.035759Z digest=sha256:daedf4d4b7293a45e6d9c907de2792bb8b18c57f7be64c9cad2f1d82faceb5dd

Observation fecd4649-0c8e-472b-b9fe-1391c4ebe95b · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.115747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.115747Z digest=sha256:066b1944c7091cf5628e9cf6bd34e9f2d5872bb4c90b2f18063dc890ad4d8a3f

Observation 68c40d27-ad2e-4f13-ad10-91e3b1feebb7 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.211465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.211465Z digest=sha256:9c683bd700241be694724689a8ae57c2ccbf5762e4800836700f6e634f54a6b4

Observation 53bb5053-9869-445c-bca5-48a438f958b9 · outbound

This paper cites Watch every step! llm agent learning via iterative step-level process refinement.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Watch every step! llm agent learning via iterative step-level process refinement

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.861644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.280991Z digest=sha256:fc8c763da565e837c2d82ee310e7700c9a8297fb6e1cf4f97b9d6e0ed5892ee1

Observation b9d66c38-59b3-44df-90a4-7ab831707a05 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.668964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.338862Z digest=sha256:ea063084f5f036e6604f6ed0250a3991176ac5a0315cbb9bf6509308e21ae1b9

Observation e2ce4d8a-523d-4628-9b25-cbf8384ce39d · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents React: Synergizing reasoning and acting in language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.421252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.421252Z digest=sha256:4005b9e127e37834da40b12e92744c1e2722ee50a9016591f6dce3071c2d433e

Observation 73c28fc9-fde8-4682-b1d7-5de80f019aea · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.471376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.471376Z digest=sha256:d5629dd88f0e9314d88ac19f41c449e419994e2dc270996c749d6c3e8458a6ab

Observation 47c97e62-e927-4e3c-b70f-fb2bdc07b4b2 · outbound

This paper cites Steptool: A step-grained reinforcement learning framework for tool learning in llms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Steptool: A step-grained reinforcement learning framework for tool learning in llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.450056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.587989Z digest=sha256:186717041a17f3ac395655303d2650aa71abfcbaf4b8a7dbc9f5d367389bde20

Observation 51204792-88b2-4000-8f53-bc77a7647bbc · outbound

This paper cites Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.676834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.676834Z digest=sha256:e073447b39cfe7a3e007750d8618cf106897ce2d8c39d0bb363dc144ba7e529c

Observation 068f6e9c-9c79-462c-83e9-a33b13845154 · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agenttuning: Enabling generalized agent abilities for llms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.284262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.740567Z digest=sha256:4440ba93b77825eef87ab2633e1d56e2f1fad88e5312e109b98b59de393d4d83

Observation 9af3635a-dd82-4606-a094-4d03986b4492 · outbound

This paper cites Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.805649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.805649Z digest=sha256:daabea1aba99226a45b8c8e2dcfaffcab24dd36b4905c9345075bef9584d7e1e

Observation fd3ffdca-1885-46b0-adfc-0d47dec6e9fb · outbound

This paper cites Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.893716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.893716Z digest=sha256:f6288cea7555b6171480c504dc743e20ee14691081f4be498149666334fe9c02

Observation b5e28e93-3f2c-446c-bb61-04b4a72ae661 · outbound

This paper cites You only look at screens: Multimodal chain-of-action agents.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents You only look at screens: Multimodal chain-of-action agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.105836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.988018Z digest=sha256:ec3792847140844f289c40be4cd64d86414705dc1cedfb820b0d6bb5ab1b52e9

Observation fab776c0-00c2-49ae-91bc-fa9fe8e225d6 · outbound

This paper cites Archer: training language model agents via hierarchical multi-turn rl.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Archer: training language model agents via hierarchical multi-turn rl

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:28.945122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:28.082695Z digest=sha256:a969c350e89ee152b4ffbc0d8359c423a943b39098ab963ec06d3dae75fa3efb

Observation ffe30d0c-eb40-4678-a19d-6972570fabd9 · outbound

This paper cites @esa (Ref.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents @esa (Ref

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:28.176007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:28.176007Z digest=sha256:30857d888a4f27a8014653def358179361b28f73681af591130070300d3d9c1e

Observation 808f4ef1-244a-4aa2-bfc5-0a214225c43a · outbound

This paper cites an unresolved cited work.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:28.285293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:28.285293Z digest=sha256:3bdd19be24b931df14331811f064c09caf071fd27a78c48f00bd0d644864e42c

Observation afe29c0e-0c7f-47c8-9952-218ab84d364d · outbound

This paper cites an unresolved cited work.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T11:21:28.775163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T11:21:28.356097Z digest=sha256:6dc8f0af8339012105b86eeed0147b9798e83a54890783ff23b1d64c9eb7735c

Pith citing papers

Observation 6287142b-4835-4649-bbb1-98811edc27ab · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.666577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:50603d0591cf4dc05241f90652d52f6b41e7fd079507a0e70044aa8ad0f8c30a

Observation f88f39fd-5a20-47bc-aadf-f41e12188cba · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:42.644462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:42.644462Z digest=sha256:ca63c7e3dbb39cf119cff1eb2ebc7a50f399db8b93b11db903edb71f0cf3c178

Observation 78e53603-8078-4053-9f3f-5712016ef43b · inbound

Differentiable Evolutionary Reinforcement Learning cites this paper.

Differentiable Evolutionary Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.689797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:36e7c1bf2cae60213e8f1b61fbdb076c0b3fce4d67304ed6d61260bb0e8dce2e

Observation cc3355c7-9b12-4dc7-a994-69ceeb212dce · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.310487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:cd0ef2252e0350a3f72b8fae1b08dcebf0631146a784dcaac3f8624836baf35a

Observation d56dd12c-0b56-454d-9b81-0981a35525cc · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.001034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:22f4bfd1873c05bd5f122f6717ff9864f08aae161c97c3c0bf12380a6376713d

Observation f6a405de-44c1-43f2-951e-dde3dce430d5 · inbound

Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning cites this paper.

Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.129260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:14:19.580071Z digest=sha256:ceaee06a8f67133952379329de7c52b0dc63fd956dfd6f1300f48d4a82e2b775

Observation 7b8b3264-a6b4-4a55-b1a6-37108d034a8e · inbound

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents cites this paper.

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:01:11.869722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:37:30.156901Z digest=sha256:076efe6f712e3b82753df5aff3ac88ffb5ec0345d8e4fa12a63af7f47fc9124f

Observation 888a8468-6fc4-4324-bb82-018071e57209 · inbound

Dynamic Mixture of Latent Memories for Self-Evolving Agents cites this paper.

Dynamic Mixture of Latent Memories for Self-Evolving Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:41:14.489608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:38:42.113849Z digest=sha256:065823f9f2896a9424b786b2ea842a97419f880489f8d3dc165c6b0bedd869d9

Observation 2e1cd5bb-ed16-4e47-87b1-71e7a414d50f · inbound

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security cites this paper.

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.653011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:18:40.244556Z digest=sha256:aba63bc2df3fea6e4e522ff2d05b243b9a241d5b79dade6ca3ec429ac2d9df3d

Observation 66fbc2e9-14e9-4686-ae43-08548ab38b29 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.863082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:24917125060764ce4abcc4ae27add5fed9c5a6839bbb079726c391bc42c24954

Observation 158c06d6-9cc9-4393-841d-356cfdfb1d41 · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.618792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:2897a290c6d111b927128e8247c6b0319bcd905cbf010af43839df05dfc0cec9

Observation 447b24f9-a4f4-4440-bff1-ce73c7e0bc3b · inbound

Diagnosing Task Insensitivity in Language Agents cites this paper.

Diagnosing Task Insensitivity in Language Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.548061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:58:11.929896Z digest=sha256:de2a6fa611efa0f69c16bd95fd5fd45ee33f0aee76401174fb9578f481125836

Observation 578e9c1e-430f-4a6a-84b3-78e2a06df079 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:51.517553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:f7de042196d70291f8c256acb56098e38f2b23ca7aef0739fe43d074ea3decb4

Observation ac00b5aa-cc46-41ed-bd85-c273d00a1419 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:b9ffefec13f036f6f2029469ba560ee6f9dd66539f3796c81b2979209528b49e

Observation f39e0090-b9af-40e9-bb51-fdc755be7ea1 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:086868e39b19569aa2cb6c6715b6e21fdf215bc0b2e1012ebcdd8f171d4b587e

Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · inbound

TAPO: Transition-Aware Policy Optimization for LLM Agents cites this paper.

TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.696164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.696164Z digest=sha256:a263e645a96049a2bf417706bbbaf767f7340739c05eea850412eacf8936f6cd

Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.506154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.506154Z digest=sha256:d84f0606f9dcb9df1ad4577ed91b5b586bc5b14171896bb9eeaa4e4f73eb1e3d