Pith. sign in

Paper Citation Record · LEDGER

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 17 inbound Pith citation observations for arXiv:2507.22844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22844 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:21:28.356097Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:45.506154Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation e7c58f48-76bb-413a-98aa-c35e51c3c875 · outbound

This paper cites write newline.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.792677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.792677Z digest=sha256:b17ad45b0cfa3832bd3f7f45a0554af50ef07733eedcf196e5ccadff1911e913

Observation 72841eb2-0aa8-436d-b40a-c3721397a918 · outbound

This paper cites Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.872638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.872638Z digest=sha256:136544f3e3264e0d7f256d0b16eecd543388fbef5b3ce953461993ade271c6cd

Observation 9dc625d8-f4ac-4e88-9ede-8385014f4c27 · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:24.987699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:24.987699Z digest=sha256:42041af39409621a441d421280c3e9a3c994c09216efbe5a6fced9826564ac94

Observation f424d337-8c80-48a4-bd9c-b09ce0e5aaff · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.067188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.067188Z digest=sha256:8f6f29f0b747fa4bc01e723af847deb04f23fb5db87fd29c7daabf118be3da2a

Observation 5674a6bd-a0c9-4192-9ca1-59e68115dd09 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.125749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.125749Z digest=sha256:54f1521d04de8452a3e53138e052538389173c3d90603cc6dd599b3d049c2040

Observation 44cf0d7b-ee2a-4185-b330-e95e284b7f69 · outbound

This paper cites Group-in-Group Policy Optimization for LLM Agent Training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Group-in-Group Policy Optimization for LLM Agent Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.329821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.329821Z digest=sha256:c5a4341bbe9e738f241a29cad75b145a72e203344fd9148b1034ef162eb8c877

Observation 6a5dedfc-061a-41d8-b712-4157991f9fe6 · outbound

This paper cites AgentRefine: Enhancing Agent Generalization through Refinement Tuning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentRefine: Enhancing Agent Generalization through Refinement Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.444678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.444678Z digest=sha256:a9182eda3fee69cf2a58fc6dafe33b3e2b80a1d78cf99a85378f6a836cd2e6c2

Observation 18f457d4-de31-4c72-9787-60e7588b2d40 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.567372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.567372Z digest=sha256:3c3d4d4e232825cb0e67491cf3106d3dbfdd96dfd73120b46466a7b1ef67cf52

Observation 5a35479b-97f9-44cd-9b9e-a06361e5caf5 · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.638167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.638167Z digest=sha256:f2d2669f77d5b4d2c562f46975d6d40e5656a00a331fc7340ef5d3b5d9f3af80

Observation 0fc1466c-d69b-4818-89f7-d6e44b0c8948 · outbound

This paper cites Metacognition: A literature review.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Metacognition: A literature review

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.727837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.727837Z digest=sha256:84302940c9052be7ad891787303b925137bf57554f71fe5db6580ca1b8c28a06

Observation cde1e6f9-84a8-428e-b73f-c8c6931e5a7a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:25.854273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:25.854273Z digest=sha256:42e0585a7edde2c26f57f3817a3c9f76030b61e45c3031ad7fe75708a7f557d1

Observation 817d8432-6a6f-4089-8efc-f47b2bf6237f · outbound

This paper cites What is metacognition? Phi delta kappan, 87 0 (9): 0 696--699, 2006.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents What is metacognition? Phi delta kappan, 87 0 (9): 0 696--699, 2006

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:30.281687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:25.985422Z digest=sha256:80bfa1fdb9083152904e9d78164fc038ae273aa9694ff1846789641cd38bdf0f

Observation 007c1f90-e20e-4186-91c2-bda3619ad0a0 · outbound

This paper cites s1: Simple test-time scaling.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.067539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.067539Z digest=sha256:72e958476732789067d13218a747eccbaf7e9e400e3790607402c79cb2348ac6

Observation 374bf8bd-f7b8-4394-ad92-958409523872 · outbound

This paper cites Training language models to follow instructions with human feedback.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Training language models to follow instructions with human feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.137344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.137344Z digest=sha256:8b6a9c6a769a3608a2cc7674c3be134ab751edcf1337ec41005d8dd2506ac2f5

Observation 061d84b4-3439-411b-9caf-109d8861021b · outbound

This paper cites Agent planning with world knowledge model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent planning with world knowledge model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:30.038442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:26.223002Z digest=sha256:3aea4d2274b117a984bbfc7d9ce3b0fe1f563d3d7e56e0e0bad64680365ea812

Observation 18d5201c-9847-4104-af5f-a46c1601b22d · outbound

This paper cites Toolllm: Facilitating large language models to master 16000+ real-world apis.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Toolllm: Facilitating large language models to master 16000+ real-world apis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.314995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.314995Z digest=sha256:400851d63c92320013a67c9fe420d8f143ea14a479bef10b30a6c84f70ba9c01

Observation 51f230ed-348e-4100-9b70-4447b6362e8a · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Direct preference optimization: Your language model is secretly a reward model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.421243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.421243Z digest=sha256:cc97e34bec80ea96014c54f81f8371571664ffe5d9251ca43cf641d1be0bbb4c

Observation 8d481db7-b25b-40e1-9bfe-246c62bd71f2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.503078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.503078Z digest=sha256:6f1b36ac565e0498ce7d4022c3c99e4675277ac1bcd9162357393c606133ae47

Observation 112a0911-c508-46da-9e21-daae599e1215 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Reflexion: Language agents with verbal reinforcement learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.592831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.592831Z digest=sha256:2f8e4ca0379bcbf2e1547c7cfff64550ac077aacb29ad29075219c8fcd60e115

Observation dfcea7cb-7364-4fa2-9355-f9cb34de8804 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.682417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.682417Z digest=sha256:e5c798e5fdf7d818ba0e44711b6ac7ca73b546342d77c938a1913e086de3ede7

Observation 2a88c7ea-1fe7-4db2-a16c-6ba094467c5d · outbound

This paper cites Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.783454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.783454Z digest=sha256:eb8482346e2a61f35064481b775c75d0e4f8ca7028cf31389199a53bbfd7e71d

Observation d29c7ebc-a2f0-4396-8f4c-2eb7f8f98d5e · outbound

This paper cites ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.873512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.873512Z digest=sha256:1de19c3a25f21466c5b2e5ba9f50f7b1f5d721102735346c6a7a2bf9d48cb07f

Observation d6640c1a-146f-451c-a166-5b508956f7d4 · outbound

This paper cites Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents, 2025 a.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Rlver: Reinforcement learning with verifiable emotion rewards for empathetic agents, 2025 a

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:26.966172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:26.966172Z digest=sha256:79ab81e71621f01828f2536c286ee7b3ec7df5da9c4e1a7b8ccf1f0001f9ba13

Observation 1167c757-ae55-4ed4-8b77-4258b14823ad · outbound

This paper cites ScienceWorld: Is your Agent Smarter than a 5th Grader?.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents ScienceWorld: Is your Agent Smarter than a 5th Grader?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.035759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.035759Z digest=sha256:650c8b1801c3b3e9381b0cd610735c9f37199a385e221e289d2624de144d0985

Observation fecd4649-0c8e-472b-b9fe-1391c4ebe95b · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.115747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.115747Z digest=sha256:e22d4bc741367d0346d4463085cc9cdf7fea618c3321ecfcd655fc9a8cfef9f5

Observation 68c40d27-ad2e-4f13-ad10-91e3b1feebb7 · outbound

This paper cites AgentGym: Evolving Large Language Model-based Agents across Diverse Environments.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents AgentGym: Evolving Large Language Model-based Agents across Diverse Environments

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.211465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.211465Z digest=sha256:ebedc3a5346cff4da5d0772f0d639b63fd5222ede9b893d58b75fcb04daf1cba

Observation 53bb5053-9869-445c-bca5-48a438f958b9 · outbound

This paper cites Watch every step! llm agent learning via iterative step-level process refinement.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Watch every step! llm agent learning via iterative step-level process refinement

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.861644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.280991Z digest=sha256:bdd17a0643d863c9d9f2152db2e292b8edd6e02c2aafa65e49e9dc6ab7183c96

Observation b9d66c38-59b3-44df-90a4-7ab831707a05 · outbound

This paper cites Gpt4tools: Teaching large language model to use tools via self-instruction.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Gpt4tools: Teaching large language model to use tools via self-instruction

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.668964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.338862Z digest=sha256:0c831200c2005d77d49d70dbde05dd1e848d6e44a275ace3327270148eb006ab

Observation e2ce4d8a-523d-4628-9b25-cbf8384ce39d · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents React: Synergizing reasoning and acting in language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.421252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.421252Z digest=sha256:860e7725e25d1d2181eb9d6ecb4a4cecffb3809f22c5ecbbf730594d4e4fc3e3

Observation 73c28fc9-fde8-4682-b1d7-5de80f019aea · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.471376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.471376Z digest=sha256:f98fda7d960e8b106bfe9e797907d38aa1f3c4570fc240888fc5ebaecfe3d461

Observation 47c97e62-e927-4e3c-b70f-fb2bdc07b4b2 · outbound

This paper cites Steptool: A step-grained reinforcement learning framework for tool learning in llms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Steptool: A step-grained reinforcement learning framework for tool learning in llms

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.450056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.587989Z digest=sha256:ed590a1ab49f64cf77439ec792f262f565e80d1270fcd4a2de92dfa5038e2729

Observation 51204792-88b2-4000-8f53-bc77a7647bbc · outbound

This paper cites Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.676834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.676834Z digest=sha256:c481dff3b9066e9802015a1878bcb0f79552cae548e0dd57ad06118de3948c00

Observation 068f6e9c-9c79-462c-83e9-a33b13845154 · outbound

This paper cites Agenttuning: Enabling generalized agent abilities for llms.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Agenttuning: Enabling generalized agent abilities for llms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.284262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.740567Z digest=sha256:1d4b6053b610e7adeea1040e4a2f26c4ea91d7723665a3b841e7fe46e13587f0

Observation 9af3635a-dd82-4606-a094-4d03986b4492 · outbound

This paper cites Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.805649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.805649Z digest=sha256:f536c37b2d794d2fd96a0439fe5b114de66ce38f185c3de4ed61fb7d844583b5

Observation fd3ffdca-1885-46b0-adfc-0d47dec6e9fb · outbound

This paper cites Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Codeagent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:27.893716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:27.893716Z digest=sha256:104d8a3b6c4788042cd3e5bb4fe5c0d044c7dc7c705b0f2fa10069525f9b71b0

Observation b5e28e93-3f2c-446c-bb61-04b4a72ae661 · outbound

This paper cites You only look at screens: Multimodal chain-of-action agents.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents You only look at screens: Multimodal chain-of-action agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:29.105836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:27.988018Z digest=sha256:fed89e7390f54ddd964d754097f4867779c40c4205a7a37f9e91e7c7ecb589b7

Observation fab776c0-00c2-49ae-91bc-fa9fe8e225d6 · outbound

This paper cites Archer: training language model agents via hierarchical multi-turn rl.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Archer: training language model agents via hierarchical multi-turn rl

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:21:28.945122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:28.082695Z digest=sha256:8a19b47335d8238f1189170586f7fed1a7625f141938f9e4f1acedd832b402fe

Observation ffe30d0c-eb40-4678-a19d-6972570fabd9 · outbound

This paper cites @esa (Ref.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents @esa (Ref

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:28.176007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:28.176007Z digest=sha256:428a18b65ebe15786c1899f136f0e6a20bd0d1cfe1dc88a10419bdb5ea269ea3

Observation 808f4ef1-244a-4aa2-bfc5-0a214225c43a · outbound

This paper cites an unresolved cited work.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T11:21:28.285293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:21:28.285293Z digest=sha256:6e16777ae0b7aea205d3e6c560d0a8bc446df955eaa0d61cce8c65c254e1f18c

Observation afe29c0e-0c7f-47c8-9952-218ab84d364d · outbound

This paper cites an unresolved cited work.

RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T11:21:28.775163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T11:21:28.356097Z digest=sha256:54c6363a9b864760341b6abd58c85dbba433200331d3a6374bee37ae8ce5a5bd

Pith citing papers

Observation 6287142b-4835-4649-bbb1-98811edc27ab · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 269

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.666577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:6a18692f74522de43c9dca222e488894febc463751f775773bdf0cc843b02636

Observation f88f39fd-5a20-47bc-aadf-f41e12188cba · inbound

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning cites this paper.

RoboGPT-R1: Enhancing Robot Task Planning with Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T09:33:42.644462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:33:42.644462Z digest=sha256:dbcf275f71fe7a3b9f39b140a3fb74da00b6a31bd4c08824cafee428466efb8d

Observation 78e53603-8078-4053-9f3f-5712016ef43b · inbound

Differentiable Evolutionary Reinforcement Learning cites this paper.

Differentiable Evolutionary Reinforcement Learning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:38:37.689797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:34:47.984401Z digest=sha256:d6aa45dc4ca971f1c02737bd7621ecd081d9384a5d31d07f548072416d82ba05

Observation cc3355c7-9b12-4dc7-a994-69ceeb212dce · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 237

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.310487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:4179002e32953e68237abadde1ca33bfa69b2a9fbab81a3248f4a945a6feeefb

Observation d56dd12c-0b56-454d-9b81-0981a35525cc · inbound

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning cites this paper.

RoboAgent: Chaining Basic Capabilities for Embodied Task Planning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:15:57.001034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:15:08.727921Z digest=sha256:9295b11d90c977cbd527dc24a798d76c02e558113ed4944f8558278677543a33

Observation f6a405de-44c1-43f2-951e-dde3dce430d5 · inbound

Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning cites this paper.

Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:21:27.129260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:14:19.580071Z digest=sha256:e213ef5178c42bb2661ab60f0cf2afebdc03e023b1e37509d78da290cc413d64

Observation 7b8b3264-a6b4-4a55-b1a6-37108d034a8e · inbound

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents cites this paper.

DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:01:11.869722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:37:30.156901Z digest=sha256:d72a7adbfdfc4782f425e5170565281187d76b359e79d95d24785425a3ab7a13

Observation 888a8468-6fc4-4324-bb82-018071e57209 · inbound

Dynamic Mixture of Latent Memories for Self-Evolving Agents cites this paper.

Dynamic Mixture of Latent Memories for Self-Evolving Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:41:14.489608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T07:38:42.113849Z digest=sha256:ccb28a5c4a1bf56e2589c6f9982f0e9b27e5775a8e60f5ec9005b9017146c4e0

Observation 2e1cd5bb-ed16-4e47-87b1-71e7a414d50f · inbound

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security cites this paper.

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 206

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.653011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:18:40.244556Z digest=sha256:cba827c7c4a2a2c09d44d16113b23b47f3182fb98200e5a2b4ea2d19a06f3fad

Observation 66fbc2e9-14e9-4686-ae43-08548ab38b29 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.863082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:c4fe43b324c806f33d46fc5259d6fc18edb5c1ec0719f1639ddb9cfd05665bbe

Observation 158c06d6-9cc9-4393-841d-356cfdfb1d41 · inbound

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents cites this paper.

Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.618792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:16:39.347676Z digest=sha256:6829f9170296720391d6d6e3c16a43e7694cd5fabd4a1504090ffce23b5746b8

Observation 447b24f9-a4f4-4440-bff1-ce73c7e0bc3b · inbound

Diagnosing Task Insensitivity in Language Agents cites this paper.

Diagnosing Task Insensitivity in Language Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:39:51.548061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:58:11.929896Z digest=sha256:4eea88ac75771a1c410e54d936a4b68ff5d2646610b1ec5307f9bd6108ab6b00

Observation 578e9c1e-430f-4a6a-84b3-78e2a06df079 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:51.517553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:09ed5032a8dbf92122c2d6002e3380813d1a6d95414ffec487ac109b4dc78038

Observation ac00b5aa-cc46-41ed-bd85-c273d00a1419 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:945695f18da686efad0fdb93c685930933947d4391abdc32612d28d4737f4add

Observation f39e0090-b9af-40e9-bb51-fdc755be7ea1 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:df1311a13baf7225b69390e2722797b8b8b8a6fe242b3b8f19b5052c7ac59c28

Observation 6198f19b-ce98-49eb-b616-8555e4fa7ad7 · inbound

TAPO: Transition-Aware Policy Optimization for LLM Agents cites this paper.

TAPO: Transition-Aware Policy Optimization for LLM Agents RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T21:44:39.696164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T21:44:39.696164Z digest=sha256:05979a6bcef8d3255143ce6421f4b8c8f782527ebbba79a58de429b81d15d5c9

Observation 0506bd30-63bc-448a-84f7-bf8d05965137 · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.506154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.506154Z digest=sha256:455f8f410b0a0e2fdb8884ed630deb14260b0f5e33e1cb527b563bb175deedcf