Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2402.00782.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:46.566747Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:39:45.380468Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c62c326f-aff4-402f-9beb-a0c31d46f666 · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 212
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89aedb79-5dcd-4cae-bd54-dd6f3cd454ff · inbound
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 134
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b3437d-e8a2-4c38-90df-e92982a2163d · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d4739b4-8fbb-4f16-9d85-ca8afd503afd · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeafc856-3917-4b89-81e4-fc7dc8fc5775 · inbound
Enhancing RLHF with Human Gaze Modeling Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86dc17e9-3288-47f0-beb1-d6f026a9a300 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 1995
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d5c9c75-420b-4e5e-8e4a-28b9c5425713 · inbound
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5392912-1faf-4e67-a8c5-ad70b0730da2 · inbound
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adaaacc2-3461-49f7-b636-482d7be0e49b · inbound
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f64647a0-71af-4baf-9b8e-a5e4a8c7f973 · inbound
Multi-Rollout On-Policy Distillation via Peer Successes and Failures Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 445c0768-7675-4872-ad45-920470cccdbd · inbound
Multi-Rollout On-Policy Distillation via Peer Successes and Failures Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65932d75-77af-4ff1-a506-947e7a565913 · inbound
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa517d00-67a1-44a2-a698-d2b0ee743b75 · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf6a8f44-ebf0-4252-9e1f-b8ab5467c684 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 290
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc916ebb-4544-40d8-a59a-ffbe1455bcca · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 291
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497c2e34-36c3-4fa8-897d-2935a505487d · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text Dense Reward for Free in Reinforcement Learning from Human Feedback
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.