Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2401.12187.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.649860Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · inbound
Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 024add06-7af1-4b65-8d0d-cbf5366dfd3d · inbound
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling WARM: On the Benefits of Weight Averaged Reward Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 608ab5de-8333-4751-8d8c-66a9fe2ffdbd · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary WARM: On the Benefits of Weight Averaged Reward Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · inbound
Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c4a7cc-89a7-40d8-941d-96d86ef9ce1e · inbound
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback WARM: On the Benefits of Weight Averaged Reward Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · inbound
Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34333c6-eb98-413a-9e3b-ea7b1f8d8977 · inbound
Mitigating Multimodal Hallucination via Phase-wise Self-reward WARM: On the Benefits of Weight Averaged Reward Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4dc419e5-df9d-4ad3-ac14-335e43c02c16 · inbound
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs WARM: On the Benefits of Weight Averaged Reward Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb16ecb7-3933-47a1-9e97-7eaa7e4bb1dc · inbound
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity WARM: On the Benefits of Weight Averaged Reward Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fc38837-3c15-43a3-bcf9-70fdafaa011f · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization WARM: On the Benefits of Weight Averaged Reward Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7de95927-6c1b-4e41-9256-d45d2ccd6eb2 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f096b6-45c0-42cb-8ad9-9578d5e75310 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d164ed71-f4cc-46b1-90dc-f45e60b3ff22 · inbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL WARM: On the Benefits of Weight Averaged Reward Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.