Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2312.09244.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:53.635786Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation f9169d0a-e4b5-4e96-b42f-62d01dd8571c · inbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe83929-5b8f-41df-8304-74f43f8010da · inbound
T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39000e2c-02cc-43e3-bb45-c39de069dde3 · inbound
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ccfc5b-4c63-4600-9e91-6f20604f8e6a · inbound
Learning a Pessimistic Reward Model in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caeb390-76d5-4292-abbc-7d1102c841bb · inbound
RewardAnything: Generalizable Principle-Following Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8325f5-f3fe-49db-925f-63231a3714a4 · inbound
Activation Reward Models for Few-Shot Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3366544c-778c-497b-a6c4-f82a7b6c2df9 · inbound
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb147b1-feb1-4f5e-ab42-0d89a6116e3c · inbound
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2eceda81-b0f9-4057-8cee-6fd326be5e8f · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1936df79-e57c-42bc-a82a-46f309452a11 · inbound
Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d88245e-cb38-489a-962f-89e09abc6976 · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2b52343-ad36-4d03-85ab-8d46ab45da4f · inbound
Towards Reliable, Uncertainty-Aware Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d33daff4-48b0-4d19-ad27-e641f2b9cd80 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2209bca8-a672-4f0c-963f-839ffae29ba4 · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53c188b7-ccb1-488a-a157-1e3c3af0e5bf · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1462243-f2f0-4201-a5d6-c14fc88c3639 · inbound
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 647e0a61-a9b6-445e-b119-5b1e41c6d209 · inbound
FUSE: Ensembling Verifiers with Zero Labeled Data Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2445045-e6d9-4182-b38a-59f9449287e6 · inbound
How Far Are Video Models from True Multimodal Reasoning? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5dba58d5-f855-4cff-9966-6e019721a62f · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91e20ee9-df19-4181-bc55-03a36ca23865 · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 595dd3df-fe53-4938-8421-b6fd18a0d474 · inbound
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b20806bf-3933-4d79-9015-e98990d3525c · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e39bdd7b-6892-43cb-a10e-85be05c618c2 · inbound
Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8342c137-ccd7-40b3-9b84-bd3eec498c32 · inbound
The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb4d80df-6aaf-4f90-8445-5265e2833eb8 · inbound
The Human-AI Delegation-Verification Dilemma: Individual Strategies, Collective Equilibria and Sociotechnical Lock-in Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c03e3297-ef02-4b98-b201-5ac4c2e9b141 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 683dab4d-0563-4100-970f-496a022b5b2b · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76842f9b-559e-43ae-8c8d-84afba033d7b · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ca641ca-ca54-47dc-9f8b-594dc499edfa · inbound
HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16667bdb-4ecb-4f3a-816d-7d452f5c4284 · inbound
A Unifying Lens on Reward Uncertainty in RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45b67a8f-0ef7-4aeb-9754-43663e65f0db · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce4e1c54-6b51-44d1-9195-3b9771a503d0 · inbound
Uncertainty-Aware Reward Modeling for Stable RLHF Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b1a7393-77d3-4d5c-8dd4-545a86f3333d · inbound
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c52b93cc-4f9a-46e5-8677-0fc57a8e894f · inbound
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1589a814-1230-4d0d-a3e6-e6e0514900d8 · inbound
What do Reward Models Memorize? Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.