Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2310.02743.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:57:54.254039Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 56747fdd-9082-453d-8616-4d7b19198df2 · inbound
InfAlign: Inference-aware language model alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3404ba26-329b-4d9c-837a-3c7f7aaaffc3 · inbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a093c96a-6bff-4f54-83cc-876ca28f9af7 · inbound
Debate Helps Weak-to-Strong Generalization Reward Model Ensembles Help Mitigate Overoptimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e538dc2d-97f4-4c23-8dd9-ab195e63b329 · inbound
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5258d64d-76db-48d4-accb-5a59a590a8eb · inbound
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Reward Model Ensembles Help Mitigate Overoptimization
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3fe5c1-4a8a-434b-865d-c3999a37b8c8 · inbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Reward Model Ensembles Help Mitigate Overoptimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 726b7ead-80f1-4df5-a9f2-480b78a979cd · inbound
RewardAnything: Generalizable Principle-Following Reward Models Reward Model Ensembles Help Mitigate Overoptimization
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1560d3f1-fc0a-4acf-8370-8e9b79abaaec · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Reward Model Ensembles Help Mitigate Overoptimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d4acc6-c2fe-47e3-8ed9-3c0a9e4d5f3e · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Reward Model Ensembles Help Mitigate Overoptimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81661bf-1992-4d60-91f9-a40889259d1e · inbound
Towards Reliable, Uncertainty-Aware Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32ede528-99af-4867-8c7c-71b624daa0fc · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Model Ensembles Help Mitigate Overoptimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7eced65-43f1-4476-b2e8-5780b93e53dc · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reward Model Ensembles Help Mitigate Overoptimization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7146a6c9-53c9-47d7-a33f-47bb6137bb49 · inbound
Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Reward Model Ensembles Help Mitigate Overoptimization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96982241-62cb-40e8-b1b3-331ca867a0ba · inbound
Reinforcement Learning via Value Gradient Flow Reward Model Ensembles Help Mitigate Overoptimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c002d6f-dab3-463b-8179-4bf45742812e · inbound
FUSE: Ensembling Verifiers with Zero Labeled Data Reward Model Ensembles Help Mitigate Overoptimization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 21372565-3424-40f8-a0fe-139255f3463c · inbound
Theoretical Limits of Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48120300-4e5a-45c6-b2a4-d5056048f549 · inbound
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Reward Model Ensembles Help Mitigate Overoptimization
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5ed0fe47-9729-45d5-b6c2-bebe997b37a6 · inbound
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 44d10863-751c-4a77-bdd7-f20068c52e47 · inbound
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e2d1bdb-71b1-4184-acae-d07b2895a53b · inbound
Variance-aware Reward Modeling with Anchor Guidance Reward Model Ensembles Help Mitigate Overoptimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cbef8efc-0e05-4646-a69f-dc02226ae345 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f2ebf4ad-b424-4b6e-a21d-1fed4734bd6e · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0c514f6-1221-492d-82a9-150c6bb749d0 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Reward Model Ensembles Help Mitigate Overoptimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 926e9c09-7e2b-4af4-b7dc-b6714b222c07 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reward Model Ensembles Help Mitigate Overoptimization
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef72fedb-2c99-414b-baf1-f5c5b96e8386 · inbound
Uncertainty-Aware Reward Modeling for Stable RLHF Reward Model Ensembles Help Mitigate Overoptimization
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6cad619-e42a-43b6-bd45-7ad7a0fa5bd8 · inbound
Against Proxy Optimization Reward Model Ensembles Help Mitigate Overoptimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9ac9227-cc96-4ce7-8e35-5affe4476085 · inbound
Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search Reward Model Ensembles Help Mitigate Overoptimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 325d1168-d703-44be-b697-bc9e1cf32054 · inbound
STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning Reward Model Ensembles Help Mitigate Overoptimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50ba71a1-39ee-4081-94da-de8cc3d99fab · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Reward Model Ensembles Help Mitigate Overoptimization
Reference 283
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d41dcab2-c667-4af5-a4dc-773fd6247303 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Reward Model Ensembles Help Mitigate Overoptimization
Reference 284
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e426a20c-f573-4c4b-a2db-4c6fccabcff7 · inbound
More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Reward Model Ensembles Help Mitigate Overoptimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.