Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2401.16335.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:54.207436Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T22:10:42.038007Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 88e81838-4d92-4525-9c47-3872bc995401 · inbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4131ff1-1cf1-4919-bf85-d5d5a7d31616 · inbound
When Can Proxies Improve the Sample Complexity of Preference Learning? Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72dd7354-9ff1-4eef-b10d-c2f5168b8a3a · inbound
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0feb13fe-dc2c-486b-8f6c-179b65f3c4b8 · inbound
RewardBench 2: Advancing Reward Model Evaluation Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7448d546-462a-47b6-bb0a-3b9a27ab379e · inbound
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99066ce0-bf66-45e2-bff9-babc6bf79869 · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8027a939-c023-4432-b829-2482dd1db6c2 · inbound
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bab6fd4f-2ca2-4852-903e-283ac5acfd32 · inbound
OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec45e4a4-4f48-4097-955a-686854b0ab13 · inbound
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d288a023-8ffe-4e7b-8cdd-cc4a31b371dc · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b3d0f08e-3cf9-4443-9f84-41c7e45cb26e · inbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
Reference 153
Source-reported events for the cited work
Unavailable: canonical work link unavailable.