Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2401.06080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:00:53.827430Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 73fa1a52-2f38-4c64-a58c-de79b310f68f · inbound
ORPO: Monolithic Preference Optimization without Reference Model Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0b09bcf4-e9eb-4ddc-840d-b79250750aaf · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6fc7a908-e2fc-48ce-bc07-3bf1539d8369 · inbound
Qwen2.5 Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95f05551-4bd2-4d8b-8c57-24adc9473adc · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c93a8b98-a211-439d-86d1-bf05ae4a3959 · inbound
Seed1.5-VL Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5fea45f8-50fa-4f09-9407-96a73b02ce3a · inbound
Incentivizing High-Quality Human Annotations with Golden Questions Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1810d4c-e22a-4a9e-b1f4-565668a357f8 · inbound
Users as Annotators: LLM Preference Learning from Comparison Mode Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3c752b3-7b51-4c17-af23-0b4694293918 · inbound
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d3ebbd09-47a7-412e-97bf-eac652a28814 · inbound
Reinforcement Learning via Value Gradient Flow Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bdc2aa2-39d8-4cc2-a6a1-fa90e214838c · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a406e49-0558-4235-92b8-ab3799e9c7d8 · inbound
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9dcd4027-37e4-456c-b1e5-32c00cd4bbd3 · inbound
Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e9cfa9b-6513-47ce-a9ae-ac8768437ab2 · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90b60aec-fea5-48c8-a609-4e07d988c717 · inbound
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8fe3cacf-d709-4e12-97df-291176f3fe31 · inbound
Boosting Self-Consistency with Ranking Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b24150e-32d1-49a0-aa08-65e809dabaa6 · inbound
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bcf61e05-b703-4838-96f2-ca6ee50f3ef2 · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 24da175c-92c6-47ff-9ec0-a4ccd107e1a7 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 84b2b752-56a8-4a70-ac76-0f528db5dc7a · inbound
Understanding helpfulness and harmless tension in reward models Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33796558-f5e7-46a1-87c9-9ca4b3415e13 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 210
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90eb243c-e305-431b-99ab-b8b2084915af · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e39d2557-3329-4f19-bc1b-9b02bb77a51b · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e433b4e7-168e-4078-ad68-e57a0bc95423 · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36eec2bb-fe64-48d6-9715-b767de896496 · inbound
Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.