Pith. sign in

Paper Citation Record · LEDGER

Secrets of RLHF in Large Language Models Part II: Reward Modeling

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2401.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06080 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:00:53.827430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73fa1a52-2f38-4c64-a58c-de79b310f68f · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.707958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:892d52e39dcccc032b44f51dd00716218082566f9e0c90ce640d4a04d320abd2

Observation 0b09bcf4-e9eb-4ddc-840d-b79250750aaf · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.483570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:1f9682c3964211bf7ab213cdd4ea4d057e08eb93db7868f97fa3ea87e8ebe3ed

Observation 6fc7a908-e2fc-48ce-bc07-3bf1539d8369 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.899604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:a4a060095a8078537fae5864c76cf71c7cb5dba55f72d1b248e7b569b6e84b8d

Observation 95f05551-4bd2-4d8b-8c57-24adc9473adc · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.575352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:9314fe620d1cbb84ba5d1fc12a7fd0ae1d0f6162a1801302886a838334f7d255

Observation c93a8b98-a211-439d-86d1-bf05ae4a3959 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.576663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ac6b36dda96badb2414f919a96321e9f573d45990871ae7d44f649af3939c407

Observation 5fea45f8-50fa-4f09-9407-96a73b02ce3a · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.352622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:8088b21adea25a2cd29ba7f9912688eb7aa3e7cbea9ece52126961922419c07b

Observation b1810d4c-e22a-4a9e-b1f4-565668a357f8 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:07.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:54d643d7593dbcb87428afceea663bcd8bd7c0faa0b77d2c982125f32b12082a

Observation b3c752b3-7b51-4c17-af23-0b4694293918 · inbound

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration cites this paper.

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.419338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:07:20.180349Z digest=sha256:c6f3a4937f88eb1e4cdf8e062786cdcfd88d4b822db0f83a6fab21ee057a721b

Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.159008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:90474fcb003c5f3f70df45359949510c84edd3f693a7f4b8e50fe3d2eca8529c

Observation d3ebbd09-47a7-412e-97bf-eac652a28814 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.604855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:932abbe6c405f621c05ec697ade36a3385520fe27dec218487e9db5b988f8d14

Observation 8bdc2aa2-39d8-4cc2-a6a1-fa90e214838c · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.869033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:29ee118af337f06e41935d22370b6f9dba5f86d7327a2f6906f42efc7cd4a7c4

Observation 3a406e49-0558-4235-92b8-ab3799e9c7d8 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.718303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:410643eab4cea5b53e612d957d6918d1380e85434edfa2ef30c1b1838419bc29

Observation 9dcd4027-37e4-456c-b1e5-32c00cd4bbd3 · inbound

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity cites this paper.

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:16:18.568531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:16:09.213612Z digest=sha256:4b4c884fff4d03d54bc95a95f495f9bee5c5155919d9e0a6588e4ae416a1d614

Observation 9e9cfa9b-6513-47ce-a9ae-ac8768437ab2 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.945673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:dd352731a0658ab46e2ecf22f8d4eacea7e6e88ca432d8d5b4c928030cf4c0bb

Observation 90b60aec-fea5-48c8-a609-4e07d988c717 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.017450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:3123f07885309f2877daa323ca75311c57d3af7fef871bb6715180e4f8b0015f

Observation 8fe3cacf-d709-4e12-97df-291176f3fe31 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.296018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:a0b3745e34a43c2adc7c264b0974f86be807a2ac6613fb808498adbb5a47d5a6

Observation 7b24150e-32d1-49a0-aa08-65e809dabaa6 · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.957727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:dc765c3151cd4b6a8e5d95768e0be2736298965b3c7be97d6be3c07e484e9144

Observation bcf61e05-b703-4838-96f2-ca6ee50f3ef2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.515245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:5043a140ec118a12261f457dde44c7291a00d0a518cbf06b963981a4653c8cdf

Observation 24da175c-92c6-47ff-9ec0-a4ccd107e1a7 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.282385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:bac80d0f1226ed3c89f384f18a08e2012729e1cd42f432a6d232efbb7af4c6ba

Observation 84b2b752-56a8-4a70-ac76-0f528db5dc7a · inbound

Understanding helpfulness and harmless tension in reward models cites this paper.

Understanding helpfulness and harmless tension in reward models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:22.991762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T07:06:23.680572Z digest=sha256:d4a3a14b23de1fb0efda91351be75434880a16320108ba050a2a746b0eb9c924

Observation 33796558-f5e7-46a1-87c9-9ca4b3415e13 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:d9a3aae929a6d57723dc5db131535b32c10fa5f87f0c40fbdbf0b37d27280c6d

Observation 90eb243c-e305-431b-99ab-b8b2084915af · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.135832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:c4d61dd6f781935fabc0346e0b58163b8f6f92fb7e4db9fc2b3eaf33f7ea28e1

Observation e39d2557-3329-4f19-bc1b-9b02bb77a51b · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:e77110342b4d32c1b02cc3f137b546123afba94323bae28068ece424b79b89a2

Observation e433b4e7-168e-4078-ad68-e57a0bc95423 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.827430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.827430Z digest=sha256:e4f99dfd2ed6420d38fc9ec699cbf2f3b721e9778a0e11724427f4c99df4c8b3

Observation 36eec2bb-fe64-48d6-9715-b767de896496 · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:12.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:12.745471Z digest=sha256:f7c7177b52cc2aecdccb5b33de85fd6b83239df20a6acfb0bd13686053fc24d0