Pith. sign in

Paper Citation Record · LEDGER

Secrets of RLHF in Large Language Models Part II: Reward Modeling

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2401.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06080 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:48:18.939588Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 73fa1a52-2f38-4c64-a58c-de79b310f68f · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.707958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:9522ba8df3abe95a78069ae5216ab96540cdde1515a69d67b5389cd27c8acf3b

Observation 347c6dc4-9bba-42d0-ba41-9ebeae1bdb67 · inbound

CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback cites this paper.

CHAI for LLMs: Improving Code-Mixed Translation in Large Language Models through Reinforcement Learning with AI Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:10:44.193768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:10:44.193768Z digest=sha256:7ef127e2245ef4467f5a6e77927665128b80730a4310a08a33da7d7e4744bb60

Observation 0b09bcf4-e9eb-4ddc-840d-b79250750aaf · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.483570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:a13721a29dfad02d7fd5a17e0cc8032f2c29662154874a13469a6b0b0fc7d78c

Observation 860f904f-f64d-495f-a931-c8ba2e8d4bba · inbound

AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data cites this paper.

AddrLLM: Address Rewriting via Large Language Model on Nationwide Logistics Data Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:09:57.843723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:09:57.843723Z digest=sha256:750a64d74872ad2406b272d5bdc414ef3ebf9686a7699e182c50fc644476a1a8

Observation 82e6b495-acfd-4179-9d86-53ba1ebc1241 · inbound

T-REG: Preference Optimization with Token-Level Reward Regularization cites this paper.

T-REG: Preference Optimization with Token-Level Reward Regularization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:56.217162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:15:56.217162Z digest=sha256:aab327fc3a925b5e4e08601e46a6fad496727501c2a6097c088ffafcddc1c806

Observation 9902bb56-b7b3-4533-99f5-7512447e3c57 · inbound

Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective cites this paper.

Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:30:38.395214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:30:38.395214Z digest=sha256:6d86692a29881884ae687942c2ec7610049e58f15d304072ce33ca27a4d4d7a0

Observation 6fc7a908-e2fc-48ce-bc07-3bf1539d8369 · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:25:27.899604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:c423a439a2329fc09f7e82fac08aeb29a7607a3bfd4e337b735a44fc7f2aef44

Observation b44864e9-f666-4875-a8b8-13dd0bb47821 · inbound

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples cites this paper.

MPPO: Multi Pair-wise Preference Optimization for LLMs with Arbitrary Negative Samples Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:21:35.486443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:21:35.486443Z digest=sha256:dbdf56190667105509f4f8d8c7d0de9d0c1b18be1aafe750e8eebdaee4e89e9e

Observation 76c5024b-02f9-4baf-a5f8-8d3e4a8d51c7 · inbound

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models cites this paper.

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:28.361176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:28.361176Z digest=sha256:ba88928015d15e9d48cb0539ab3a43cb5fae0c898967cd69f6754d4bedb41e89

Observation c464e561-8a36-493e-9a92-9e34e96cbeae · inbound

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems cites this paper.

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T22:51:52.569666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:51:52.569666Z digest=sha256:9163d13e6f7a8fc3d4f6bd7ef9ba5dffbb99b0afb82ab015e7d363b47c939c35

Observation 6e25ac54-8e66-427a-82bc-72e3558795b2 · inbound

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information cites this paper.

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:56.524700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:56.524700Z digest=sha256:ca69aa8eaef4d7931d088d6e2d7201c7fa3c747d0f4047b1b6652bc7c171f49e

Observation 418f59c4-f5d4-4007-8983-11201a86c7f9 · inbound

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament cites this paper.

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:31.249038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:39:31.249038Z digest=sha256:4998e5bd02d70133f27b99a483b16dd07aa096f1241a1b6f4c641315a35b0cab

Observation 98c2e227-9c84-4ac3-b7b7-dc299e2ecb9b · inbound

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking cites this paper.

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.351960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.351960Z digest=sha256:37808879c6619a907fd6aee4d69c0c20de5847006896821e743b5ad0e775a3b1

Observation 791982b8-e588-4887-bc77-7e90d866e91f · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.634354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.634354Z digest=sha256:e951a3ee2885f67c6fce5c66d8ddc5a10ad29a807b508881485b89366458541b

Observation e5842996-e851-443c-bc8d-630bf3aa9e5c · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.959132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.959132Z digest=sha256:325c44eca6a0ba4ab9e85e8fa6393823ead6555606abca31268ddc48f93c96f4

Observation 95f05551-4bd2-4d8b-8c57-24adc9473adc · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.575352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:acdaa98d2b2cc7db7068af1e55cf2507674ceaa796157745d597e6170e012f8a

Observation 67d604db-332f-4d2e-a985-bce45b67d65c · inbound

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions cites this paper.

Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:35.445137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:35.445137Z digest=sha256:6b4a332f37b6b872ce861d30d6996429ded199e078dd25d006c6cee98c2c8d65

Observation b94a2143-d519-4a37-9fa7-fbb3ac0946b1 · inbound

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples cites this paper.

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T11:57:11.983807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:57:11.983807Z digest=sha256:53e81570f7001531e84f9396e4f6ff68c1192724dc8b0a9c0c95ebe5569ba819

Observation 36ebacfc-d809-4df6-8546-1096a4895a80 · inbound

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law cites this paper.

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-16T00:48:18.939588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:48:18.939588Z digest=sha256:01f5a9e746f718859c841cc4f2acd9e0c69cc0bfeb790fec9bcdcd1458bee3b2

Observation ff05ea61-5cf2-4448-bb54-cb05ad9c6d9f · inbound

Learning Guarantee of Reward Modeling Using Deep Neural Networks cites this paper.

Learning Guarantee of Reward Modeling Using Deep Neural Networks Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:53.893955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:53.893955Z digest=sha256:b785a64a221052c67935919f1a98f812dc0544f704b60e60eeffa69f3550901f

Observation c93a8b98-a211-439d-86d1-bf05ae4a3959 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.576663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:73345bed28e7a20b9ca0729620e29246f1c7be3aa9f54d48c1cebabfcc7c8cd5

Observation 0246dc54-b726-4841-9133-fb45a8521d82 · inbound

On the Robustness of Reward Models for Language Model Alignment cites this paper.

On the Robustness of Reward Models for Language Model Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T22:26:59.422007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:26:59.422007Z digest=sha256:eb2bc362f042694767ff4d976232bb054f40f08020d5327ad36f99989a9bbad5

Observation 21ad8af1-130a-4f9f-bb6f-4d01402cdb91 · inbound

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment cites this paper.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.292618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.292618Z digest=sha256:55a780610a23b6fdb401d9c270f392ce83a394ac63bfcb6c4129afffe39331d1

Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624148Z digest=sha256:5fe5a415b28a1461f5a1a3838053e0c0273a0214ffb33aef36650673ef5e4de4

Observation 987cc69d-84de-4dab-9a7e-544ac5e73b2c · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:31.942943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:31.942943Z digest=sha256:9deb1b0f5d9a51953a62bcec23b81f0c669cf46be0059f817ff0b48d6ca8a9fd

Observation 5fea45f8-50fa-4f09-9407-96a73b02ce3a · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.352622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:396b194aae1050ef9548be693d51d9cd405512381fcaf098d8f14f6449edcbd2

Observation 3952763d-ace7-4824-9572-3163cc78be12 · inbound

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing cites this paper.

Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:15:55.913770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:15:55.913770Z digest=sha256:d212334549c537e2823aaa7b902577264b439931bd63ba47e12a5b15c361be8c

Observation 0b0cd5c5-7278-4c5a-a7b3-0ea7ca8dfc02 · inbound

Pairwise Calibrated Rewards for Pluralistic Alignment cites this paper.

Pairwise Calibrated Rewards for Pluralistic Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T20:48:01.949682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:48:01.949682Z digest=sha256:1479e095f7935bc126a51a16c6ddcb8cd4571b845fb4723b69b62cd86672dffe

Observation 712e93bf-a4be-4b77-8f3b-32299be42706 · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:56.965198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:33:56.965198Z digest=sha256:7ec9b53af546f6aff12df28f39bf9fee8f60ba4966bb08ce70bbfd6747c274f7

Observation 39dc39f2-8afa-497d-ba50-1883950b60d0 · inbound

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding cites this paper.

SymbolicThought: Integrating Language Models and Symbolic Reasoning for Consistent and Interpretable Human Relationship Understanding Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:57:28.485425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:57:28.485425Z digest=sha256:0b5dc32a7e9c26a1fd92d24f33b892d81c76b40814797f5052cb3288836a6830

Observation 27fdfeb5-0952-486b-a2b7-02ddb351d597 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:30.462615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:30.462615Z digest=sha256:9e51245f17bbfb0692c1047f27fbad8ca7a87bb34e9b3d878071da9bbe92dc86

Observation 6e2155e5-e488-4033-a71f-79949deb1a4c · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:22.806003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:22.806003Z digest=sha256:0f98a01adcb89d86482f317ca3dc50fabd2c2529ca28755390c31495019c872d

Observation b1810d4c-e22a-4a9e-b1f4-565668a357f8 · inbound

Users as Annotators: LLM Preference Learning from Comparison Mode cites this paper.

Users as Annotators: LLM Preference Learning from Comparison Mode Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:21:07.012995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T08:19:58.093621Z digest=sha256:7f9332a0684f9ff9811f191634771d44534cbbde7f91714933fd1404caa65483

Observation b3c752b3-7b51-4c17-af23-0b4694293918 · inbound

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration cites this paper.

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.419338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T16:07:20.180349Z digest=sha256:0b5b0af0cbe3876c0687637e9a7321fad4fc4291abb48b80f5913184d31464de

Observation ee674794-cfa2-4c6b-abdd-e3641071a541 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.159008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:04d05061b4fc57d81c34a98883d32fabe698f648d2e2440d793288a6c7efddd8

Observation d3ebbd09-47a7-412e-97bf-eac652a28814 · inbound

Reinforcement Learning via Value Gradient Flow cites this paper.

Reinforcement Learning via Value Gradient Flow Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:25.604855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T13:18:16.532434Z digest=sha256:31f4de2512a151a0013f3fbd314a1328763d389fa1f0930fe6b317f55378fd18

Observation 8bdc2aa2-39d8-4cc2-a6a1-fa90e214838c · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.869033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:3c51fab007816abbdbfca51650bf2588a2c69d47e00c533c580b4cf59b34a949

Observation 3a406e49-0558-4235-92b8-ab3799e9c7d8 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.718303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:0e238978714a3901592636a52efcff77759f5780097eb0500af334561f3bc5ef

Observation 9dcd4027-37e4-456c-b1e5-32c00cd4bbd3 · inbound

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity cites this paper.

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:16:18.568531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:16:09.213612Z digest=sha256:1f1b0f8cc070422e73c540ce3b71f9f0b3ef20a9e3fa697924dcfb06f6987a2d

Observation 9e9cfa9b-6513-47ce-a9ae-ac8768437ab2 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.945673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:ac89a5ebbc6f9a8fdae97b795b936a9e30659546a002dec1fb6cfe5a8272a976

Observation 90b60aec-fea5-48c8-a609-4e07d988c717 · inbound

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance cites this paper.

Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.017450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T03:18:26.590871Z digest=sha256:020e661e4249a861c3732d3765718c6291d0cd02c500a8f1f79bfd33fdc76dd4

Observation 8fe3cacf-d709-4e12-97df-291176f3fe31 · inbound

Boosting Self-Consistency with Ranking cites this paper.

Boosting Self-Consistency with Ranking Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 185

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T06:51:44.296018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T06:49:58.051659Z digest=sha256:2aa8820e8a7dc020202fc045c8a377f6ce43dfd3928743ed6070933a6bb074f3

Observation 7b24150e-32d1-49a0-aa08-65e809dabaa6 · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.957727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:31a31e26d149137314301be1232f47d7cfe251794c2929410617ef7445027e98

Observation bcf61e05-b703-4838-96f2-ca6ee50f3ef2 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 115

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.515245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:479d4b145441244dc0a594074dd4d37568ad8f2b176d9d701de1d7b022c1f5d2

Observation 24da175c-92c6-47ff-9ec0-a4ccd107e1a7 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.282385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:5358ee3496188fe5aac95ebfb029bf1a98b18c4452d912497407de1722296759

Observation 84b2b752-56a8-4a70-ac76-0f528db5dc7a · inbound

Understanding helpfulness and harmless tension in reward models cites this paper.

Understanding helpfulness and harmless tension in reward models Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:22.991762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T07:06:23.680572Z digest=sha256:e271200861ae49585d55db342ad960586370bca970855ac0465a4449ec28cc2a

Observation 33796558-f5e7-46a1-87c9-9ca4b3415e13 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.105153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:efebdd2cf9c08c92ce61725e282b61243c9563e30ee929567d466acf66a06c85

Observation 90eb243c-e305-431b-99ab-b8b2084915af · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.135832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:2e32d36c301475da0d5c00ae380f0cbada119b41d97c8dc8301c418153a87cd4

Observation e39d2557-3329-4f19-bc1b-9b02bb77a51b · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:365e3c4d4406e138862cad63303fddf9b51fe690299245545afb49e0af91b0f2

Observation e433b4e7-168e-4078-ad68-e57a0bc95423 · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.827430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.827430Z digest=sha256:03a6f5451e7da1781a7d01a4510f25c8a3bdefd04488f1d251a93a463a0cc982

Observation 36eec2bb-fe64-48d6-9715-b767de896496 · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:12.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:12.745471Z digest=sha256:bb916866ea3f48a1b310b4d3b24f261fdffe9265cbb15b7d1e41b8f1b8859b30