Pith. sign in

Paper Citation Record · LEDGER

RRM: Robust Reward Model Training Mitigates Reward Hacking

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2409.13156.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.13156 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:53.954441Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:30.665306Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31781822-1fb8-43b9-b342-af1f9554021d · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.954441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.954441Z digest=sha256:67b88346b81090f4735085daabda359acb109ab6d4d99a1e5ad665c9a2733163

Observation 3e72c6b5-9edd-4f59-b902-d1674575148e · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.490143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.490143Z digest=sha256:21bea7ebbeabc7a493ad0f93da376bcf0e786dfd1a4d8068582d6453babe9df5

Observation e836238a-5e6c-4473-87fd-580985d69e41 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.920441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:f642e5e44961d821940b993b82194b37b33fe20e478f6327f84b34bc62bac4a5

Observation 9236ba57-eb05-42d2-a5e3-a1d44d921573 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.095336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.095336Z digest=sha256:1b0da3e965dceaa13279b9600aa57ca9cce0bbb784f6ecd1da150c110ceb78a2

Observation 7f11790f-2146-4f4b-b693-e3c2f5d2a95e · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 246

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.186910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.186910Z digest=sha256:a8571baeeaa580136c56c1c12b11331a8f47fe7e756be22170fd357b7daa48f4

Observation 151a3599-716c-4f98-ad09-f762cebc36b2 · inbound

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction cites this paper.

Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:49:56.466791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:49:56.466791Z digest=sha256:20ec55ff3d1df1dbc49a7cc6eee13efa09d9ed20ec4a3e040e09d35a7655b7bb

Observation 38324b6b-7ff6-4626-aae3-e722bec70186 · inbound

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning cites this paper.

Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:14.851208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:44:14.851208Z digest=sha256:e0ac534978afaedffe533b6b20427511a532f8e114e159a79144c5fb8380e247

Observation a28b8faa-a12c-4835-9c20-62263506f7aa · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.682918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:b528de470647baf0a30dc89d4dde4e28ee2f10edb7620396ce9e5b3d23b5fc9d

Observation 6ef86500-44be-4aa3-b01e-412e34b3bb22 · inbound

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling cites this paper.

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:33.261612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:14:33.261612Z digest=sha256:c35c0a59bc1129d6aadef437f462f8fa99ea9587f0436407da6b6d51d92767bd

Observation 9f864ed2-9d40-4b23-ba63-4e6f83bd6e96 · inbound

Optimal Transport for LLM Reward Modeling from Noisy Preference cites this paper.

Optimal Transport for LLM Reward Modeling from Noisy Preference RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.323685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:5533983dcbf1cdd65ebce264591cc46638d91212bddf46ef0502c73f7c1623ea

Observation fa2ac8bf-e749-4359-8e2b-63111cc12bbb · inbound

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning cites this paper.

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:30.056298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:07:06.227106Z digest=sha256:284d1e79ddceda0f4208c1ddafe16c1cf65ae334356a0af2ce00b257ab1ba864

Observation 7d0e2e98-01af-477f-a1c4-d3df1ae905df · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.507576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:65a0d2e00a97583954a8d7c25065fd67cf9223acd936d87cbb0c5ef77fd17c39

Observation 4f0df22f-a659-4a08-9687-546640696a52 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:30.668236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:82eeec4308fc0ff22364188d27a512c6cc0fd56dee02c8db85240658efd12966

Observation 979b804f-5af7-42f0-8aa5-5ca53274541e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:b2b7818a800bba4b399105b26ba2e2c7ff4e3d816cbaa8239474902833a97346

Observation 3c81f251-c07b-4059-a2c3-79fcf80c5546 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:34.864629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:34.864629Z digest=sha256:5b05edc3ea55919affce2ae9fc83bcbca59316f6c0dc2487960b21e1cfc6c38f

Observation 62b0f854-1733-48f8-9274-0cffbcf5f435 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:18.584158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:18.584158Z digest=sha256:b66f7b98cbbbb77d4a56f7ecab0bb1db934537da02ea96e8c601706181427333