Pith. sign in

Paper Citation Record · LEDGER

ODIN: Disentangled Reward Mitigates Hacking in RLHF

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2402.07319.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.07319 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:01:57.056053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:46.669571Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 45abfca5-0ae6-49c3-943e-6ac160b9dfdd · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.979266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:c07838f266daca6cc016ad2c7b03b9bc7306a745c7093fe9fddbd3e0ef76f5bb

Observation 2e00a555-b139-4013-8ce2-8655891135f1 · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.862282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:4d48878320438803782a958dc3310f0d28962d69abd333c64056bc48ad12fd05

Observation 163fec65-6cc6-462f-b789-08baac871e75 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:25:45.371973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:a9cf072abed5c1f7ac23e745db9c461574793443ae0006caa1152465809feb95

Observation 674e5b0e-29ed-44e1-8a23-8126f4dbeea9 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.711581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:a61ff007cd7271cdf014658cff844b2d0fd60aceb169f847581db139858b39ed

Observation 29b958c3-ab11-43ba-bbbc-8b7b9dceb1b5 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.555672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:e0d3f623ef2f6b507b53d334d8e6bf83d7ad8eede2a207e03a1a23b0de7234a2

Observation 0ae14f58-bc00-4810-96d7-48ceae5e544f · inbound

RVPO: Risk-Sensitive Alignment via Variance Regularization cites this paper.

RVPO: Risk-Sensitive Alignment via Variance Regularization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:08.473849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T15:00:11.237293Z digest=sha256:57ce1b0cfd8e5065321b73847d834d85b5bf29f6567bc609c33b76d4e9ba16ae

Observation e6bc8d08-6f0e-4077-aa87-a0c5207aa80b · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.073790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:e2482b564f900061e042300498ba2a7b4e9ee8bc4e276642193a5643e2c93b8b

Observation ec71211d-e745-4e88-9e62-9c33004311fc · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.527008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:8fadc97f89b9890e1eaf7a22aef6eb2f592e2eb75ea92676f992a39d42cec4cd

Observation 5eca7ad6-eb1a-4f45-9ac3-6ff0ab240b1b · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.735216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:aea086908112cf1ff1dd3e88962815f4eacb026cce89178865b975288bdbc522

Observation e9f9d502-d305-4086-9927-bedb017f52d8 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:54:02.859083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:16f2ec42222d0f2b5c9d55708dbff86055822ba937c7f748e32d3a7b11238c53

Observation 184967ed-9231-46dc-8553-08bb083bfc90 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:24:45.746637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:5fc3e0780986d5fa97d9cab5867950e2766174b783cf2bcf807e4963321bef32

Observation 86b8fada-495c-4ee4-8766-514451a80c4e · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.505006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:abf5917d65de9fc9f27c2afb5c8c5fef9681d9dcb3b71e1c1352a3ed4aec7064

Observation 16d00a7b-976c-4d46-909b-80e40d11902d · inbound

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling cites this paper.

Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:46.671372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T17:56:48.024298Z digest=sha256:38d859038ec8f5e9ed64d69d620fc0984aae49470e30d94d2f6d00a2f0941b66

Observation e09f281b-e0e9-4d96-8bb5-3d7d060122fa · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:1672ee8fc71e9974abe4e16634bfe4edfebf79220eb62707baed49ea555a868e

Observation 768ec640-82d1-4148-9f71-238efeaac97c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:38.315887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:38.315887Z digest=sha256:4d81c5487e96cc676b42c016ec768b50a73dcde3f124ad3c110c43f38ccc70de

Observation de695394-6551-43ea-8b6d-e119fa33956d · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:57.056053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:57.056053Z digest=sha256:16b242821d9f785894a67f9074c7e470aa7a151b2e5d1bd02e9bff0517fffe02