Pith. sign in

Paper Citation Record · LEDGER

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2501.09620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09620 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:17:58.970049Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.312996Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5f310af-a24e-4a49-bc3c-7c61467b464f · inbound

Token-Level LLM Collaboration via FusionRoute cites this paper.

Token-Level LLM Collaboration via FusionRoute Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:26:31.463111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T12:25:59.747665Z digest=sha256:f4f7d8e894d36908fa6e409428351935f29af4fb49b5110106eba7e825371e96

Observation 1652529e-63d4-48c6-b9f1-18c2e2152821 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:20:13.610834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:161e273b02afdba61431438479bfae8861ddc995b7c583c441dccfccb3a81d3a

Observation 179f691f-920f-4e6c-940f-1558f0e63770 · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:22:31.155448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:e1c6dd61f09920b6ebd8e96597108f951fc2bb3794fd1b8dd7cffda74b70c298

Observation cfc6e527-9550-452b-97d5-ed00910a49c1 · inbound

Robust Reward Modeling for Large Language Models via Causal Decomposition cites this paper.

Robust Reward Modeling for Large Language Models via Causal Decomposition Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:29.618255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:50:47.902911Z digest=sha256:c784e94cdcdd77d22054ab4576f533bd762cde1b3ac5fa98fad8924879e61b58

Observation 13eda639-c3b6-4123-b9ec-b038e7e81e47 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:49:10.189118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:05a6f32e5ec4e9c9fc02e9b2be7c6b6c1d110457344db51e0ff4f82308d01eb6

Observation e9c2965c-b75e-43b0-9ebb-2b896aa8d476 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:48:17.738781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:274b00ae10063d83d8f2fc61544f5bf634980a489a27a6cca9792a284ab399c9

Observation 2cb5d842-b19a-408f-ae1c-5c2321cf1067 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:54:02.919308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:d4f30a8abee921df3e5d27d418fb4c3c34afc90fb730bbdec1795fd244f244aa

Observation 6141da83-70ab-495c-873f-1f9cadf5b143 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:24:45.673645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:97f02dcc1f64b8b1da964e11cac9d14b8e2be224abc034d2e2a931de445ed580

Observation 1e74ceca-1986-4219-9dab-cf06241319fe · inbound

Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines cites this paper.

Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:14:47.427916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T15:06:58.205536Z digest=sha256:9160bd06fb454d796b5aa7d9fc579245687a6bfdcf2bff998279b1e3e8f6b240

Observation bcaadb17-13e9-4e69-b5b7-dd25a1675cfd · inbound

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure cites this paper.

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.235837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-29T12:22:06.635622Z digest=sha256:f7cc3a2f53b01e1f8aa08b023111631e28b59614b683f5b6f7c53b57b6381ec0

Observation 187781a2-cf6c-4319-a5dc-22d5ffae41ed · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.314429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:237a1573e47082b65af706ca04cf6eda1d31d58766c5a71822085d23e02a44a6

Observation 0112c743-f68e-4928-a191-06466a8189c1 · inbound

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation cites this paper.

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:17:58.970049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:17:58.970049Z digest=sha256:a9c3051544c8e45e7fb1b7e0114f544c85ffae8c82e9568d968b34410153be64

Observation 043a5c66-d191-449a-afca-960a7c2be8ca · inbound

What do Reward Models Memorize? cites this paper.

What do Reward Models Memorize? Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.895479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.895479Z digest=sha256:52bad8774447e87d7716ef41e750e07dc7f33c426c2abc7bfc30247e5fb56d5d