Pith. sign in

Paper Citation Record · LEDGER

Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2204.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.06601 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:06.259098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.237337Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc961c3e-c66e-430a-b4e3-ddb6d8c0a42c · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:46:56.819389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:8161b4222437ecc38ef0de5ad3da8ce74b1fe6d3f16f3b23a0faeed1b2c5cd63

Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.259098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.259098Z digest=sha256:aa03424fa4f29bfa1c5aa9c922e21bee06cd9fca35e7068a504b920263a3cc19

Observation 3a528559-9b07-4aa3-8710-8ad63f060415 · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:17.017683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:17.017683Z digest=sha256:60c96bd6e0db33eeae42de22371fe28984e85e8454834d989ca6567bc8bfca69

Observation 8873c3f5-2282-4a05-aae8-6661ff942edf · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.046448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.046448Z digest=sha256:bdbe179dbe2671d98e6c1eaab0bea7d93add77e9b73bd38234923b2b0121a9c8

Observation c39d7299-8903-4d26-857d-5f897ba91aaa · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:2e532dd7ddcbac6174b1aab2771e53defd01d1007d5cf77351285db99939765b

Observation d30fb4d1-e33d-41f5-81fb-75d86b60e1bd · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:24.068249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:f7ccc8c953406d999c0d8db6b89327f6cc61646db4412be556fa6e0a5e82f2aa

Observation eba63cda-02e1-4808-a11d-e0a19d880f55 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.193120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:bfac8556c24ab2d37e9ee7b7014c08364d2b39e09efb364b34a68ffdc25ce776

Observation 09df6ec3-a0e2-473b-add8-634482ab72bb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 173

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:e8ea706d6356d9e5a3c3e6566c2e1b36e892a96efe45f354521e9380cb3d3664

Observation a339ab8c-88e3-443c-9121-9a77aef70368 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.349555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.349555Z digest=sha256:b13002073824d5717a602509280d63b7989a20e5bf465db9cc899a1fdce75ecd