Pith. sign in

Paper Citation Record · LEDGER

Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2204.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2204.06601 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:45:12.378181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.237337Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc961c3e-c66e-430a-b4e3-ddb6d8c0a42c · inbound

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment cites this paper.

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:46:56.819389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T00:46:56.664582Z digest=sha256:5472932d4b053c544cbb97dc216875869c5242585e84c2f15e386c02a82c7c92

Observation ef303f7c-fb2a-4fb1-8fde-c11ed3a7cf1d · inbound

Aligning LLMs with Domain Invariant Reward Models cites this paper.

Aligning LLMs with Domain Invariant Reward Models Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:45:12.378181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:45:12.378181Z digest=sha256:02640b502460b0fde4cf9e6f1fd23a85b6e8a7680ecfe3f49b03f6dbae709b6d

Observation 2188078f-3ec4-47cd-8dcd-da45e6a778b8 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.259098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.259098Z digest=sha256:c0a6c88ea75db224a42eb4e183fb8a8cfb1439db7762b2705d0575b048da62b9

Observation 3a528559-9b07-4aa3-8710-8ad63f060415 · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:17.017683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:17.017683Z digest=sha256:286629993114a21ec25ee9006df8e2ffe03d4554559f53340a7297e49cdcfd6c

Observation 8873c3f5-2282-4a05-aae8-6661ff942edf · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.046448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.046448Z digest=sha256:28cc6b9b01e6e2240ff6bb0c20eb229ab7f8b3f0ca6d28d56d7950965463655a

Observation c39d7299-8903-4d26-857d-5f897ba91aaa · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.240433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:ec394706732a16271d809b2ea5d3839772a92837d2523d1d750eccceb04f0869

Observation d30fb4d1-e33d-41f5-81fb-75d86b60e1bd · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:11:24.068249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:4fbfc1c4cb3552e69301f4d0da9ba40d1eea22b8023e7d0691916c15a31cef95

Observation eba63cda-02e1-4808-a11d-e0a19d880f55 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.193120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:1fb0af3cf3fa4d2496cab414a8be26349d8a3e7b4c32d0d074ccbef9f5a21dd1

Observation 09df6ec3-a0e2-473b-add8-634482ab72bb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 173

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:ebfec2c53c9dfc5c855176f35f9ccb48435ee78134c563684c88739144e0a409

Observation a339ab8c-88e3-443c-9121-9a77aef70368 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Causal Confusion and Reward Misidentification in Preference-Based Reward Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.349555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.349555Z digest=sha256:451a67a73ec5e1be8436ee6bf38199479ea1a5b96fba2b0ce85a966b05b6f8a0