Pith. sign in

Paper Citation Record · LEDGER

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

As of 22 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 18 inbound Pith citation observations for arXiv:2501.09620.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09620 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:55:50.816995Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.946749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:37:30.312996Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3a6c421a-c83d-4302-85ea-1f24dc104656 · outbound

This paper cites write newline.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.441701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.441701Z digest=sha256:6b19abf9866cacff3f0d634f63bb165a6f43e569658bfa80f6d95d5810eaf41f

Observation 38b2f93b-ab11-4cdf-9c78-7431c1d62087 · outbound

This paper cites @esa (Ref.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment @esa (Ref

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.450259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.450259Z digest=sha256:58fea1d2b39ec3a4213781fbb013bc5065b508e10f21dd5cf7a15bf3fbc492fb

Observation 82ee2420-aa68-4854-95aa-eb33d8d6a67b · outbound

This paper cites an unresolved cited work.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.456732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.456732Z digest=sha256:ad8d1e6c34676a1bfae986d1ca3e6ecd8dbc05f29c202c70bfdd8c5b75a78260

Observation 63dbe0ca-683c-4034-b2d3-42bb3ea68842 · outbound

This paper cites an unresolved cited work.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T19:55:52.099876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.463959Z digest=sha256:833f993c2fdd90cfadf562fe787889e5ec774624a3d5d57441d5c261de7b6ac2

Observation 8d920591-75fd-4083-9629-c344ede01d84 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Maximum a Posteriori Policy Optimisation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.469992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.469992Z digest=sha256:c01bf5eb6bec2a7293460301be36f80a7fdf7ec4599b0b0c5188f926706cadce

Observation 77dd62a2-c20a-4a1d-a003-6ceb872ac68b · outbound

This paper cites BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.475790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.475790Z digest=sha256:c47d5066222598a9daca18042c6561c65d0f8e33820d9e7a93f345266924bbf4

Observation 2b5a344f-138c-4e2f-9a6d-51e9b53018d3 · outbound

This paper cites Concrete Problems in AI Safety.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Concrete Problems in AI Safety

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.481808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.481808Z digest=sha256:28c91900ad9ece94ac51cc1e474506bca0b1b95e94470ccdaf196636e37a5581

Observation e88b23e9-baaa-4345-a6a3-3101215e6698 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.488108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.488108Z digest=sha256:8f3b2b21a85a0c0b229c7a6e3372a82ac3e38143c8b8413474ba21e45c907997

Observation d2540dc7-257e-45ac-b316-c2ab19d2dae7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.494311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.494311Z digest=sha256:1b630996cd2f9105d118bb6efcf3b062f3034520162d12c6cf8d3447594e0b34

Observation 90dd914a-fa80-4438-866d-65a00fbd5eae · outbound

This paper cites Language models are few-shot learners.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Language models are few-shot learners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:52.083876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.500570Z digest=sha256:b9414075a023ffcf27d0ee65f0975787e5e997274ccd3db420c405034584f612

Observation 769e38ec-1831-4bfe-a1e9-89e1706b2c4f · outbound

This paper cites Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.506817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.506817Z digest=sha256:a6b4b857c0c0c3d51678f80a63432d7a88ad93d86b784f483e2a8dfdd5bf966a

Observation c15d8687-b085-4eab-a787-d23bd899b104 · outbound

This paper cites ODIN: Disentangled Reward Mitigates Hacking in RLHF.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.512927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.512927Z digest=sha256:ef849fdd3c14ff2decbb6258d14bca60d51c69f23be7770a081bebed2288ce03

Observation 46408abd-c50f-4c06-a372-618b79890ef3 · outbound

This paper cites MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.518773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.518773Z digest=sha256:5b4c7cce9ce45d9574cbb6d9154cddbf7549537b3a3b0e2e703041863da6acbf

Observation 6244ab2f-b6b6-4069-a198-a47a7d4a8d81 · outbound

This paper cites SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.524544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.524544Z digest=sha256:30f5b069cde8a7400cfcff2b44c42505bc0c012bbd98a76ce48db6433b385ca3

Observation 28c74b67-2321-4a39-b155-cd3080f3a81a · outbound

This paper cites AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.530651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.530651Z digest=sha256:286a9e073d83b0226a3d432e812d97b6f6785cf38d8a8d01797e611c57610d9c

Observation 6cff7834-fa54-4946-a4d0-037075d1eb66 · outbound

This paper cites Safe reinforcement learning via hierarchical adaptive chance-constraint safeguards.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Safe reinforcement learning via hierarchical adaptive chance-constraint safeguards

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:52.067315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.536595Z digest=sha256:c4be90ac230a55473a30142d1345ea29100288b691d155dd7089fd043408ac7e

Observation cc70e555-6f66-4030-95a2-4bc0fe774036 · outbound

This paper cites Autoprm: Automating procedural supervision for multi-step reasoning via controllable question decomposition.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Autoprm: Automating procedural supervision for multi-step reasoning via controllable question decomposition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:52.050547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.543604Z digest=sha256:d0f577cf1aee5b0fda79711ef3498f062a19b25da59e9549684665d05633d931

Observation 6dd8f194-7b4b-4765-956a-db6620ff9bdd · outbound

This paper cites Deep reinforcement learning from human preferences.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Deep reinforcement learning from human preferences

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.554122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.554122Z digest=sha256:5d55d613007af0a87bcc00447df20d31a2c4d6d02f948a713c29e95cc44ea1a2

Observation 3404ba26-329b-4d9c-837a-3c7f7aaaffc3 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.559707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.559707Z digest=sha256:c19abcf8b206db20907e3f61adfe9a409d458240291f00970bac9d035161962e

Observation d14dcc32-55e8-4f20-a704-2ab909e015c6 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.565621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.565621Z digest=sha256:85f4e65781c725d674af26a434ff555a5bb8764ed2a931ebb32ab6b11faf205c

Observation 8d999c63-65f3-4590-8f6b-99196f291410 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The Llama 3 Herd of Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.571884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.571884Z digest=sha256:8d4b31f03ef361edfc5f9c5daafb4c7ad2db1a1a338a83f772b7bfb5eca07d80

Observation 7338bc79-c272-42ac-9bb1-8b5b832eedf0 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.576818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.576818Z digest=sha256:b8d708f79d17f3342ce3aefd58bd4dc60606f390e52e69a287a13bf282eca8ce

Observation c1ba73f6-2bdd-49da-af6a-727bb0eb8a0e · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:52.023222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.582577Z digest=sha256:e169ad152a6a60d2294657037697f0f184cb070ba75eea2f5cd8498aaabed3d2

Observation 89e1677c-5e42-42ec-bb3a-0f46186df670 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.588363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.588363Z digest=sha256:28cfe52d6a1e1c6d508c3014eeaf5411f4842168791a5d5484ad92fc23c8eb4f

Observation 34bfadea-eeeb-4c28-8695-0a95bc61f0eb · outbound

This paper cites Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:52.006833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.593967Z digest=sha256:5410102577406593f56af7c589d0fc55ae7f6b8a64beed5982cb4c85302f4e53

Observation a9433bc1-9081-4d35-9c8c-bc664127606c · outbound

This paper cites Shortcut learning in deep neural networks.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Shortcut learning in deep neural networks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.598702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.598702Z digest=sha256:9e1493dd144eb630be23f25ba38d2b6738f415331d41cb80b4ecca04128f2b2e

Observation cbf7c90a-8834-46dc-9ac9-af313b313a79 · outbound

This paper cites Neural networks and the bias/variance dilemma.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Neural networks and the bias/variance dilemma

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.979172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.603264Z digest=sha256:630c6d41963cab47f06ee0f3fafd93a7ed79a51323cf9c9857326c02fd065149

Observation 1deb7543-d357-474e-bfd5-f1cc5cfabb73 · outbound

This paper cites A kernel two-sample test.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A kernel two-sample test

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.607999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.607999Z digest=sha256:93740d45678c6570eb7dd386b75146e5f0a34a1de4a032e791b3b8e2f38c84cd

Observation d398aafd-8526-41b6-987f-44d8bb9e7caa · outbound

This paper cites Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.951609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.612768Z digest=sha256:0bc9d0cf04fcab4a89fee0f4349b1ebfc3e535c487d9f87644e21f2a7f847b7d

Observation cfff32fc-e56d-40b0-8ce9-264e346a4371 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.618160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.618160Z digest=sha256:4da9fe54a654e27f5df2a0800eabbfac731dc6ab4b7cc82e43117820a06458a4

Observation e41ad49a-c8ed-4369-ad34-bb25b2c6aecf · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.623157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.623157Z digest=sha256:cbcc17c4fab33b9bdfe6d76f222a96b434b162116d4cefe47c6b432934772e40

Observation 17741692-7502-49f9-89e7-24b3d285c676 · outbound

This paper cites Post-hoc reward calibration: A case study on length bias.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Post-hoc reward calibration: A case study on length bias

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.628206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.628206Z digest=sha256:4f039129d8f9cdd88b1a93989cecab1959d2e73e4ccf01ba42ae2facb5b539ec

Observation cc78dde2-7e0f-4400-bfec-fde963f7e795 · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.632894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.632894Z digest=sha256:881a2600dc775ae3a18d5d1419a3718fd6094e4456905f9ee742bc979b2b7376

Observation 9d568f85-6c94-4e6a-b226-9166371c21e7 · outbound

This paper cites A survey of reinforcement learning from human feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A survey of reinforcement learning from human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.638057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.638057Z digest=sha256:36fecc0fb448c4bad609c51d4a293a5a7e7c1f81877f936f6bfa5b5910eedebd

Observation 8a70a8d1-eb70-4760-aa90-fb1f2a78ec92 · outbound

This paper cites Learning deep kernels for non-parametric two-sample tests.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning deep kernels for non-parametric two-sample tests

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.935036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.642619Z digest=sha256:b2aa716ea0245527c1943cf35d2a476d4934a595d5ce5282ba2ca76bb93d8be7

Observation 858d4867-4002-40b7-af1b-531e7e2fa7d3 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The flan collection: Designing data and methods for effective instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.918390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.647184Z digest=sha256:4428be103486797a86aafd7fbb3105a5e256d3f1ce70a159dd0812dfdc0f246f

Observation 2047617a-d80c-4de2-92ef-cb094551dda8 · outbound

This paper cites Learning word vectors for sentiment analysis.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning word vectors for sentiment analysis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.652035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.652035Z digest=sha256:7f92b62834440018d64df87323ca3b6898ab4fe5e78335421fcf05064ea8a4d0

Observation 6ac9bc59-1cd3-4dc9-b197-dc3518ab9b47 · outbound

This paper cites Selection Bias Induced Spurious Correlations in Large Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Selection Bias Induced Spurious Correlations in Large Language Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:55:51.239262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.656637Z digest=sha256:e6dc447902e68d7cadbd7a1d68c33af4312f2b770c46ecec0cae67a9cfa91bf1

Observation c21f253c-5104-420d-86da-2cc3ba9f3680 · outbound

This paper cites Human-level control through deep reinforcement learning.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Human-level control through deep reinforcement learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.661541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.661541Z digest=sha256:920d6e6d21656c35cb89bb6fb2e0854682d9cbbbe1388f5249ef08b2454b2361

Observation dbe20df6-8fda-4895-b674-164c01b6e758 · outbound

This paper cites Confronting Reward Model Overoptimization with Constrained RLHF.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Confronting Reward Model Overoptimization with Constrained RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.666316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.666316Z digest=sha256:0511b4f692b3fbdc400cad9d604ad30d480e56bd60e418af240d1c4b831f0b60

Observation fd23faf4-4584-414f-a82c-f4b61699f793 · outbound

This paper cites Understanding the Failure Modes of Out-of-Distribution Generalization.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Understanding the Failure Modes of Out-of-Distribution Generalization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.671247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.671247Z digest=sha256:c127f497ee794817555c9b3355d7809f018d55e021bcad4c8a3b5eb7e4228178

Observation a65dd5ed-6ebf-4b3a-99a9-bace86b688ce · outbound

This paper cites GPT -4 technical report, 2023.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment GPT -4 technical report, 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.875257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.676220Z digest=sha256:422d91461e333262fe2ff677a6d3ade87e020e626cc3ae06646e9f6b35b5c026

Observation 0833cf30-9ae1-4ac7-8acd-87ebe0c21a76 · outbound

This paper cites an unresolved cited work.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.680873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.680873Z digest=sha256:87ed492d7b7f2c867b01539e91d67463eab60a50f94b273860de73e88eb6efc2

Observation 932b32c1-dc43-4b34-a1c5-07d554322208 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.685830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.685830Z digest=sha256:028b8f8636574f9e92b85dc50d3148c620f9f5535fac0e85605c138922f29403

Observation 2ed95578-aa42-4051-afbe-d7589554ecf2 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Discovering Language Model Behaviors with Model-Written Evaluations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.691103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.691103Z digest=sha256:ae7b9e684fbc43a9edf89f445d3ef2c7eff55ef2a97a164775974bf2d3d1a0c7

Observation 50ffadf8-a737-471c-9530-de0293a994cf · outbound

This paper cites Quantifying Generalization Complexity for Large Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Quantifying Generalization Complexity for Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.696642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.696642Z digest=sha256:af8aa508596d5535feb89f5081ba9966b46d3753582f00287faab8a3466bc386

Observation 69b90a22-bfcb-4f7c-824f-caadc3693d20 · outbound

This paper cites Learning Counterfactually Invariant Predictors.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning Counterfactually Invariant Predictors

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:55:51.137086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.701703Z digest=sha256:f2486198ae7119e20e3327fd8ccb66cea8f19c79e6d93339e56b0c46f41ca606

Observation 0cc014de-00bb-47ba-90f6-8da9e69882ea · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.707005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.707005Z digest=sha256:5853741132eb20a55305d21ae0b18c58d21e5cae3bdaf47dedf5799ecf9919e1

Observation ee3b711b-9044-4903-b495-354df54afd59 · outbound

This paper cites When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.712217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.712217Z digest=sha256:15241ac6c6ce9537cd7bb837096d7800d6e6cf82d5266023095d367d0421183e

Observation 009dd6da-996c-469c-bb53-e8c79b2328ac · outbound

This paper cites why should i trust you?.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment why should i trust you?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.717585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.717585Z digest=sha256:6428c77582df85ab4f5e65f8bd67abd8fa8f098c16f752db1fc78ef3b0e7568d

Observation 2be3743c-59f0-4871-b68b-dedf19916374 · outbound

This paper cites The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.837589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.722457Z digest=sha256:d89508f0964fbec388fecf11bffbc4c197239d0b410ba8eb45e175272c961fd3

Observation b3f05908-d960-4fd7-94c9-b243523cc1c9 · outbound

This paper cites Trust Region Policy Optimization.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Trust Region Policy Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.727156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.727156Z digest=sha256:e395a92e9f3a9dc81c3c533fa697d6a882e2c89c10eb55f04b86b1f8bd9dcaa8

Observation 269aeae3-1825-4c3b-9bb2-c283ee0a828a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Proximal Policy Optimization Algorithms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.732285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.732285Z digest=sha256:d3fa58b310fe6de123a7d2f61a1f69a581f18c091ff31581caa4dc9ef90c1009

Observation b82dd49e-e50e-46c1-9825-f29f98496273 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Towards Understanding Sycophancy in Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.737082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.737082Z digest=sha256:4d181873103d61f6512164a3f61f1c3229dda8db9660a3d00eb0e2fc85f23f44

Observation 4f3b8895-aa70-41ab-a6f6-33d222789660 · outbound

This paper cites Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.742656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.742656Z digest=sha256:a9d2a7ccf4de91a08e6784ca8229f592f037092bb5f2279857f7cc8276a29211

Observation 9d594349-e429-4e62-a5f7-ff968947da6e · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A Long Way to Go: Investigating Length Correlations in RLHF

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.747644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.747644Z digest=sha256:ef9cf395d251b618e7d07ef66921fbc689edb1e20303e17e22881dedd12375ac

Observation 7fd0adb2-76ca-4158-9e6a-3fa3ef21be1b · outbound

This paper cites Length bias in Encoder Decoder Models and a Case for Global Conditioning.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Length bias in Encoder Decoder Models and a Case for Global Conditioning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:55:51.000468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.752651Z digest=sha256:6a194f1cd8685608eb4c1aaf901d59a58760dcd88dd2d6116917116d39cfb0f4

Observation 1ba22360-6a4c-4e71-9329-f12a59134437 · outbound

This paper cites Learning to summarize from human feedback.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning to summarize from human feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.757716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.757716Z digest=sha256:935e088e2ecb3f896fcd06ccb68d85c743f0ffc5bd3d5928bc0907be6d69fec9

Observation 2ba12e12-63c1-4a06-ac77-e9e2c2d650ac · outbound

This paper cites Evaluating and Mitigating Discrimination in Language Model Decisions.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Evaluating and Mitigating Discrimination in Language Model Decisions

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.762598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.762598Z digest=sha256:301f9494ce248006339b65b178e9179d4668c880b8f00050879acd3dba6f70b3

Observation 93758d4f-2e7b-4759-8d3b-c2f0be987a06 · outbound

This paper cites Minimax estimation of maximum mean discrepancy with radial kernels.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Minimax estimation of maximum mean discrepancy with radial kernels

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.767422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.767422Z digest=sha256:6eec9f2328be71fd60e95c75fdab6a075fc02ab61a415e2e91a878491a9f491b

Observation 51606330-7cbb-4b6b-b647-d4c6a93075b4 · outbound

This paper cites Counterfactual invariance to spurious correlations in text classification.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Counterfactual invariance to spurious correlations in text classification

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.809403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.772093Z digest=sha256:69a8715956efdf87eec94c322c72c4613df85b6e1583ef639e5f32eb57ec8555

Observation 17a0ddab-9425-4747-922f-b6e90b486444 · outbound

This paper cites Preference Optimization with Multi-Sample Comparisons.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Preference Optimization with Multi-Sample Comparisons

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:55:50.944111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.776578Z digest=sha256:b215c6382a97fda969012d541676106753bbb58cb453d2dcc699ea92d48c7175

Observation 6bfcad3d-662c-401d-a3df-ce6194fb54ba · outbound

This paper cites Reward Hacking in Reinforcement Learning , 11 2024.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Hacking in Reinforcement Learning , 11 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:55:51.791895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T19:55:50.782347Z digest=sha256:1b269bd97d6498197e2bfc573700bbee33071ab697ce784dcf0a081ae3b3124a

Observation 23cf8539-7052-461b-966c-6ec7be5df0a5 · outbound

This paper cites Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.786867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.786867Z digest=sha256:e7ecf1fa140d989c0715d8aeb4a79f245a46911c1dad07a0eb15b50d4f375fd7

Observation 5b68ea24-ad1f-4bd3-b7e7-2431f6fb10cb · outbound

This paper cites Character-level convolutional networks for text classification.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Character-level convolutional networks for text classification

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.792052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.792052Z digest=sha256:041db0873cc1514cdd3f68d6f04f3b96f68bd113c2aa5930172e16f3eaa94480

Observation d92e50b6-da95-4b76-a598-0a96b768e662 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.796899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.796899Z digest=sha256:dc582fef9ae139a4f46ef79cc1eebc63f19237d71468382af576d05a3e22e7ce

Observation d7d5a7b0-3972-49ef-ab98-408310fb65e8 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.801859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.801859Z digest=sha256:06e40553f2203fe684a36b3a4ecd06eb983e23c3190360f5ca713ef2a5493b93

Observation 2ad59a22-7b4e-4f63-aea7-0ffdbb5787cf · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.806681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.806681Z digest=sha256:1a84c0469093c8b64608dfb64a625f2216730a7fdd5e896ac5cef6daec585cd4

Observation 5fada9a1-8fd4-4eda-854d-45d8663a82c9 · outbound

This paper cites Explore Spurious Correlations at the Concept Level in Language Models for Text Classification.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.812045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.812045Z digest=sha256:ea775d4283efc8aeae65d99ca524eedd4543505a3081b2d33413060fc5858340

Observation 21d068e7-08ab-4d97-9bf7-ef9a775ddd34 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Fine-Tuning Language Models from Human Preferences

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.816995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.816995Z digest=sha256:478eff05fc5740522d12f83f6a1f9983b786e4a4456bc0ed56f2b322416ba257

Pith citing papers

Observation e1116a38-d5a5-4fda-93dd-2bd5414ad7a1 · inbound

GRAPE: Generalizing Robot Policy via Preference Alignment cites this paper.

GRAPE: Generalizing Robot Policy via Preference Alignment Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:23:15.764291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:23:15.764291Z digest=sha256:9d818d801a673887f38bbe409810325751fc066c606d7fcee57f2dd3eddf37d0

Observation 9a272143-534e-48a8-8fcd-3a0e9f42ce18 · inbound

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective cites this paper.

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:17.355112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:17.355112Z digest=sha256:6c8beccdc74a9ac7bdd7eb38673e53acf8fe965bbe061eb4d7c92529ef8833df

Observation 5f6c9ef9-e9d4-4d1b-8d73-7571e98f163e · inbound

Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization cites this paper.

Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:32:47.849443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:32:47.849443Z digest=sha256:12f0cae600ddbf6d984b69705cd3a17b528eac6f15eed5b86b5fb55c89cc2f7e

Observation 666bbd5d-4176-40a4-904b-9923fbbc7e61 · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:32.462708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:32.462708Z digest=sha256:bf2f323344d6be6687de3c1343c46f110234d29fdd1bbd91a67f70c407fdb32f

Observation c5f310af-a24e-4a49-bc3c-7c61467b464f · inbound

Token-Level LLM Collaboration via FusionRoute cites this paper.

Token-Level LLM Collaboration via FusionRoute Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:26:31.463111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T12:25:59.747665Z digest=sha256:c2a04c48378bfdf23b946234c8ad01ff38fda1285d19cf44fe3f8999e1a4a4c4

Observation 1652529e-63d4-48c6-b9f1-18c2e2152821 · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T14:20:13.610834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:cde01dd634a2088cc15e272a997429cac1389d676af18cee51b68fa373d7f044

Observation 179f691f-920f-4e6c-940f-1558f0e63770 · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:22:31.155448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:55849bd0d3f4a2d331075ac433cd670734f023b8e7545f16a3da30ed8710f1a0

Observation cfc6e527-9550-452b-97d5-ed00910a49c1 · inbound

Robust Reward Modeling for Large Language Models via Causal Decomposition cites this paper.

Robust Reward Modeling for Large Language Models via Causal Decomposition Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:29.618255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:50:47.902911Z digest=sha256:a72fdda686a36a8782406a8bc509e43f0875e0cf3b83b8bc45ca240cf3440111

Observation 13eda639-c3b6-4123-b9ec-b038e7e81e47 · inbound

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders cites this paper.

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:49:10.189118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T22:48:54.238767Z digest=sha256:03fc5bdfcc35e16d3930dc63e5e4f0a710ead4e39c848649d2766e5b58f62540

Observation e9c2965c-b75e-43b0-9ebb-2b896aa8d476 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:48:17.738781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:f461d6064d854063b0bc40dd890db7fcbc5fe8609a3858f1823bec7b536d2987

Observation 2cb5d842-b19a-408f-ae1c-5c2321cf1067 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T07:54:02.919308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:b754a21c34192f72eb925dc76905c869f2184031d6432fada825709c2d159076

Observation 6141da83-70ab-495c-873f-1f9cadf5b143 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:24:45.673645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:4940c4f5fe9545a6123d16b54af6c6f79147118813aee562c39552044f107d94

Observation 1e74ceca-1986-4219-9dab-cf06241319fe · inbound

Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines cites this paper.

Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:14:47.427916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:06:58.205536Z digest=sha256:31bfe3a55560a3a4fd414d2dcd4854daafa8e41691a6176c1b4701155148fc5b

Observation bcaadb17-13e9-4e69-b5b7-dd25a1675cfd · inbound

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure cites this paper.

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:23:24.235837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T12:22:06.635622Z digest=sha256:fe15b5b3a1f0c34cb8f6ada65d8613d48fc507a280d0d4af475ffcc28fe8af00

Observation 187781a2-cf6c-4319-a5dc-22d5ffae41ed · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 135

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.314429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:5ddb24da52b154050afc2dd5f360f9336599f2184e7c4d1d28ec97958fe9990f

Observation 0112c743-f68e-4928-a191-06466a8189c1 · inbound

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation cites this paper.

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:17:58.970049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:17:58.970049Z digest=sha256:4b9729786b35f6d0b9108cc412f981d527b7e69a4871c0600b7dc2087328e8dc

Observation 043a5c66-d191-449a-afca-960a7c2be8ca · inbound

What do Reward Models Memorize? cites this paper.

What do Reward Models Memorize? Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T13:41:42.895479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T13:41:42.895479Z digest=sha256:a42c033bb7ec5f9bd5f4f9279b72f42b881c70f44cb5b0a8ab0f2527e6fe5c72

Observation f904c68f-6879-4859-8ce7-4640b3670e3c · inbound

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation cites this paper.

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T14:57:53.946749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:57:53.946749Z digest=sha256:753e719806cb614b65816a3dcb47b0874bf3777550fe5629121ef4613bb976b4