Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:55:50.816995Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 18 inbound Pith citation observations for arXiv:2501.09620.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:55:50.816995Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T14:57:53.946749Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:37:30.312996Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3a6c421a-c83d-4302-85ea-1f24dc104656 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38b2f93b-ab11-4cdf-9c78-7431c1d62087 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment @esa (Ref
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ee2420-aa68-4854-95aa-eb33d8d6a67b · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dbe0ca-683c-4034-b2d3-42bb3ea68842 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8d920591-75fd-4083-9629-c344ede01d84 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Maximum a Posteriori Policy Optimisation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77dd62a2-c20a-4a1d-a003-6ceb872ac68b · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5a344f-138c-4e2f-9a6d-51e9b53018d3 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Concrete Problems in AI Safety
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88b23e9-baaa-4345-a6a3-3101215e6698 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2540dc7-257e-45ac-b316-c2ab19d2dae7 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90dd914a-fa80-4438-866d-65a00fbd5eae · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Language models are few-shot learners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 769e38ec-1831-4bfe-a1e9-89e1706b2c4f · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15d8687-b085-4eab-a787-d23bd899b104 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment ODIN: Disentangled Reward Mitigates Hacking in RLHF
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46408abd-c50f-4c06-a372-618b79890ef3 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6244ab2f-b6b6-4069-a198-a47a7d4a8d81 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c74b67-2321-4a39-b155-cd3080f3a81a · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cff7834-fa54-4946-a4d0-037075d1eb66 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Safe reinforcement learning via hierarchical adaptive chance-constraint safeguards
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cc70e555-6f66-4030-95a2-4bc0fe774036 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Autoprm: Automating procedural supervision for multi-step reasoning via controllable question decomposition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6dd8f194-7b4b-4765-956a-db6620ff9bdd · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Deep reinforcement learning from human preferences
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3404ba26-329b-4d9c-837a-3c7f7aaaffc3 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Model Ensembles Help Mitigate Overoptimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14dcc32-55e8-4f20-a704-2ab909e015c6 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d999c63-65f3-4590-8f6b-99196f291410 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The Llama 3 Herd of Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7338bc79-c272-42ac-9bb1-8b5b832eedf0 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1ba73f6-2bdd-49da-af6a-727bb0eb8a0e · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 89e1677c-5e42-42ec-bb3a-0f46186df670 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34bfadea-eeeb-4c28-8695-0a95bc61f0eb · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a9433bc1-9081-4d35-9c8c-bc664127606c · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Shortcut learning in deep neural networks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf7c90a-8834-46dc-9ac9-af313b313a79 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Neural networks and the bias/variance dilemma
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1deb7543-d357-474e-bfd5-f1cc5cfabb73 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A kernel two-sample test
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d398aafd-8526-41b6-987f-44d8bb9e7caa · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cfff32fc-e56d-40b0-8ce9-264e346a4371 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment LoRA: Low-Rank Adaptation of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41ad49a-c8ed-4369-ad34-bb25b2c6aecf · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17741692-7502-49f9-89e7-24b3d285c676 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Post-hoc reward calibration: A case study on length bias
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc78dde2-7e0f-4400-bfec-fde963f7e795 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d568f85-6c94-4e6a-b226-9166371c21e7 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A survey of reinforcement learning from human feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a70a8d1-eb70-4760-aa90-fb1f2a78ec92 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning deep kernels for non-parametric two-sample tests
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 858d4867-4002-40b7-af1b-531e7e2fa7d3 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The flan collection: Designing data and methods for effective instruction tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2047617a-d80c-4de2-92ef-cb094551dda8 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning word vectors for sentiment analysis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac9bc59-1cd3-4dc9-b197-dc3518ab9b47 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Selection Bias Induced Spurious Correlations in Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c21f253c-5104-420d-86da-2cc3ba9f3680 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Human-level control through deep reinforcement learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe20df6-8fda-4895-b674-164c01b6e758 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Confronting Reward Model Overoptimization with Constrained RLHF
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd23faf4-4584-414f-a82c-f4b61699f793 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Understanding the Failure Modes of Out-of-Distribution Generalization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65dd5ed-6ebf-4b3a-99a9-bace86b688ce · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment GPT -4 technical report, 2023
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0833cf30-9ae1-4ac7-8acd-87ebe0c21a76 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 932b32c1-dc43-4b34-a1c5-07d554322208 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed95578-aa42-4051-afbe-d7589554ecf2 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Discovering Language Model Behaviors with Model-Written Evaluations
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ffadf8-a737-471c-9530-de0293a994cf · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Quantifying Generalization Complexity for Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69b90a22-bfcb-4f7c-824f-caadc3693d20 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning Counterfactually Invariant Predictors
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0cc014de-00bb-47ba-90f6-8da9e69882ea · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment WARM: On the Benefits of Weight Averaged Reward Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3b711b-9044-4903-b495-354df54afd59 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 009dd6da-996c-469c-bb53-e8c79b2328ac · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment why should i trust you?
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be3743c-59f0-4871-b68b-dedf19916374 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment The unequal opportunities of large language models: Examining demographic biases in job recommendations by chatgpt and llama
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b3f05908-d960-4fd7-94c9-b243523cc1c9 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Trust Region Policy Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 269aeae3-1825-4c3b-9bb2-c283ee0a828a · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Proximal Policy Optimization Algorithms
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82dd49e-e50e-46c1-9825-f29f98496273 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Towards Understanding Sycophancy in Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3b8895-aa70-41ab-a6f6-33d222789660 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d594349-e429-4e62-a5f7-ff968947da6e · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment A Long Way to Go: Investigating Length Correlations in RLHF
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd0adb2-76ca-4158-9e6a-3fa3ef21be1b · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Length bias in Encoder Decoder Models and a Case for Global Conditioning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ba22360-6a4c-4e71-9329-f12a59134437 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Learning to summarize from human feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ba12e12-63c1-4a06-ac77-e9e2c2d650ac · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Evaluating and Mitigating Discrimination in Language Model Decisions
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93758d4f-2e7b-4759-8d3b-c2f0be987a06 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Minimax estimation of maximum mean discrepancy with radial kernels
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51606330-7cbb-4b6b-b647-d4c6a93075b4 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Counterfactual invariance to spurious correlations in text classification
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 17a0ddab-9425-4747-922f-b6e90b486444 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Preference Optimization with Multi-Sample Comparisons
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6bfcad3d-662c-401d-a3df-ce6194fb54ba · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Reward Hacking in Reinforcement Learning , 11 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23cf8539-7052-461b-966c-6ec7be5df0a5 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b68ea24-ad1f-4bd3-b7e7-2431f6fb10cb · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Character-level convolutional networks for text classification
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92e50b6-da95-4b76-a598-0a96b768e662 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment GRAPE: Generalizing Robot Policy via Preference Alignment
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d5a7b0-3972-49ef-ab98-408310fb65e8 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ad59a22-7b4e-4f63-aea7-0ffdbb5787cf · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fada9a1-8fd4-4eda-854d-45d8663a82c9 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Explore Spurious Correlations at the Concept Level in Language Models for Text Classification
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d068e7-08ab-4d97-9bf7-ef9a775ddd34 · outbound
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment Fine-Tuning Language Models from Human Preferences
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1116a38-d5a5-4fda-93dd-2bd5414ad7a1 · inbound
GRAPE: Generalizing Robot Policy via Preference Alignment Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a272143-534e-48a8-8fcd-3a0e9f42ce18 · inbound
Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6c9ef9-e9d4-4d1b-8d73-7571e98f163e · inbound
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666bbd5d-4176-40a4-904b-9923fbbc7e61 · inbound
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f310af-a24e-4a49-bc3c-7c61467b464f · inbound
Token-Level LLM Collaboration via FusionRoute Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1652529e-63d4-48c6-b9f1-18c2e2152821 · inbound
Factored Causal Representation Learning for Robust Reward Modeling in RLHF Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 179f691f-920f-4e6c-940f-1558f0e63770 · inbound
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cfc6e527-9550-452b-97d5-ed00910a49c1 · inbound
Robust Reward Modeling for Large Language Models via Causal Decomposition Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 13eda639-c3b6-4123-b9ec-b038e7e81e47 · inbound
Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e9c2965c-b75e-43b0-9ebb-2b896aa8d476 · inbound
General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2cb5d842-b19a-408f-ae1c-5c2321cf1067 · inbound
General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6141da83-70ab-495c-873f-1f9cadf5b143 · inbound
General Preference Reinforcement Learning Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1e74ceca-1986-4219-9dab-cf06241319fe · inbound
Causality as the Statistical Conscience of Artificial Intelligence: From Pearl's Ladder to Trustworthy Machines Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bcaadb17-13e9-4e69-b5b7-dd25a1675cfd · inbound
Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 187781a2-cf6c-4319-a5dc-22d5ffae41ed · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0112c743-f68e-4928-a191-06466a8189c1 · inbound
Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043a5c66-d191-449a-afca-960a7c2be8ca · inbound
What do Reward Models Memorize? Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f904c68f-6879-4859-8ce7-4640b3670e3c · inbound
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.