{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:2NZEOYYHNW56QZH6ASGBYRYHUC","short_pith_number":"pith:2NZEOYYH","schema_version":"1.0","canonical_sha256":"d3724763076dbbe864fe048c1c4707a09c4ec3d5a6e0bb29b0d2ce20be38cb60","source":{"kind":"arxiv","id":"2409.13156","version":2},"attestation_state":"computed","paper":{"title":"RRM: Robust Reward Model Training Mitigates Reward Hacking","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Abe Ittycheriah, Anastasiia Makarova, Aviral Kumar, Bilal Piot, Daniel Sohn, Jeremiah Liu, Jiaming Shen, Jie Ren, Junru Wu, Lichang Chen, Mohammad Saleh, Rishabh Joshi, Tianhe Yu, Tianqi Liu, Wei Xiong, Yang Gao, Yuan Liu, Zhen Qin","submitted_at":"2024-09-20T01:46:07Z","abstract_excerpt":"Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences inde"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2409.13156","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-09-20T01:46:07Z","cross_cats_sorted":[],"title_canon_sha256":"ae68b1fe4ccfcb9d210c520dd01b2a65ef27b945538f11c44135b54efbfcf063","abstract_canon_sha256":"50761d2109bcefefca08f3f152f8c6b0689ec38af078e9b1dd2b36ab879170d7"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:21:04.427793Z","signature_b64":"4heL3jxt244pSYfUoYcWQplJ24UpE5osAyOtRkTymKyHfiUvbCaAP4BIMAovtlENIWKuEL5IO961mEyIzs5DCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d3724763076dbbe864fe048c1c4707a09c4ec3d5a6e0bb29b0d2ce20be38cb60","last_reissued_at":"2026-07-05T10:21:04.427259Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:21:04.427259Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"RRM: Robust Reward Model Training Mitigates Reward Hacking","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Abe Ittycheriah, Anastasiia Makarova, Aviral Kumar, Bilal Piot, Daniel Sohn, Jeremiah Liu, Jiaming Shen, Jie Ren, Junru Wu, Lichang Chen, Mohammad Saleh, Rishabh Joshi, Tianhe Yu, Tianqi Liu, Wei Xiong, Yang Gao, Yuan Liu, Zhen Qin","submitted_at":"2024-09-20T01:46:07Z","abstract_excerpt":"Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences inde"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2409.13156","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2409.13156/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2409.13156","created_at":"2026-07-05T10:21:04.427328+00:00"},{"alias_kind":"arxiv_version","alias_value":"2409.13156v2","created_at":"2026-07-05T10:21:04.427328+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2409.13156","created_at":"2026-07-05T10:21:04.427328+00:00"},{"alias_kind":"pith_short_12","alias_value":"2NZEOYYHNW56","created_at":"2026-07-05T10:21:04.427328+00:00"},{"alias_kind":"pith_short_16","alias_value":"2NZEOYYHNW56QZH6","created_at":"2026-07-05T10:21:04.427328+00:00"},{"alias_kind":"pith_short_8","alias_value":"2NZEOYYH","created_at":"2026-07-05T10:21:04.427328+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19818","citing_title":"Uncertainty-Aware Reward Modeling for Stable RLHF","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09711","citing_title":"Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09078","citing_title":"The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2601.21350","citing_title":"Factored Causal Representation Learning for Robust Reward Modeling in RLHF","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2507.06419","citing_title":"Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.06036","citing_title":"Optimal Transport for LLM Reward Modeling from Noisy Preference","ref_index":212,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC","json":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC.json","graph_json":"https://pith.science/api/pith-number/2NZEOYYHNW56QZH6ASGBYRYHUC/graph.json","events_json":"https://pith.science/api/pith-number/2NZEOYYHNW56QZH6ASGBYRYHUC/events.json","paper":"https://pith.science/paper/2NZEOYYH"},"agent_actions":{"view_html":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC","download_json":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC.json","view_paper":"https://pith.science/paper/2NZEOYYH","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2409.13156&json=true","fetch_graph":"https://pith.science/api/pith-number/2NZEOYYHNW56QZH6ASGBYRYHUC/graph.json","fetch_events":"https://pith.science/api/pith-number/2NZEOYYHNW56QZH6ASGBYRYHUC/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC/action/timestamp_anchor","attest_storage":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC/action/storage_attestation","attest_author":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC/action/author_attestation","sign_citation":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC/action/citation_signature","submit_replication":"https://pith.science/pith/2NZEOYYHNW56QZH6ASGBYRYHUC/action/replication_record"}},"created_at":"2026-07-05T10:21:04.427328+00:00","updated_at":"2026-07-05T10:21:04.427328+00:00"}