{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:ZGHDMJINES6QP4L72YYOOCBRMT","short_pith_number":"pith:ZGHDMJIN","schema_version":"1.0","canonical_sha256":"c98e36250d24bd07f17fd630e7083164ee83aa323b39514abc1c025a0520a82c","source":{"kind":"arxiv","id":"2309.06256","version":4},"attestation_state":"computed","paper":{"title":"Mitigating the Alignment Tax of RLHF","license":"http://creativecommons.org/publicdomain/zero/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Hangyu Lin, Hanning Zhang, HanZe Dong, Han Zhao, Haoxiang Wang, Heng Ji, Jianmeng Liu, Jipeng Zhang, Nan Jiang, Renjie Pi, Rui Pan, Shizhe Diao, Tong Zhang, Wei Xiong, Wenbin Hu, Yong Lin, Yuan Yao","submitted_at":"2023-09-12T14:16:54Z","abstract_excerpt":"LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. To investigate alignment tax, we conducted experiments with existing RLHF algorithms using OpenLLaMA-3B, which revealed a pronounced alignment tax in NLP tasks. Whereas, despite various techniques to mitigate forgetting, they are often at odds with the RLHF performance, leading to a trade-off between alignment performance and forgetting mitigation, leading to an alignment-forg"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2309.06256","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/publicdomain/zero/1.0/","primary_cat":"cs.LG","submitted_at":"2023-09-12T14:16:54Z","cross_cats_sorted":[],"title_canon_sha256":"c01d0956b9fdf046bcf7c956980d437c068d9b9c8a64f396fd7f4ce9d4f70a86","abstract_canon_sha256":"23a0baf6b0f2ca6e9953cca7704bc3e20254b763729480d4590eaa20715448ae"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:19:51.969818Z","signature_b64":"95PgP8ng/jCd2p9p8sAPGDDiHMarMznT3reNX4vb4611s4DobSC+NWfUfdoJyRuJs+vo4QAgeb4GsS8nL8sfBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"c98e36250d24bd07f17fd630e7083164ee83aa323b39514abc1c025a0520a82c","last_reissued_at":"2026-07-05T09:19:51.969308Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:19:51.969308Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Mitigating the Alignment Tax of RLHF","license":"http://creativecommons.org/publicdomain/zero/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Hangyu Lin, Hanning Zhang, HanZe Dong, Han Zhao, Haoxiang Wang, Heng Ji, Jianmeng Liu, Jipeng Zhang, Nan Jiang, Renjie Pi, Rui Pan, Shizhe Diao, Tong Zhang, Wei Xiong, Wenbin Hu, Yong Lin, Yuan Yao","submitted_at":"2023-09-12T14:16:54Z","abstract_excerpt":"LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. To investigate alignment tax, we conducted experiments with existing RLHF algorithms using OpenLLaMA-3B, which revealed a pronounced alignment tax in NLP tasks. Whereas, despite various techniques to mitigate forgetting, they are often at odds with the RLHF performance, leading to a trade-off between alignment performance and forgetting mitigation, leading to an alignment-forg"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2309.06256","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2309.06256/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2309.06256","created_at":"2026-07-05T09:19:51.969378+00:00"},{"alias_kind":"arxiv_version","alias_value":"2309.06256v4","created_at":"2026-07-05T09:19:51.969378+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2309.06256","created_at":"2026-07-05T09:19:51.969378+00:00"},{"alias_kind":"pith_short_12","alias_value":"ZGHDMJINES6Q","created_at":"2026-07-05T09:19:51.969378+00:00"},{"alias_kind":"pith_short_16","alias_value":"ZGHDMJINES6QP4L7","created_at":"2026-07-05T09:19:51.969378+00:00"},{"alias_kind":"pith_short_8","alias_value":"ZGHDMJIN","created_at":"2026-07-05T09:19:51.969378+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.26102","citing_title":"Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29706","citing_title":"ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2603.06610","citing_title":"CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2509.03403","citing_title":"Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12484","citing_title":"Learning, Fast and Slow: Towards LLMs That Adapt Continually","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11679","citing_title":"Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12484","citing_title":"Learning, Fast and Slow: Towards LLMs That Adapt Continually","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11679","citing_title":"Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19087","citing_title":"OLLM: Options-based Large Language Models","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.17497","citing_title":"Generative AI Technologies, Techniques & Tensions: A Primer","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT","json":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT.json","graph_json":"https://pith.science/api/pith-number/ZGHDMJINES6QP4L72YYOOCBRMT/graph.json","events_json":"https://pith.science/api/pith-number/ZGHDMJINES6QP4L72YYOOCBRMT/events.json","paper":"https://pith.science/paper/ZGHDMJIN"},"agent_actions":{"view_html":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT","download_json":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT.json","view_paper":"https://pith.science/paper/ZGHDMJIN","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2309.06256&json=true","fetch_graph":"https://pith.science/api/pith-number/ZGHDMJINES6QP4L72YYOOCBRMT/graph.json","fetch_events":"https://pith.science/api/pith-number/ZGHDMJINES6QP4L72YYOOCBRMT/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT/action/timestamp_anchor","attest_storage":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT/action/storage_attestation","attest_author":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT/action/author_attestation","sign_citation":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT/action/citation_signature","submit_replication":"https://pith.science/pith/ZGHDMJINES6QP4L72YYOOCBRMT/action/replication_record"}},"created_at":"2026-07-05T09:19:51.969378+00:00","updated_at":"2026-07-05T09:19:51.969378+00:00"}