{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:V7DIRKHF63DV4UZRGGQ3RZHDUF","short_pith_number":"pith:V7DIRKHF","schema_version":"1.0","canonical_sha256":"afc688a8e5f6c75e533131a1b8e4e3a1549661de586e241ce589ec2255a08724","source":{"kind":"arxiv","id":"2402.08925","version":2},"attestation_state":"computed","paper":{"title":"MaxMin-RLHF: Alignment with Diverse Human Preferences","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG","cs.RO"],"primary_cat":"cs.CL","authors_text":"Alec Koppel, Amrit Singh Bedi, Dinesh Manocha, Furong Huang, Hui Yuan, Jiahao Qiu, Mengdi Wang, Souradip Chakraborty","submitted_at":"2024-02-14T03:56:27Z","abstract_excerpt":"Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data. However, such an approach overlooks the rich diversity of human preferences inherent in data collected from multiple users. In this work, we first derive an impossibility result of alignment with single reward RLHF, thereby highlighting its insufficiency in representing diverse human preferences. To provide an equitable solution to the problem, we learn a mixture of preference distributions via an expectation-maximization algorithm and"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.08925","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-02-14T03:56:27Z","cross_cats_sorted":["cs.AI","cs.LG","cs.RO"],"title_canon_sha256":"ce17d20d5e051d979c0f9718567004098839f8a9c0c00e9b16aed7cd5051eb75","abstract_canon_sha256":"51ba2dfdf9f4f0b249e052bae31dfbf2a07d67af960f9f29358f8afe0746466b"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:54:01.061438Z","signature_b64":"I9Ai+6w3Tm2GIdAtb0oHPi/i4GcAkfB9q38Q1uRq4ytt53EFeaY0Oa3YMwGFyVJHOuTitEE/wPyuBN0xwG8CAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"afc688a8e5f6c75e533131a1b8e4e3a1549661de586e241ce589ec2255a08724","last_reissued_at":"2026-07-05T09:54:01.060976Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:54:01.060976Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"MaxMin-RLHF: Alignment with Diverse Human Preferences","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.LG","cs.RO"],"primary_cat":"cs.CL","authors_text":"Alec Koppel, Amrit Singh Bedi, Dinesh Manocha, Furong Huang, Hui Yuan, Jiahao Qiu, Mengdi Wang, Souradip Chakraborty","submitted_at":"2024-02-14T03:56:27Z","abstract_excerpt":"Reinforcement Learning from Human Feedback (RLHF) aligns language models to human preferences by employing a singular reward model derived from preference data. However, such an approach overlooks the rich diversity of human preferences inherent in data collected from multiple users. In this work, we first derive an impossibility result of alignment with single reward RLHF, thereby highlighting its insufficiency in representing diverse human preferences. To provide an equitable solution to the problem, we learn a mixture of preference distributions via an expectation-maximization algorithm and"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.08925","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.08925/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.08925","created_at":"2026-07-05T09:54:01.061031+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.08925v2","created_at":"2026-07-05T09:54:01.061031+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.08925","created_at":"2026-07-05T09:54:01.061031+00:00"},{"alias_kind":"pith_short_12","alias_value":"V7DIRKHF63DV","created_at":"2026-07-05T09:54:01.061031+00:00"},{"alias_kind":"pith_short_16","alias_value":"V7DIRKHF63DV4UZR","created_at":"2026-07-05T09:54:01.061031+00:00"},{"alias_kind":"pith_short_8","alias_value":"V7DIRKHF","created_at":"2026-07-05T09:54:01.061031+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.10569","citing_title":"Hidden Consensus:Preference-Validity Compression in Human Feedback","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30323","citing_title":"In-Context Reward Adaptation for Robust Preference Modeling","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09038","citing_title":"Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs","ref_index":170,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20408","citing_title":"Spectral Souping: A Unified Framework for Online Preference Alignment","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.13875","citing_title":"Common-agency Games for Multi-Objective Test-Time Alignment","ref_index":186,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12288","citing_title":"TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching","ref_index":145,"is_internal_anchor":false},{"citing_arxiv_id":"2603.27141","citing_title":"Routing Sensitivity Without Controllability: A Diagnostic Study of Fairness in MoE Language Models","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12288","citing_title":"TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching","ref_index":145,"is_internal_anchor":false},{"citing_arxiv_id":"2604.25895","citing_title":"Three Models of RLHF Annotation: Extension, Evidence, and Authority","ref_index":12,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF","json":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF.json","graph_json":"https://pith.science/api/pith-number/V7DIRKHF63DV4UZRGGQ3RZHDUF/graph.json","events_json":"https://pith.science/api/pith-number/V7DIRKHF63DV4UZRGGQ3RZHDUF/events.json","paper":"https://pith.science/paper/V7DIRKHF"},"agent_actions":{"view_html":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF","download_json":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF.json","view_paper":"https://pith.science/paper/V7DIRKHF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.08925&json=true","fetch_graph":"https://pith.science/api/pith-number/V7DIRKHF63DV4UZRGGQ3RZHDUF/graph.json","fetch_events":"https://pith.science/api/pith-number/V7DIRKHF63DV4UZRGGQ3RZHDUF/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF/action/timestamp_anchor","attest_storage":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF/action/storage_attestation","attest_author":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF/action/author_attestation","sign_citation":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF/action/citation_signature","submit_replication":"https://pith.science/pith/V7DIRKHF63DV4UZRGGQ3RZHDUF/action/replication_record"}},"created_at":"2026-07-05T09:54:01.061031+00:00","updated_at":"2026-07-05T09:54:01.061031+00:00"}