{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:MWZG7PTELHSEQEHF6QRUENIQOD","short_pith_number":"pith:MWZG7PTE","schema_version":"1.0","canonical_sha256":"65b26fbe6459e44810e5f42342351070f2a5a147290a57714875b0bdb8eb3bb5","source":{"kind":"arxiv","id":"2505.18531","version":1},"attestation_state":"computed","paper":{"title":"Generative RLHF-V: Learning Principles from Multi-modal Human Preference","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.CV"],"primary_cat":"cs.AI","authors_text":"Boyuan Chen, Donghai Hong, Jiaming Ji, Jiapeng Sun, Jiayi Zhou, Sirui Han, Wenqi Chen, Yaodong Yang, Yike Guo","submitted_at":"2025-05-24T05:50:07Z","abstract_excerpt":"Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, e.g., reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that in"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2505.18531","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.AI","submitted_at":"2025-05-24T05:50:07Z","cross_cats_sorted":["cs.CV"],"title_canon_sha256":"d9f5d6412a51f02a9cd11602d34b6d6dffa1b5ba6ae229a451cb410ea3324ba7","abstract_canon_sha256":"94e6c41ac7eda42f88626992cd8c08f8ccc108304e6f477e9366f88ae4a4b12f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:08:47.813949Z","signature_b64":"y+IZWkr1BJxvNIxL6jrMONx9ZwEYDVlD5QM+GRBl6GlqpJ5HGbHXXimD+Z4PAitHNkOjXmSFKzuGOzo0jcUMAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"65b26fbe6459e44810e5f42342351070f2a5a147290a57714875b0bdb8eb3bb5","last_reissued_at":"2026-07-05T11:08:47.813445Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:08:47.813445Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Generative RLHF-V: Learning Principles from Multi-modal Human Preference","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.CV"],"primary_cat":"cs.AI","authors_text":"Boyuan Chen, Donghai Hong, Jiaming Ji, Jiapeng Sun, Jiayi Zhou, Sirui Han, Wenqi Chen, Yaodong Yang, Yike Guo","submitted_at":"2025-05-24T05:50:07Z","abstract_excerpt":"Training multi-modal large language models (MLLMs) that align with human intentions is a long-term challenge. Traditional score-only reward models for alignment suffer from low accuracy, weak generalization, and poor interpretability, blocking the progress of alignment methods, e.g., reinforcement learning from human feedback (RLHF). Generative reward models (GRMs) leverage MLLMs' intrinsic reasoning capabilities to discriminate pair-wise responses, but their pair-wise paradigm makes it hard to generalize to learnable rewards. We introduce Generative RLHF-V, a novel alignment framework that in"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2505.18531","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2505.18531/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2505.18531","created_at":"2026-07-05T11:08:47.813505+00:00"},{"alias_kind":"arxiv_version","alias_value":"2505.18531v1","created_at":"2026-07-05T11:08:47.813505+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2505.18531","created_at":"2026-07-05T11:08:47.813505+00:00"},{"alias_kind":"pith_short_12","alias_value":"MWZG7PTELHSE","created_at":"2026-07-05T11:08:47.813505+00:00"},{"alias_kind":"pith_short_16","alias_value":"MWZG7PTELHSEQEHF","created_at":"2026-07-05T11:08:47.813505+00:00"},{"alias_kind":"pith_short_8","alias_value":"MWZG7PTE","created_at":"2026-07-05T11:08:47.813505+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.09711","citing_title":"Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization","ref_index":277,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07872","citing_title":"VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05660","citing_title":"Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation","ref_index":290,"is_internal_anchor":false},{"citing_arxiv_id":"2503.03480","citing_title":"SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19544","citing_title":"DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling","ref_index":65,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13602","citing_title":"Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges","ref_index":184,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD","json":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD.json","graph_json":"https://pith.science/api/pith-number/MWZG7PTELHSEQEHF6QRUENIQOD/graph.json","events_json":"https://pith.science/api/pith-number/MWZG7PTELHSEQEHF6QRUENIQOD/events.json","paper":"https://pith.science/paper/MWZG7PTE"},"agent_actions":{"view_html":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD","download_json":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD.json","view_paper":"https://pith.science/paper/MWZG7PTE","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2505.18531&json=true","fetch_graph":"https://pith.science/api/pith-number/MWZG7PTELHSEQEHF6QRUENIQOD/graph.json","fetch_events":"https://pith.science/api/pith-number/MWZG7PTELHSEQEHF6QRUENIQOD/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD/action/storage_attestation","attest_author":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD/action/author_attestation","sign_citation":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD/action/citation_signature","submit_replication":"https://pith.science/pith/MWZG7PTELHSEQEHF6QRUENIQOD/action/replication_record"}},"created_at":"2026-07-05T11:08:47.813505+00:00","updated_at":"2026-07-05T11:08:47.813505+00:00"}