{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:DUSVRFFVRMO6IS5EV5XBU2LRRS","short_pith_number":"pith:DUSVRFFV","schema_version":"1.0","canonical_sha256":"1d255894b58b1de44ba4af6e1a69718caeeb2dab144e477bd1b2bce57a3562df","source":{"kind":"arxiv","id":"2402.04792","version":2},"attestation_state":"computed","paper":{"title":"Direct Language Model Alignment from Online AI Feedback","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.CL","cs.HC"],"primary_cat":"cs.AI","authors_text":"Alexandre Rame, Biao Zhang, Bilal Piot, Felipe Llinares, Johan Ferret, Mathieu Blondel, Misha Khalman, Shangmin Guo, Thomas Mesnard, Tianlin Liu, Tianqi Liu, Yao Zhao","submitted_at":"2024-02-07T12:31:13Z","abstract_excerpt":"Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not require a separate reward model. However, the preference datasets used in DAP methods are usually collected ahead of training and never updated, thus the feedback is purely offline. Moreover, responses in these datasets are often sampled from a language model distinct from the one being aligned, and since the model evolves over training, the alignment phase is inevitably off-policy. In this study, we posit that online fe"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.04792","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.AI","submitted_at":"2024-02-07T12:31:13Z","cross_cats_sorted":["cs.CL","cs.HC"],"title_canon_sha256":"beca7be5dbbffd9c57e6fcbd8067692fe940f30a6238ffd8556679004e0563b6","abstract_canon_sha256":"b99e82ca4c61ab67a0ad86437f11f13455387d9e06068c35d8249e432f29c3f0"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:50:55.375766Z","signature_b64":"WhDkHrDwGoKikeo8wCqfCorzhzNSaLLq3GKopVICc+sFK6CTyABij+iihyweNcqii/W9wLRzHCWQt+SRbGdVAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"1d255894b58b1de44ba4af6e1a69718caeeb2dab144e477bd1b2bce57a3562df","last_reissued_at":"2026-07-05T07:50:55.375196Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:50:55.375196Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Direct Language Model Alignment from Online AI Feedback","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.CL","cs.HC"],"primary_cat":"cs.AI","authors_text":"Alexandre Rame, Biao Zhang, Bilal Piot, Felipe Llinares, Johan Ferret, Mathieu Blondel, Misha Khalman, Shangmin Guo, Thomas Mesnard, Tianlin Liu, Tianqi Liu, Yao Zhao","submitted_at":"2024-02-07T12:31:13Z","abstract_excerpt":"Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not require a separate reward model. However, the preference datasets used in DAP methods are usually collected ahead of training and never updated, thus the feedback is purely offline. Moreover, responses in these datasets are often sampled from a language model distinct from the one being aligned, and since the model evolves over training, the alignment phase is inevitably off-policy. In this study, we posit that online fe"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.04792","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.04792/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.04792","created_at":"2026-07-05T07:50:55.375253+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.04792v2","created_at":"2026-07-05T07:50:55.375253+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.04792","created_at":"2026-07-05T07:50:55.375253+00:00"},{"alias_kind":"pith_short_12","alias_value":"DUSVRFFVRMO6","created_at":"2026-07-05T07:50:55.375253+00:00"},{"alias_kind":"pith_short_16","alias_value":"DUSVRFFVRMO6IS5E","created_at":"2026-07-05T07:50:55.375253+00:00"},{"alias_kind":"pith_short_8","alias_value":"DUSVRFFV","created_at":"2026-07-05T07:50:55.375253+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":26,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.24937","citing_title":"The Hitchhiker's Guide to Agentic AI: From Foundations to Systems","ref_index":209,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09711","citing_title":"Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization","ref_index":219,"is_internal_anchor":false},{"citing_arxiv_id":"2606.04807","citing_title":"BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01249","citing_title":"Trust Region On-Policy Distillation","ref_index":203,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04477","citing_title":"Data-dependent Exploration for Online Reinforcement Learning from Human Feedback","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2605.26293","citing_title":"CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.23398","citing_title":"TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2504.12501","citing_title":"Reinforcement Learning from Human Feedback","ref_index":203,"is_internal_anchor":false},{"citing_arxiv_id":"2602.06239","citing_title":"Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15604","citing_title":"VSPO: Vector-Steered Policy Optimization for Behavioral Control","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2507.06419","citing_title":"Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2511.01014","citing_title":"IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2512.16378","citing_title":"Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2602.06239","citing_title":"Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15012","citing_title":"Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2601.08584","citing_title":"Ministral 3","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2604.02766","citing_title":"Random Is Hard to Beat: Active Selection in online DPO with Modern LLMs","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27733","citing_title":"Mind the Gap: Structure-Aware Consistency in Preference Learning","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09291","citing_title":"dFlowGRPO: Rate-Aware Policy Optimization for Discrete Flow Models","ref_index":142,"is_internal_anchor":false},{"citing_arxiv_id":"2605.08283","citing_title":"HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2412.05579","citing_title":"LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods","ref_index":76,"is_internal_anchor":false},{"citing_arxiv_id":"2604.23809","citing_title":"LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.04477","citing_title":"Data-dependent Exploration for Online Reinforcement Learning from Human Feedback","ref_index":69,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13602","citing_title":"Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges","ref_index":164,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15847","citing_title":"CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization","ref_index":12,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS","json":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS.json","graph_json":"https://pith.science/api/pith-number/DUSVRFFVRMO6IS5EV5XBU2LRRS/graph.json","events_json":"https://pith.science/api/pith-number/DUSVRFFVRMO6IS5EV5XBU2LRRS/events.json","paper":"https://pith.science/paper/DUSVRFFV"},"agent_actions":{"view_html":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS","download_json":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS.json","view_paper":"https://pith.science/paper/DUSVRFFV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.04792&json=true","fetch_graph":"https://pith.science/api/pith-number/DUSVRFFVRMO6IS5EV5XBU2LRRS/graph.json","fetch_events":"https://pith.science/api/pith-number/DUSVRFFVRMO6IS5EV5XBU2LRRS/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS/action/timestamp_anchor","attest_storage":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS/action/storage_attestation","attest_author":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS/action/author_attestation","sign_citation":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS/action/citation_signature","submit_replication":"https://pith.science/pith/DUSVRFFVRMO6IS5EV5XBU2LRRS/action/replication_record"}},"created_at":"2026-07-05T07:50:55.375253+00:00","updated_at":"2026-07-05T07:50:55.375253+00:00"}