{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:3QOXZ4BQBXZEZSFX25YRBQ7KKR","short_pith_number":"pith:3QOXZ4BQ","schema_version":"1.0","canonical_sha256":"dc1d7cf0300df24cc8b7d77110c3ea547af089f65b0a5b9c677693a7a59262b0","source":{"kind":"arxiv","id":"2405.19107","version":1},"attestation_state":"computed","paper":{"title":"Offline Regularised Reinforcement Learning for Large Language Models Alignment","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Aliaksei Severyn, Bernardo Avila Pires, Bilal Piot, Daniele Calandriello, Daniel Guo, Eugene Tarassov, Gil Shamir, Jonathan Mallinson, Lior Shani, Lucas Spangher, Mohammad Gheshlaghi Azar, Pierre Harvey Richemond, Rafael Rafailov, Remi Munos, Rishabh Joshi, Tianqi Liu, Will Ellsworth, Yunhao Tang","submitted_at":"2024-05-29T14:11:29Z","abstract_excerpt":"The dominant framework for alignment of large language models (LLM), whether through reinforcement learning from human feedback or direct preference optimisation, is to learn from preference data. This involves building datasets where each element is a quadruplet composed of a prompt, two independent responses (completions of the prompt) and a human preference between the two independent responses, yielding a preferred and a dis-preferred response. Such data is typically scarce and expensive to collect. On the other hand, \\emph{single-trajectory} datasets where each element is a triplet compos"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2405.19107","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2024-05-29T14:11:29Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a6331ee315f03d7d4b7c88e28046509d55bd2c95b0aec8245c064ba245fd3681","abstract_canon_sha256":"004025d79f6453aaab78db2e6e36fe5bacf1cf4a43185844e3b4539dd4c28853"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:24:47.994557Z","signature_b64":"9Wre/ZIkHjIK51Vc/jHeYspMcuunHNj3AIDNe1Qglhq6d1w4sbeYl0LpMmAJ19rc6Ta+Cj2YIo78AJGu5m5ZAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"dc1d7cf0300df24cc8b7d77110c3ea547af089f65b0a5b9c677693a7a59262b0","last_reissued_at":"2026-07-05T08:24:47.994059Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:24:47.994059Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Offline Regularised Reinforcement Learning for Large Language Models Alignment","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Aliaksei Severyn, Bernardo Avila Pires, Bilal Piot, Daniele Calandriello, Daniel Guo, Eugene Tarassov, Gil Shamir, Jonathan Mallinson, Lior Shani, Lucas Spangher, Mohammad Gheshlaghi Azar, Pierre Harvey Richemond, Rafael Rafailov, Remi Munos, Rishabh Joshi, Tianqi Liu, Will Ellsworth, Yunhao Tang","submitted_at":"2024-05-29T14:11:29Z","abstract_excerpt":"The dominant framework for alignment of large language models (LLM), whether through reinforcement learning from human feedback or direct preference optimisation, is to learn from preference data. This involves building datasets where each element is a quadruplet composed of a prompt, two independent responses (completions of the prompt) and a human preference between the two independent responses, yielding a preferred and a dis-preferred response. Such data is typically scarce and expensive to collect. On the other hand, \\emph{single-trajectory} datasets where each element is a triplet compos"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2405.19107","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2405.19107/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2405.19107","created_at":"2026-07-05T08:24:47.994120+00:00"},{"alias_kind":"arxiv_version","alias_value":"2405.19107v1","created_at":"2026-07-05T08:24:47.994120+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2405.19107","created_at":"2026-07-05T08:24:47.994120+00:00"},{"alias_kind":"pith_short_12","alias_value":"3QOXZ4BQBXZE","created_at":"2026-07-05T08:24:47.994120+00:00"},{"alias_kind":"pith_short_16","alias_value":"3QOXZ4BQBXZEZSFX","created_at":"2026-07-05T08:24:47.994120+00:00"},{"alias_kind":"pith_short_8","alias_value":"3QOXZ4BQ","created_at":"2026-07-05T08:24:47.994120+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.21943","citing_title":"Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning","ref_index":160,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01249","citing_title":"Trust Region On-Policy Distillation","ref_index":204,"is_internal_anchor":false},{"citing_arxiv_id":"2408.15339","citing_title":"UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2602.07832","citing_title":"rePIRL: Learn PRM with Inverse RL for LLM Reasoning","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2512.15605","citing_title":"Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09214","citing_title":"Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability","ref_index":45,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR","json":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR.json","graph_json":"https://pith.science/api/pith-number/3QOXZ4BQBXZEZSFX25YRBQ7KKR/graph.json","events_json":"https://pith.science/api/pith-number/3QOXZ4BQBXZEZSFX25YRBQ7KKR/events.json","paper":"https://pith.science/paper/3QOXZ4BQ"},"agent_actions":{"view_html":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR","download_json":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR.json","view_paper":"https://pith.science/paper/3QOXZ4BQ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2405.19107&json=true","fetch_graph":"https://pith.science/api/pith-number/3QOXZ4BQBXZEZSFX25YRBQ7KKR/graph.json","fetch_events":"https://pith.science/api/pith-number/3QOXZ4BQBXZEZSFX25YRBQ7KKR/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR/action/timestamp_anchor","attest_storage":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR/action/storage_attestation","attest_author":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR/action/author_attestation","sign_citation":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR/action/citation_signature","submit_replication":"https://pith.science/pith/3QOXZ4BQBXZEZSFX25YRBQ7KKR/action/replication_record"}},"created_at":"2026-07-05T08:24:47.994120+00:00","updated_at":"2026-07-05T08:24:47.994120+00:00"}