{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:GZC6EAZPDRDSCBQFTTMDI4WVCO","short_pith_number":"pith:GZC6EAZP","schema_version":"1.0","canonical_sha256":"3645e2032f1c472106059cd83472d513b39c695f4f75a2ca9cd2c1b6294c7c30","source":{"kind":"arxiv","id":"2507.06892","version":3},"attestation_state":"computed","paper":{"title":"Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Hongyao Tang, Jianye Hao, Jing Liang, Jinyi Liu, Lei Bai, Shuyue Hu, Yan Zheng, Yi Ma","submitted_at":"2025-07-09T14:29:45Z","abstract_excerpt":"Reinforcement Learning (RL) has demonstrated its potential to improve the reasoning ability of Large Language Models (LLMs). One major limitation of most existing Reinforcement Finetuning (RFT) methods is that they are on-policy RL in nature, i.e., data generated during the past learning process is not fully utilized. This inevitably comes at a significant cost of compute and time, posing a stringent bottleneck on continuing economic and efficient scaling. To this end, we launch the renaissance of off-policy RL and propose Reincarnating Mix-policy Proximal Policy Gradient (ReMix), a general ap"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.06892","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","primary_cat":"cs.LG","submitted_at":"2025-07-09T14:29:45Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"709bcd98843228718f5f149a7677eaf705385f62e21ed9b7353ad94f2e82f70e","abstract_canon_sha256":"665761fadaaaf49c690dcf290b9fc12b2b11cbf486d34dcf88d6ff9fd3f6b9a6"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:35:36.380187Z","signature_b64":"PT5vF+6CMAvFzsSujtWltQ/oBg3PcMHOa1nNZN1lwlIikzjs1euFf0f3MM+uC5wFdgV4swgPa1xx3TmW5eCzCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"3645e2032f1c472106059cd83472d513b39c695f4f75a2ca9cd2c1b6294c7c30","last_reissued_at":"2026-07-05T11:35:36.379677Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:35:36.379677Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Hongyao Tang, Jianye Hao, Jing Liang, Jinyi Liu, Lei Bai, Shuyue Hu, Yan Zheng, Yi Ma","submitted_at":"2025-07-09T14:29:45Z","abstract_excerpt":"Reinforcement Learning (RL) has demonstrated its potential to improve the reasoning ability of Large Language Models (LLMs). One major limitation of most existing Reinforcement Finetuning (RFT) methods is that they are on-policy RL in nature, i.e., data generated during the past learning process is not fully utilized. This inevitably comes at a significant cost of compute and time, posing a stringent bottleneck on continuing economic and efficient scaling. To this end, we launch the renaissance of off-policy RL and propose Reincarnating Mix-policy Proximal Policy Gradient (ReMix), a general ap"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.06892","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.06892/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.06892","created_at":"2026-07-05T11:35:36.379732+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.06892v3","created_at":"2026-07-05T11:35:36.379732+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.06892","created_at":"2026-07-05T11:35:36.379732+00:00"},{"alias_kind":"pith_short_12","alias_value":"GZC6EAZPDRDS","created_at":"2026-07-05T11:35:36.379732+00:00"},{"alias_kind":"pith_short_16","alias_value":"GZC6EAZPDRDSCBQF","created_at":"2026-07-05T11:35:36.379732+00:00"},{"alias_kind":"pith_short_8","alias_value":"GZC6EAZP","created_at":"2026-07-05T11:35:36.379732+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.21943","citing_title":"Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning","ref_index":112,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01281","citing_title":"RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2509.08827","citing_title":"A Survey of Reinforcement Learning for Large Reasoning Models","ref_index":298,"is_internal_anchor":false},{"citing_arxiv_id":"2604.04142","citing_title":"OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12004","citing_title":"Learning Agentic Policy from Action Guidance","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18530","citing_title":"OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14142","citing_title":"From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02913","citing_title":"Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning","ref_index":73,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO","json":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO.json","graph_json":"https://pith.science/api/pith-number/GZC6EAZPDRDSCBQFTTMDI4WVCO/graph.json","events_json":"https://pith.science/api/pith-number/GZC6EAZPDRDSCBQFTTMDI4WVCO/events.json","paper":"https://pith.science/paper/GZC6EAZP"},"agent_actions":{"view_html":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO","download_json":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO.json","view_paper":"https://pith.science/paper/GZC6EAZP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.06892&json=true","fetch_graph":"https://pith.science/api/pith-number/GZC6EAZPDRDSCBQFTTMDI4WVCO/graph.json","fetch_events":"https://pith.science/api/pith-number/GZC6EAZPDRDSCBQFTTMDI4WVCO/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO/action/timestamp_anchor","attest_storage":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO/action/storage_attestation","attest_author":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO/action/author_attestation","sign_citation":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO/action/citation_signature","submit_replication":"https://pith.science/pith/GZC6EAZPDRDSCBQFTTMDI4WVCO/action/replication_record"}},"created_at":"2026-07-05T11:35:36.379732+00:00","updated_at":"2026-07-05T11:35:36.379732+00:00"}