Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:38.738111Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 13 inbound Pith citation observations for arXiv:2507.06892.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:57:38.738111Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:39:03.382391Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:09:40.703268Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2610c449-92c3-4237-b68f-837cb95d3efd · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Courville, and Marc G
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c77a04f0-b335-4aca-835e-cf99a9b82da2 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb985f7-c2bf-461d-b3b7-ddd22a85de4a · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Back to basics: Revisiting reinforce-style optimization for learning from human feedback in llms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80e48911-6d5a-4f96-b991-58bd963bb06b · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Polaris: A post-training recipe for scaling reinforcement learning on advanced reasoning models, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad64ff2-e8d6-4a4e-8250-b65159d4fdfd · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Asymmetric reinforce for off-policy reinforcement learning: Balancing positive and negative rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0eee0b-866b-48b0-bc3d-b6799b1b46d5 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c988a462-281d-4c88-acf0-f800fc8c8105 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd95391-ebaf-41cc-ac93-6aefd7da941e · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Randomized ensembled double q-learning: Learning fast without a model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 162c4416-8a67-4ce8-9d07-a7313e31f590 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea8bd9b-b671-4073-ae7e-c9bc97ee0670 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Soft Policy Optimization: Online Off-Policy RL for Sequence Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70a7c6b4-3295-4808-be17-964e2c5f0e00 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Reinforcement learning for reasoning in small llms: What works and what doesn't
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393233f6-6531-4951-acab-e6f7ab0deac5 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16e72a72-22cd-4a63-95f9-04c69ebf38df · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Concise reasoning via reinforcement learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5e93d1-783f-4001-ad8c-aa6e1b1e5aa8 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Fujimoto, H
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a900bc4e-07ba-4cc7-83f7-964739ef6ac6 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Omni-math: A universal olympiad level mathematic benchmark for large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70b85926-1f13-48d4-84d2-16cf5324ca0b · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9c56bf-41f2-45f6-b4da-2a02afdfffcb · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12e65803-a1fd-4ef8-a8de-6ab262a11134 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model O lympiad B ench: A challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eac0ca50-77da-4c6d-b617-3f70d76a2fe7 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Skywork Open Reasoner 1 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b5bf4b-de32-4e61-97f0-9d47fd639b78 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Measuring mathematical problem solving with the MATH dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3328c8a4-0c39-45d6-89c0-85eb10fc3582 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Rainbow: Combining improvements in deep reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a378b549-95f8-4ce0-962c-75f61d17b505 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Dropout Q-Functions for Doubly Efficient Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cfd7d7e-9d71-4916-8f80-8d1013ceb891 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Ii-thought
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 683bc374-3d9e-40c1-aa52-0171f9784ca8 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model OpenAI o1 System Card
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354f7b51-4215-4f86-b33f-d8765cb4fac0 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Towards Mitigating Hallucination in Large Language Models via Self-Reflection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381f4a2a-5caf-4152-a782-427d95387ff5 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Kakade and John Langford
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd2317f9-283c-4d90-addd-5786421cb595 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9e9e1f-94cb-4ccc-8aab-1cf469a1c64d · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Solving quantitative reasoning problems with language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a63b81c-201f-4b0c-8e2b-08465d99b10d · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model RePO: Replay-Enhanced Policy Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef5b0fc-e36b-4d2a-89a9-c32b6b72ce9c · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77054f07-b2a7-4362-8750-e8324c27da23 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Enhancing Robotic Manipulation with AI Feedback from Multimodal Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19eb0e03-baf4-4282-a5d3-d0a919c91fd5 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model From Chaos to Order: The Atomic Reasoner Framework for Fine-grained Reasoning in Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6be53a0b-78aa-4601-aa1a-3b61d86d2755 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf3a4479-b0c9-4279-bdc4-58d442b206d5 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Reining generalization in offline reinforcement learning via representation distinction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3aa58f0-4432-4700-8bd8-5d8c5fc0f2eb · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Iteratively refined behavior regularization for offline reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ade3e9f-5330-4261-b1b4-e77e1e4854aa · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9a26ac5-95fb-4a62-b7a7-1b638571f434 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8d6d070-020a-4d9e-a63b-1ae8a827ee77 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Cassandras
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b7553c1-3787-4eb4-bf64-4a81543fe049 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Speq: Offline stabilization phases for efficient q-learning in high update-to-data ratio reinforcement learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24c1ddbc-e70b-4886-9465-ec048c185c8d · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d11505c-4543-4ea3-adc5-2d9f491d471d · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Jordan, and Philipp Moritz
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22f9bee0-732a-4bae-ae6d-42f6cb655df2 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Jordan, and Pieter Abbeel
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dcffece2-cacf-4009-9d27-8f172a16c27c · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Proximal Policy Optimization Algorithms
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d95310c-4341-4732-a79a-c1e3bcae33cd · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5439f3bd-5b71-454b-9616-ee3804e18e7c · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08bf06f9-9683-4cc3-b47e-4cd2db8693cb · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 451cb349-4f0e-4ef1-9500-466989e44479 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd9063f1-88c2-4cca-962d-8342577d4b07 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Sutton and Andrew G
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 013e6b96-1d08-4581-a879-967f60f4dbef · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model RL-finetuning LLMs from on- and off-policy data with a single algorithm
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a51081-e03b-4d3c-891b-ecf941f8de30 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Reft: Reasoning with reinforced fine-tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a646dfcc-7460-4a39-807b-28dd653aa209 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d7c104-2b2d-4e9b-a5a0-310b1a3e7b01 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Truly proximal policy optimization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37f24888-f153-448b-8e6f-c7327f052b92 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Chain-of-thought prompting elicits reasoning in large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4d0928f-6f77-40d6-9eaf-e34070df6e62 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9d34cd-b56e-4fc3-ba0e-28c5d0378d74 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Learning to Reason under Off-Policy Guidance
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f9f242-705e-4d1b-8ae8-cc19763f253b · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Qwen3 Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bd9843a-1ad8-4571-9604-2100e99d3afc · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea91d13-5412-42af-b37b-2f39d9ddb1de · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Star: Self-taught reasoner bootstrapping reasoning with reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67976d1d-e23a-47ed-b422-8588f9368580 · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model Rest-mcts*: Llm self-training via process reward guided tree search
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5bd67773-33ef-40e1-8585-8922dcb3f4fd · outbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AdaptThink: Reasoning Models Can Learn When to Think
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a88053-8484-4f20-8e2e-d2afacf86fab · inbound
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55dcc61-54c0-42e7-b913-376652390503 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 298
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70c6515c-dc0f-448b-9b73-832627027942 · inbound
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 387033bb-6cb1-4f23-9b62-1a342820f39f · inbound
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b66fa036-e05a-4695-81bc-fef0a840ab7d · inbound
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b03ad2d-f6af-45dc-92fa-3ad994e68e0a · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92c5b202-9bfe-4c5c-b64a-08617c072a60 · inbound
Learning Agentic Policy from Action Guidance Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d775364-fee8-466c-b897-6c43696ec376 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8710f2d7-6436-4e9b-b31c-6062714716a8 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 894710be-2893-4b06-bcb2-48495fcf0ba9 · inbound
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4830619b-67bb-4586-920e-a111b79b15f6 · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ba5694-badf-4fad-8793-20ff8e40f96b · inbound
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a5b6de-1ab0-41b3-90ec-3cb6462a186e · inbound
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.