Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:57:03.100519Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2604.16995.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T06:57:03.100519Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:30.920209Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T00:20:31.411452Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3ec81a05-1c77-4b03-9511-0a94af6757e6 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 77497899-a400-4eed-af00-ff405fc54f83 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models write newline
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0975fb6a-f936-43a8-a3fd-b0305a165f23 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d36710cd-2d66-45b6-858c-68a541f2b5d7 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ffb3f9a7-bae6-4387-a39d-f03a74925aa4 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1771139-640d-4c79-b493-af854a62b4c7 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 33ffb170-3324-48bc-a311-c6d1a2b3b52a · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a1f5723-2969-4d0b-af72-950b586f1e0c · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 566cc16d-5202-4ecb-8586-b45f49c7df01 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0c42ba5-f993-4ce8-a57f-e8ff472c2653 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Mathematical exploration and discovery at scale
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 15d98f44-dae8-4feb-9b38-332a03a75865 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8241e895-015b-4158-968c-e2b2c8c22082 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7e810b50-e00e-4419-a74e-2208df121415 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 828fdfcf-0e14-4045-a842-0079075aadc5 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b9f4dcbc-5880-46cb-8691-fba541979f9c · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7b109b5-68e7-43ee-b788-bbfdb906e21b · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Gonzalez, Haotong Zhang, and Ion Stoica
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 084d5fc7-7dfe-47e3-8c35-eb9ec42f3423 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models S*: Test Time Scaling for Code Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7513236-bdc9-4733-933c-919b15a33438 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3bd4f56d-4dc1-463d-9567-a72a04affa89 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50f8c7e8-bcf7-4ffe-8582-1721a190e9a4 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Liu, and Jialu Liu
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 267c52bd-43f3-40d2-b00e-ee636677471a · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7b07e0c8-3e02-4f7a-82c4-bd952e3ac418 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f35a81de-7c7a-4c00-8068-85bb814299a7 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ec9e707-fc55-4dcd-a479-7aa308443d9f · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fc835128-3a56-42a2-86bb-8a2e16499593 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning Dynamics of LLM Finetuning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57b873b7-101f-45e4-9ad2-c2024e6d23f1 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa85f6be-7ebd-4018-b803-58e7402a18b6 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e97a95ed-936f-4b32-b3c7-02c611786cd1 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning to summarize from human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c8f68e1-facc-4105-b947-2b61dfe24128 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Supervised Fine-Tuning as Inverse Reinforcement Learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 338164bb-6721-4a23-909a-40b141775f2c · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8296079b-4ebf-40a0-8935-3c49212b7bd7 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aca2a5bf-65e9-4038-9db7-09d3e536cd12 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c1e9508d-5c33-4d73-9f4e-e8843132a5be · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83ffb24a-c66c-4536-969b-f4492e885316 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2abed577-c436-4d51-afd6-15a186e90882 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4d8f7b00-749d-4690-87b3-5fd3a02809e1 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8af3911-7863-44ad-87a0-472bec56d515 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 987896e9-8a23-4fba-9a5d-b302312a374a · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a05add80-9231-4bbe-8eed-f59ba42e59cc · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 964beff2-08cf-4683-916e-3735a1a8ba7c · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models It Takes Two: Your GRPO Is Secretly DPO
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4a79a9c6-e7c4-4681-a0d7-92eb0ac8bdeb · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64027dc7-37f5-4b37-9f66-ed52fc03ee4e · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning to Reason under Off-Policy Guidance
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 224994c8-493a-40c4-a9a2-cfe69a5746ac · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8363d4e-eee7-401f-bc58-e1865fd4cee3 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd8b9091-2d6c-4354-941c-51ef0de7d0bc · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24495761-22ec-45de-8521-e71f187b863a · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dde4905d-5d91-48c2-8471-467b3586a07f · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e13566b5-4de1-4557-8884-d490fc1d1882 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 844d362b-ca1a-40d3-85d7-cffa359a81c5 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Group Sequence Policy Optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d08ba6a7-c231-4cca-ad25-bb2cc357cf19 · outbound
SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f1bdb373-645c-4c2d-b478-0a7e8fc05de0 · inbound
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.