Pith. sign in

Paper Citation Record · LEDGER

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2604.16995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.16995 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T06:57:03.100519Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:20:30.920209Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:20:31.411452Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact26
  • verified fuzzy4
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ec81a05-1c77-4b03-9511-0a94af6757e6 · outbound

This paper cites online" 'onlinestring :=.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models online" 'onlinestring :=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:55:15.525804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:b100f80cfb183cebdb0d2b71d2aa6bcb5712bf3cd6d28ab7e52dadaa45c5464b

Observation 77497899-a400-4eed-af00-ff405fc54f83 · outbound

This paper cites write newline.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models write newline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:55:15.530524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:606d9eec771ad88cf27264d6cdbb7bba06cdbb5645eac172804797248f9611ef

Observation 0975fb6a-f936-43a8-a3fd-b0305a165f23 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.532366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:cf45c071bb61f36639db1c274b3f061367ba88bc3324c42f5a54509aa6908594

Observation d36710cd-2d66-45b6-858c-68a541f2b5d7 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.536500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:e3624297b956d1f0dffe8b2706525bf4348b8af72631bbb6bc81e3b9c6b7b1dc

Observation ffb3f9a7-bae6-4387-a39d-f03a74925aa4 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.470700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:7cdd5ec50a59232c788489f11afcdf45ec5076103c278dc3cd4d45eb0a95457e

Observation f1771139-640d-4c79-b493-af854a62b4c7 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:31:34.026635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:09db2ceeabfce562c6137ffd3393185f3e1ce827d2044b14f007ff4b61f06133

Observation 33ffb170-3324-48bc-a311-c6d1a2b3b52a · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.561608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:3b47bf94b710d914bafcf37c5aee178c704d5d141a808e7a1ef1f712518009dc

Observation 8a1f5723-2969-4d0b-af72-950b586f1e0c · outbound

This paper cites From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.475714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:629e1de8def2e90370aed3b5e42c83a2b4bdb3779a76c6f87b479703f9c3b5c6

Observation 566cc16d-5202-4ecb-8586-b45f49c7df01 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.563452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:fd4c8fb519e4c6d648beabaa32bb7c704bb7b9cb6d090c3184872308799f4cdc

Observation d0c42ba5-f993-4ce8-a57f-e8ff472c2653 · outbound

This paper cites Mathematical exploration and discovery at scale.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Mathematical exploration and discovery at scale

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.446323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:08914670babb684901cc46c273bb434761bbb4b996454133e4e028f50aace42b

Observation 15d98f44-dae8-4feb-9b38-332a03a75865 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:43:44.770948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:15cc2bbd71eac87e2a89e7d30bccad7f69e8724cdf04c0d0440de831ae54c10f

Observation 8241e895-015b-4158-968c-e2b2c8c22082 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.528293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:c0a34bbc9370d89b6de5427e424c658979ee8e3ff89063aa5141defef5f876ad

Observation 7e810b50-e00e-4419-a74e-2208df121415 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.534403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:1e61cb5d47222116c1028ba27c45eb4f7f21fdfb4167de1eae5fd817fd5cb8cb

Observation 828fdfcf-0e14-4045-a842-0079075aadc5 · outbound

This paper cites Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.441454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:eb268b672216180e397dc277ec210a44f173fabd11e18b2190974573c2b00dcc

Observation b9f4dcbc-5880-46cb-8691-fba541979f9c · outbound

This paper cites CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.478225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:806a7b8808cabe981b18686b642da2673124580d0d8cf9ac52d8b83b7f584953

Observation e7b109b5-68e7-43ee-b788-bbfdb906e21b · outbound

This paper cites Gonzalez, Haotong Zhang, and Ion Stoica.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Gonzalez, Haotong Zhang, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:55:15.569301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:3dabbb940c7fd7b5d75ce51042c91418e9f83b4b4155abfcf61e49136e74351f

Observation 084d5fc7-7dfe-47e3-8c35-eb9ec42f3423 · outbound

This paper cites S*: Test Time Scaling for Code Generation.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models S*: Test Time Scaling for Code Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.460726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:f3a01855675bc736af4e381111160b93f96b95fc8c9a13dd09eb143288a9742a

Observation f7513236-bdc9-4733-933c-919b15a33438 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.542991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:5200fe7343d952f13cd5dc8487a63af8e2ef3e8c7f1e3c8a42e82373e1542797

Observation 3bd4f56d-4dc1-463d-9567-a72a04affa89 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:52:33.970587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:0da04f004a02f946ef174ade667f2ca9b401f7709aaa3090096fb1a393d2a2bc

Observation 50f8c7e8-bcf7-4ffe-8582-1721a190e9a4 · outbound

This paper cites Liu, and Jialu Liu.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Liu, and Jialu Liu

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T14:55:15.555714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:7a0c907496511bae52f1c69cf039265fac1e60a390c8ea014ee52f6465c1e4f9

Observation 267c52bd-43f3-40d2-b00e-ee636677471a · outbound

This paper cites Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.455674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:ce62fb5bf3931a6a2d3d728176dcc1c328597b4641937b8f42ce3c90dfd7520a

Observation 7b07e0c8-3e02-4f7a-82c4-bd952e3ac418 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.574693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:8970924a8a848e5ebdfff7496f02faedfb6f36423d5c9b952a933519ec7f3100

Observation f35a81de-7c7a-4c00-8068-85bb814299a7 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.576846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:2a1dd8ecd7bb04ce59d2ec46e355a7be20955d275937a20383178484876a2a2f

Observation 4ec9e707-fc55-4dcd-a479-7aa308443d9f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:34.729655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:00411050fac3c8be4aa8173e48b2f2e6fc508dcc1f2b426dbec7e4f52216248d

Observation fc835128-3a56-42a2-86bb-8a2e16499593 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning Dynamics of LLM Finetuning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.448711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:a3421ed9d2b14c86c6954006f1d265e871196f602058546bd10e807065200396

Observation 57b873b7-101f-45e4-9ad2-c2024e6d23f1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Proximal Policy Optimization Algorithms

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.450983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:41996e911ca4b58aed9cc0c096545f9482adf926d57bd1bf7f117fb8674bef6f

Observation aa85f6be-7ebd-4018-b803-58e7402a18b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.453306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:83a9479f98e47431a1f62ba159fbdddcb05159620055edf0dc680a04c03a85b7

Observation e97a95ed-936f-4b32-b3c7-02c611786cd1 · outbound

This paper cites Learning to summarize from human feedback.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning to summarize from human feedback

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:46:18.838660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:08ac6172a2d539f10e6a8afdbbe2d0938e22f33fbc243acc851d5ec42c016905

Observation 4c8f68e1-facc-4105-b947-2b61dfe24128 · outbound

This paper cites Supervised Fine-Tuning as Inverse Reinforcement Learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Supervised Fine-Tuning as Inverse Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.443854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:7a42336392191c8e828e74045fbb9aede558f4b3cf4d3ac63f3246c3a58a9704

Observation 338164bb-6721-4a23-909a-40b141775f2c · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.567485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:2226118c57214d01302e786a7aa9551de07a940b592555dec4c377627cdd2a90

Observation 8296079b-4ebf-40a0-8935-3c49212b7bd7 · outbound

This paper cites Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Inverse-RLignment: Large Language Model Alignment from Demonstrations through Inverse Reinforcement Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.500518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:99bc2d3c383b123d7eae0d45e6718f51bfb728acf34158b9e31e286d5a137b4e

Observation aca2a5bf-65e9-4038-9db7-09d3e536cd12 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.540840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:0996ef5ed69f7e6db0bd1af8789f9a3e61df77d648f838efc36fa7c208a1022a

Observation c1e9508d-5c33-4d73-9f4e-e8843132a5be · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.559671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:5264a5c290f61f874b33be1d553846ba17a66955841ba3d5bb7715b0e2c0d7f0

Observation 83ffb24a-c66c-4536-969b-f4492e885316 · outbound

This paper cites Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.483237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:275214dbf48b2750fbeae2ee132bd88bad9f3e4b187d8dd4c3d1188036628ae1

Observation 2abed577-c436-4d51-afd6-15a186e90882 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.571089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:045b3f07bbae23e3ae5701a313cf1f736dba95b1c29ad648b49163a9668dad27

Observation 4d8f7b00-749d-4690-87b3-5fd3a02809e1 · outbound

This paper cites Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.485580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:ab3dc37de2d69b5c3879cb30523f59d2c5b44037f8dde3516fd3753410557e73

Observation c8af3911-7863-44ad-87a0-472bec56d515 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.545253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:cc4646190ce00afd95932cae2405c493c4188b2c19f08c870eba15fd3014d82a

Observation 987896e9-8a23-4fba-9a5d-b302312a374a · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.551479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:24dd805070ffb569012fda7bbef560f2cfe0c0cc135fe0a489adfee1f3a7c281

Observation a05add80-9231-4bbe-8eed-f59ba42e59cc · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.538844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:773b9736b0628695dc291a1e78bc2571fbeb9ed25ea3731ad3e388b12a71babb

Observation 964beff2-08cf-4683-916e-3735a1a8ba7c · outbound

This paper cites It Takes Two: Your GRPO Is Secretly DPO.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models It Takes Two: Your GRPO Is Secretly DPO

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:11.007781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:0121acdd398d77d66a10f3f7a5656663d4f2931b2d9ed741e210f9b6a7a7176a

Observation 4a79a9c6-e7c4-4681-a0d7-92eb0ac8bdeb · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:57:51.136140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:9447d6bfbed19afa05349dbb4a107a637074378b5f38023fcf651b706a8b30ed

Observation 64027dc7-37f5-4b37-9f66-ed52fc03ee4e · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Learning to Reason under Off-Policy Guidance

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:03.075876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:913849ff5fd1502efacd989cb6bdb1caee1b0529a67dab772f40c983ac6519c5

Observation 224994c8-493a-40c4-a9a2-cfe69a5746ac · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:15:55.036587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:8c7f25baebc3d3478c5f065c74a8b7d0d09b7ece0e33c872ac99598f961c03e7

Observation c8363d4e-eee7-401f-bc58-e1865fd4cee3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.495394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:889a48a0845387cda2c90965a86d72a40ea67126b1cd8b7f28d2bdfc7baa5e89

Observation bd8b9091-2d6c-4354-941c-51ef0de7d0bc · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.487917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:edaa61c0c573ac05e3db86e1abad31e55e12e66b85c3820e69b1d5a2ca231ac4

Observation 24495761-22ec-45de-8521-e71f187b863a · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.565324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:134ce29a19ab39c3391265b92029f4cb809543a043ec75fad6e6c31d6ae15e16

Observation dde4905d-5d91-48c2-8471-467b3586a07f · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.572846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:f6d8db763365ef33ee0120521c5240ee84aebed8301e04e0037768764cc90535

Observation e13566b5-4de1-4557-8884-d490fc1d1882 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.557760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:5d25cd2938dba956b8924ccb44134521d206b32dc932282a078bc9d9a901f917

Observation 844d362b-ca1a-40d3-85d7-cffa359a81c5 · outbound

This paper cites Group Sequence Policy Optimization.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Group Sequence Policy Optimization

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:22:54.227899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:2173a5052eb12b86db7954f411fc4ad64ec9c70688b424686fa9367075aaa67e

Observation d08ba6a7-c231-4cca-ad25-bb2cc357cf19 · outbound

This paper cites an unresolved cited work.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-21T14:55:15.553670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:b41872e69a682d9c5d4917fd4de700db52205281ea4e765478e21ea0c26c4c2d

Pith citing papers

Observation f1bdb373-645c-4c2d-b478-0a7e8fc05de0 · inbound

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning cites this paper.

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:20:31.418525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T00:20:30.920209Z digest=sha256:136d073e18ff17103f27f11182d098bc3d891ee10da6ef1423357e430069566d