Pith. sign in

Paper Citation Record · LEDGER

Spurious Rewards: Rethinking Training Signals in RLVR

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 69 inbound Pith citation observations for arXiv:2506.10947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10947 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T13:37:51.086217Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:05.462042Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact3
  • verified fuzzy1
  • unresolved10
  • parse uncertain2
  • malformed identifier2
  • metadata mismatch3

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4a385ed9-c2c4-4bf4-9ed3-c95874899ef5 · outbound

This paper cites doi: 10.1038/s41586-025-09422-z.

Spurious Rewards: Rethinking Training Signals in RLVR doi: 10.1038/s41586-025-09422-z

Reference 1

Resolution
metadata mismatch
doi, observed 2026-05-16T13:37:51.109158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:58ca5ce07dcb6e054ae8b9aa5fb067d4a20d3341cf82c81e1a0234659d97678a

Observation 2dea4de5-be44-48d8-b64c-93c47805eaa6 · outbound

This paper cites 2 OLMo 2 Furious.

Spurious Rewards: Rethinking Training Signals in RLVR 2 OLMo 2 Furious

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:37:51.117231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:633395f4d1804d1411f46eb9346a743b785cc043ca5b432233f9e04ae39b1770

Observation 2a5e8549-3eef-4533-a2e1-d79e8df51518 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

Spurious Rewards: Rethinking Training Signals in RLVR Maximizing Confidence Alone Improves Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.121149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:dbc9ba304bf96326bd30ce4240b412097b07033948cf28279750a4ac1b0da157

Observation c9d64740-703c-4c47-b418-e60d75ab1a8f · outbound

This paper cites ISBN 979-8-89176-288-6.

Spurious Rewards: Rethinking Training Signals in RLVR ISBN 979-8-89176-288-6

Reference 4

Resolution
verified exact
doi, observed 2026-05-16T13:37:51.104488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:843078ab7d32b2f185c707b7f7222352f9ac233ea77d6dd10bb8a6abebb20c2d

Observation a1f7f2c2-4234-4f6a-ad6b-3b505e044238 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Spurious Rewards: Rethinking Training Signals in RLVR Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:57:51.136140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:777c5b926bfc7f1953f78d1bad3dbedc4bb57e4c2ef489b34ffa8e695ea28eac

Observation 59b052cd-f9ff-4880-9dd4-617d1ce15ab0 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Spurious Rewards: Rethinking Training Signals in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T13:37:51.113502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:4569d107d37a4c60427f5ec32f9c992c5f140baeb53922f209ae00161c3ad7ab

Observation bde87d3a-33ec-4a7e-95c3-0ba6794ae34c · outbound

This paper cites clipping bias.

Spurious Rewards: Rethinking Training Signals in RLVR clipping bias

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T13:37:51.151752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:8522fdf0f198b4cb2cbc680cbe481c3bcded177575e3fc071c626b2357bb54bc

Observation 973f233d-76be-4905-85e2-892b00cd5d91 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.153992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:3e3b05aee2b9c006076ff0e4cdf9ec93a4210c46d04352c10fa4c4adf8ca3340

Observation 600501ef-5d6f-4b5a-8405-621481040869 · outbound

This paper cites r={r},θ= {theta}.

Spurious Rewards: Rethinking Training Signals in RLVR r={r},θ= {theta}

Reference 9

Resolution
malformed identifier
raw_fallback, observed 2026-05-16T13:37:51.156186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:5cf08e4d9cab5a97eb82cb797c60790132538b8cc0023062b7e5169e14047d26

Observation 4a01737a-1ad0-401c-ab0d-6290fa19e4f9 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.158345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:25347a9b81e54b8411e2aae18a51d0fa8ad015c40f45b61d00a01c25134857b0

Observation 41602f8a-22e4-4d04-b61a-fbd9a478714f · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.160343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:40437bcaeb98132b49f127298d12661778e407ed24ecefe459d4c326530bda3d

Observation a629e801-ab4c-410a-b36b-3827bc20669b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.127702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:8d452cf16fcd3d4d5e6c0146273b90b2a0c32eda2b1ffeef4fc310f7894fac6b

Observation 7076f53d-f1f8-474b-956c-aaddbd3892dc · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.130453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:349e177fb92263f31ea6c414d65985c065a4315066c9b8e371adceb6f3888421

Observation da9ee582-868e-45f1-a06d-751106e9c93a · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.132731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:3082ce7776fea773d42c4cd4e39d2262d32cec97d316750790a821b3434a7af0

Observation d5deb67a-ce15-48fa-ac4b-50ebb92cbe7d · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.135158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:71b128c16c3a858150df66d0305cf8a00110cb73534bceb56dd54b267c1a7360

Observation 6b536b42-0532-4243-8aa7-e0cc01c4a81a · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.137643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:3995dd7010e71972862e868e54347c99a8b66ce82de11cae861351688e4bdac1

Observation 057e586d-008b-4832-b47e-429b76f7e8f2 · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 20

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T13:37:51.140204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:1a42373f5f4f50b3ad5ffc99319b8a0ba8566d11431643aaacc5a643d62a4fff

Observation 7916eca4-cbc1-4fb4-945c-9b1933db954b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 21

Resolution
parse uncertain
raw_fallback, observed 2026-05-16T13:37:51.142465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:e2e0bfc3ad75957e57100b8b87ae3348962b180ddc52b326c7573cb18ae1fef0

Observation af755b84-41b2-4e3a-be8a-3ff8821da76b · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.144688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:443ed403609e439c7ed0f9824ca85c1fd7d18791df5434bb23cd0ba63bab1d15

Observation a16ecced-68fd-4e4e-8c10-7cb5b99b973e · outbound

This paper cites an unresolved cited work.

Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T13:37:51.146842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:d3452892c10214739997336ea94ca70364b3e3ff4ab050713eed9e4b101b84da

Observation f6993bc0-5fda-48f1-99e0-57f59dacc218 · outbound

This paper cites Let’s convert10010 to base six using Python.

Spurious Rewards: Rethinking Training Signals in RLVR Let’s convert10010 to base six using Python

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T13:37:51.149197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:37:51.086217Z digest=sha256:70d8efe70358ff6eafa489254cdbc8ac923f0a33d5996cb09129697ad4f58d82

Pith citing papers

Observation 41dbb91b-0dd0-4c7b-b90f-4ef002dcd815 · inbound

PRL: Prompts from Reinforcement Learning cites this paper.

PRL: Prompts from Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T14:26:40.402303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:25:22.002977Z digest=sha256:4e124b8697d0c503073cc7afa270129fa30a27a88124c04ebd7640fe64412f6c

Observation 4c0b8fcf-9b99-4446-8d7e-3af410306ce9 · inbound

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling cites this paper.

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:52:05.462042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:52:05.462042Z digest=sha256:f3bbc7aaa43fed6549d638156303460d2ee1158b47e581e33abf842f5967e6ff

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · inbound

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation cites this paper.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:bcb24f697d7fe5ee8b9570849b6266fb17874fd7ebd2f622c601ea7918d8d1d0

Observation 4cef041f-37dd-4886-a7f6-26e174124b9c · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.942088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.942088Z digest=sha256:1ec3c57d179d7900102ddae972c2b3e02e971eea4d748a87e9e8b82e79ef35eb

Observation bf7edc7d-f230-4347-9a5c-06a11217689c · inbound

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess cites this paper.

Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:13:15.096866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:13:15.096866Z digest=sha256:12a65c9bf6fbc15bff150ea255a7197729d5e2fb238e25632addebb3a3aa45f3

Observation 349ccc8e-5008-48bf-8d04-fdc34e30e3e7 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:5e1553eb7bffe2ec4eac63d821bd8826422be88a39c4ee20d1f08cfa644a4585

Observation 9ebc9ced-6798-441c-9ce9-e5b486c46400 · inbound

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning? cites this paper.

KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning? Spurious Rewards: Rethinking Training Signals in RLVR

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:17:06.284535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:17:06.284535Z digest=sha256:0fe95c4e71b8c5f4984fe1f0555634be9a943b004ff81b3544eeeb43b9471d3a

Observation c615db8a-b98d-4bb0-aa2a-419d1f96996b · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Spurious Rewards: Rethinking Training Signals in RLVR

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.176219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.176219Z digest=sha256:8d85a176976459c27ea5cc10affe59393db72f3a57901e8441419d0f1a9cac43

Observation 24d1b1ce-3df5-47ed-81f1-c779200c85d5 · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding Spurious Rewards: Rethinking Training Signals in RLVR

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:13.444701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:13.444701Z digest=sha256:abb8b2b02991f673117f44333b88ee853d21c816d0fe9f983e1737d09393f941

Observation e6e576f3-303d-4653-9fa9-de1668690b16 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.501625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.501625Z digest=sha256:e735f229711c831c6adc89c211c5803496858bb405419aecc9953bd7519f436a

Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · inbound

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning cites this paper.

ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:11.863783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:12:11.863783Z digest=sha256:88c55161bb7ad2773e52429371ad7aea0787f81897fd26bfeaf6dae130dd02eb

Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.167577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.167577Z digest=sha256:788d97637bab1db91b1882da8f0ee00951fd46d23d0c8c14b113fa0780437c99

Observation 936e5643-2132-4e42-b7ff-7ab06ed3bc6b · inbound

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism cites this paper.

ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:02.795840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:02.795840Z digest=sha256:a45fb80e7a40cfde5ad93e58ee00ae2724bdd4ca156ba9f2b591ab0d29bb4e9b

Observation e84b70b4-5be6-47c1-8bf3-377ccab63756 · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T21:06:50.837840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:b8a514be20964e35eb1b3862a9ada3dd5ddf00e6864d99d4fd5a8f8b9075c7a9

Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.869115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.869115Z digest=sha256:465fb3aeb1244a8aa642e1902be5c633f0d1650ac46b54f034bed90749ec401f

Observation 37a43eb2-b4c8-43e4-a591-5430be5eb915 · inbound

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies cites this paper.

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Spurious Rewards: Rethinking Training Signals in RLVR

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:31:29.537032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:31:11.974395Z digest=sha256:148a7dbc66c24e86444cf55c74ee32d0894a40e8ff6738001ebafcc15524334a

Observation d69aa6bd-aea1-46e4-914e-affbf4746d30 · inbound

A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining cites this paper.

A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:59:01.622788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:59:01.622788Z digest=sha256:0316a97870bb3691b4bd7b1564ab56c7863533c1ec771498c11d0e47749b34ad

Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · inbound

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning cites this paper.

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:23:20.230179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:23:20.230179Z digest=sha256:5666e701870e5a0a6ec26d7824b76a6bbf31883d3d0270a6cb6c8809a166c721

Observation cdfd8145-d278-4c87-bc4a-31f25d7eed51 · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:31.017232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:31.017232Z digest=sha256:059d56d028e88a5c59aba566a5a4ab33e7a46a50a9d241f7a40895e541c0281b

Observation 74485d24-0a0e-45a5-88f1-77a2265e499a · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:12:23.642649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:11:29.205366Z digest=sha256:4f301f7b4c7e8f4207960c171efa11606e70f08898e7474ed860dc84bcc75eff

Observation 335e52f9-56cd-4d6e-b4f0-1f8cd628784a · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:10:34.994167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:06:16.172916Z digest=sha256:4aecc2662706b9e85ed64739cfb7ce1149040ce9094f4ed6a9b6f9859eea6a03

Observation 1d569e59-5604-43f3-ab7f-8d9ffc7f311f · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:07.162212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:07.162212Z digest=sha256:2f69a3077740f262c6aa98189c4a354ffcf604a4c715f5f62d384259010fa96e

Observation cd0c33e4-7ff5-4b20-b972-aed121ca57f9 · inbound

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards cites this paper.

Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Spurious Rewards: Rethinking Training Signals in RLVR

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:40:17.700706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:37:54.702010Z digest=sha256:c54acd6ee9b25aa55c672d2aae8976ce4fee14ddbd590dffedd049109355838f

Observation 1b0b2d8a-1d75-447c-9764-7a64e8d2eed6 · inbound

ThetaEvolve: Test-time Learning on Open Problems cites this paper.

ThetaEvolve: Test-time Learning on Open Problems Spurious Rewards: Rethinking Training Signals in RLVR

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:16:27.293449Z digest=sha256:153dad0f3c8896de1692f81e50d8b46876ad0302ea87c449a2eea24213791fe9

Observation 5ec94834-9f43-4502-8714-d84f68b76924 · inbound

What Is Preference Optimization Doing, and Why? cites this paper.

What Is Preference Optimization Doing, and Why? Spurious Rewards: Rethinking Training Signals in RLVR

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:40:28.955681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T18:37:48.161545Z digest=sha256:aaa05edcf528876e3611649f072227d048f8d1a4a0e77cc81cbda2b9dfeb058d

Observation b03b36f4-d86f-4452-be41-ff46f16d8cf7 · inbound

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs cites this paper.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Spurious Rewards: Rethinking Training Signals in RLVR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.212310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.212310Z digest=sha256:8c4a202cdb79abf4c496017c2c0836856b5bbf06a9686e5254166f8849dbf41d

Observation 2a3f5c8f-3e25-4b3a-b2a9-fe248c574759 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.283684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.283684Z digest=sha256:d8f57b14c1abd1f7cf5981088521394e0291acdf57432216bf3e154107ac776f

Observation 0359888a-b996-489a-a866-91570c1013e8 · inbound

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics cites this paper.

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Spurious Rewards: Rethinking Training Signals in RLVR

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T23:11:55.977680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:11:55.977680Z digest=sha256:c22b5bdbce6446454260ddd57fdeefad92fe2e45c0145da2ed0c638b16cfd6e8

Observation 33539e0f-63bf-4dd7-99df-98f992a9cac4 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T19:23:10.901441Z digest=sha256:570b27b7e9616cad78a3ee5dcd46868a273f0a2bdd6b926edec666795aa4509d

Observation 00440b3d-9da3-4d3c-9dec-6331025970ee · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:29.593103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:29.593103Z digest=sha256:ed01de5101843026a6a673c285ab8ab7a3918721ad7cd62a2f93ea41d955f88f

Observation 888f70ca-4c29-4f1e-b9cb-33d0d227567c · inbound

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning cites this paper.

Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:30:40.368494Z digest=sha256:3f2f04dcad5a25bc0ac0f97f0a706e047c675d3c22a2230eea17bc0f210ac108

Observation 766f1fee-28d3-414c-892a-3872cab38956 · inbound

Beyond Distribution Sharpening: The Importance of Task Rewards cites this paper.

Beyond Distribution Sharpening: The Importance of Task Rewards Spurious Rewards: Rethinking Training Signals in RLVR

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T08:07:14.691463Z digest=sha256:e9d67faaec7d010699d81147e94445c4b6fbbf93d37682a980a0d9ed656c7729

Observation fa7f8d2c-efbb-4c43-bc3a-a4de58daa56a · inbound

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions cites this paper.

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions Spurious Rewards: Rethinking Training Signals in RLVR

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:38:22.041842Z digest=sha256:41e0db0852b908c3bf82ca5b3df9620e81a7d1decd8fbd18ae607e6ad297e4f9

Observation c4a486c9-bda1-4833-9791-dc67fc342346 · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills Spurious Rewards: Rethinking Training Signals in RLVR

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:019e75291d7198d5523590146799baca4e862cd12b33cacee190c398ded8be4a

Observation 8597c541-ade7-44b2-990d-16382fbae164 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Spurious Rewards: Rethinking Training Signals in RLVR

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:3b8a793f83a15f84a0f1733b1442dced99f7d0f72a18030d72d239c80374d8f7

Observation c955551c-4937-4047-b106-13062012e8e4 · inbound

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning cites this paper.

SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T06:44:34.968571Z digest=sha256:6d12970e6ad76f63ace4221d92c1711da5af38ecec339c5edac32a96e3a310b9

Observation 9d6aea07-dd05-4d32-abd4-019fd8f4337e · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:de8fedb60a30b224ac30683f47e951318cf4cca708c29016e99c3ac2fa934696

Observation f0af7455-69d4-43f4-a332-c2eef49e8601 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:538135c2c89c80d49237746c76ccb7eb00d4517a5db6f1b02bec003976b9b29b

Observation c23d3da4-be60-49fc-9edb-5f45b380306f · inbound

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling cites this paper.

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T18:58:22.865080Z digest=sha256:24cf70a25fa16730ec8df80f4055f738951da30ecf68e45ad23b8edf8e97df4d

Observation c802e97d-8c70-4acd-bf93-69ab6ad87b2a · inbound

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling cites this paper.

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:14:17.870184Z digest=sha256:b35bbe314ab4f8c498d9a628ae5dad6fac3dfae390a603a978b7f511653bc51b

Observation 3c315a2b-e9eb-49b3-ae7e-cf7e89165677 · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:064a0d8755b801c39e54ecb604514d67bd5ecf73232f7798750a2d425d23d552

Observation 3d0b2257-17ee-4361-83c5-0afc8b6c230a · inbound

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation cites this paper.

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T16:53:00.860162Z digest=sha256:a2fccf8022d23cd0c90c5754da3463d6e03674f79246e9e151e34f043d2bea80

Observation 50827bcc-6bfd-4371-ba76-e13d04f92767 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:a575e408ed2c9947c7b595206f25b779217a85ed70146c680c2c230096a92119

Observation 1f867728-d6a6-491d-b6aa-2af9f8caa81e · inbound

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models cites this paper.

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:01:15.799220Z digest=sha256:2b37354089e833658b990135968164fca036d9f8abb9ad9e6b1813536893d72b

Observation cfdb2cf9-250d-4247-80ff-9601d85bf7b4 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:df560836ab66b42784ab5ff372d02ce807fcbf80bf0d09414bec012ff7c08f44

Observation 726a0883-aeaf-4dd3-a01c-6345737dee68 · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T10:01:22.952772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:8bf8dc07af392beb196aceb34d5bdad61316d89a1b652a73dc1fc6e6a5f86a18

Observation dc8355c6-b564-4f6c-a35c-fce4c1af9507 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:0c58d782274558074457c89508783b172796f99c05cd95ee8600930715ec8ae6

Observation 2a608b4e-ddbc-43ba-8a18-56c3a284fbe7 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:37:51.161107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:62e575fa3725fd509c420b08f17b51f2ef237fdcd40ea5ed8c5a22838f14206f

Observation 982276db-f770-4500-a269-4fe318c9a47a · inbound

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP cites this paper.

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP Spurious Rewards: Rethinking Training Signals in RLVR

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:33:03.662351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c68b7f0d8a2fb76314920785f26114e6f9a2aa6fd5ec9aef47e0dc471aa66014

Observation a9bcd287-bdb8-46ac-af5f-ae14cf350702 · inbound

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR cites this paper.

Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:03:03.508869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:02:35.271960Z digest=sha256:306653aa9c3542398de9ee89bb9c028d619d9d424c8e34af277b93a0211a5409

Observation 25f5bb25-b4c2-4e0b-9f3c-1341ed627632 · inbound

Label-Free Reinforcement Learning via Cross-Model Entropy cites this paper.

Label-Free Reinforcement Learning via Cross-Model Entropy Spurious Rewards: Rethinking Training Signals in RLVR

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T14:23:30.616183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:18:38.733276Z digest=sha256:dafb3a300b5ec3941040783c129dc3e544c67b2d4c65082f0c210b103794afc4

Observation 43107580-6423-41fe-8bb9-d97f04f04b5d · inbound

Reasoning with Sampling: Cutting at Decision Points cites this paper.

Reasoning with Sampling: Cutting at Decision Points Spurious Rewards: Rethinking Training Signals in RLVR

Reference 3

Resolution
malformed identifier
local_arxiv, observed 2026-06-29T08:23:15.534726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:17:08.488565Z digest=sha256:bb6be96e645452216012dceb180c0bcbe91a2966a994c70691a5353f64679849

Observation fb845395-bada-49c1-aeb5-8b25ae86e72a · inbound

Consolidating Rewarded Perturbations for LLM Post-Training cites this paper.

Consolidating Rewarded Perturbations for LLM Post-Training Spurious Rewards: Rethinking Training Signals in RLVR

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:42:46.870525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:37:34.673402Z digest=sha256:f17e24766a2472e6082b9d2ef26785ed5cb8a10cfaa49051b18f1aa98a9a7d01

Observation 03f90375-e681-4bc1-915a-4f83010bb428 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.873180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:c46c6396dd44cf6565155083334914b229a0a8f8b82d15331d5ebe2ebe579967

Observation c9a30e50-d8dc-478f-a9ed-d53e814b1198 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.467969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:f4780999dd9055f7e0b78379d504d7482ee9123fbb05e88abbc31fddd90b26cd

Observation e8950c1b-8edd-464d-bc8d-23c354c61ced · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Spurious Rewards: Rethinking Training Signals in RLVR

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:40.763205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:578c5f8420d9d2a642f6575a82e86e5f1ce23aa5cac824b0163f5a9eb3622df2

Observation e3aadc56-5a37-423e-a0c8-87ff3bbb679c · inbound

A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR cites this paper.

A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:06:59.082200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:35:26.306648Z digest=sha256:134f86d45f146a2a789f9de28bdd40ec132c49a1214f8004d35b15c8d39b8f23

Observation e497ec34-05e1-4055-b5dd-2c784799cdb0 · inbound

RREDCoT: Segment-Level Reward Redistribution for Reasoning Models cites this paper.

RREDCoT: Segment-Level Reward Redistribution for Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 137

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:57.646397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T02:12:16.029526Z digest=sha256:d1327c91c31a6f65bfe369f168bdfdc3fbf481800b9a083d165d1edc0e346f41

Observation 4ca2855a-fccd-46ad-a9e0-0db11187f3a5 · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T13:20:56.495398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:cd8feab966f3c6360bcd2d5ebbd3acf7654a4d061cfc8bdebdb31d3189a3f188

Observation 028afbad-e9eb-470a-b0e9-1f03e1a6cc86 · inbound

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO cites this paper.

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T11:54:46.435397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:54:46.435397Z digest=sha256:8ef092fbbb4ffb3f9a19cbf4cfbd451aa44858bcf9e0d7ec3a3ce0178acf8df5

Observation 701020a6-8764-4e8c-886e-c8bbc61bc510 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:38:55.978596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:8e513a61e1c3fa74f01837a09ba7269f78db082b2eaf62ad465c5ca4537a3e22

Observation 45b40a9c-d615-4e3b-b516-ed6555117f7a · inbound

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents cites this paper.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-04T20:40:07.890098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:55:51.114244Z digest=sha256:8ab3243155d95032f51a5435e081b2144a128bb9dfe7f330a522204d31852751

Observation 520d274b-f92c-4b51-a10f-b89807016227 · inbound

RLVP: Penalize the Path, Reward the Outcome cites this paper.

RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-09T11:16:11.379832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T11:15:06.569421Z digest=sha256:d2fa6e5a0bce36916e960a4463caffc8ff60884885af44b7f5438d5265427e97

Observation 07f73784-95e0-4cbd-8408-e75d5948032c · inbound

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning cites this paper.

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T09:26:08.755245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-09T09:20:08.212837Z digest=sha256:5f12355db71ac4ca149fb782fe769d35e8a7b7454d8eb465a960dee78de0f59a

Observation 353c0b54-a1fe-4d06-adee-9740cc3fc99d · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:059176bf8eb61a3b3d90488519e9eeb70b372cda6ad912106601dc007834a661

Observation 35fffbe0-93cb-4d5a-8829-f4f83000f389 · inbound

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning cites this paper.

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T17:39:27.292075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:39:27.292075Z digest=sha256:7d3d2d771e0a8877fc48265a4ff70d5e84b8da81a21b1a8e6b7130c390871b83

Observation fb1244d4-f004-4d78-b606-b1105598a081 · inbound

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR cites this paper.

When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T07:36:47.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:36:47.336670Z digest=sha256:c8b0224e3578591945be2d1954254fae70c8901dd00b8cccc5993895520fcb01

Observation 3348cd36-1e17-4155-90d2-f7211246e2ce · inbound

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift cites this paper.

Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:41.835722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:41.835722Z digest=sha256:56257bd1ab063edc0fa0a8e474ae82d80dc61644409842086fda136c399a5b74

Observation 66b24e57-486c-443f-a922-b8654f6369e8 · inbound

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models cites this paper.

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T02:03:05.632468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T02:03:05.632468Z digest=sha256:9fccd799573de271b45f47a7276b44c5a6c30126d609471559dc16375be7a587