Pith. sign in

Paper Citation Record · LEDGER

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

As of 13 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 6 inbound Pith citation observations for arXiv:2601.11061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.11061 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:15.647182Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T03:13:07.963343Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T03:14:31.665483Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f5a0e18-e609-4c03-8bcd-d7415847e76a · outbound

This paper cites an unresolved cited work.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.280726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.280726Z digest=sha256:0f392f5a23a0e6b25043b09607e21a445e2ddcb6b35aa3e253e6a436a1c01dc9

Observation d77764c6-16dc-4013-baae-24893b47c2f9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.452071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.452071Z digest=sha256:71432fa9f402b1287cbcaea7d71e3931d8d370993deb5b492712e722b491a6f6

Observation 0b09f264-9685-4506-981d-1900be54018f · outbound

This paper cites The reasoning-memorization interplay in language models is mediated by a single direction.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs The reasoning-memorization interplay in language models is mediated by a single direction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.523326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.523326Z digest=sha256:11966484b50dc0e7f5986d1b06eff8997c009ee5c556603df9638c0920997bcc

Observation ed9ed0df-d29b-4f28-be89-c8d8198ed734 · outbound

This paper cites Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.630860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.630860Z digest=sha256:6c2c6044ae1a91e71d5d43f8f59d32f717e5ddfde065400161ef980fbcaa5b77

Observation fdf2b441-794e-4350-bc82-1d0146fe3a2f · outbound

This paper cites HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.814920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.814920Z digest=sha256:e24ee713d3a8b432c20b7dafbd330f75b2f25debf9781f25458a15dceda34dd9

Observation 0ffc9584-3a50-421c-9018-91003611e43e · outbound

This paper cites Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.951041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.951041Z digest=sha256:1d77f5844a1b1ec755f658ad62c3a28497cc96da1a130d1ed348b859f3423310

Observation d3ae321a-ec46-4c73-a723-a87ea545fefd · outbound

This paper cites 2 OLMo 2 Furious.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs 2 OLMo 2 Furious

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.145479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.145479Z digest=sha256:b1cd502d4725c39a440e6bba2e96391f739373fadec08118c0e1267cf9fc84c1

Observation b03b36f4-d86f-4452-be41-ff46f16d8cf7 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Spurious Rewards: Rethinking Training Signals in RLVR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.212310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.212310Z digest=sha256:a63e82435f7111b3f44b7862c509695ca6f005759ef8168b1e007347df455817

Observation 7db093bf-05df-4210-883f-92377ccec7e2 · outbound

This paper cites Detecting Memorization in Large Language Models.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Detecting Memorization in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.310993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.310993Z digest=sha256:3c91252e46a709a6213e4399859aec3147c409263b9ddf5b3a414aa4a6e49202

Observation efa8a444-62f3-40aa-86a1-443620487913 · outbound

This paper cites Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.414661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.414661Z digest=sha256:a4fde7ff8f1ac455b583f1ee598ed83256e19da1ff20789f9f50f0382836783e

Observation a6a6d943-7412-49b1-8e8d-079f10e2ebf3 · outbound

This paper cites Qwen3 Technical Report.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Qwen3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.553819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.553819Z digest=sha256:97990807227b527bd1dd8c6a4710bc96b878ae89d2ff47afdf6679d6ac663718

Observation 6dde6f1f-5748-492b-baf2-b2d3e2dd3021 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.647182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.647182Z digest=sha256:b61ab59afd011d54a738220e066b025ddf29e27bd94eb51b357031cd365c5d76

Observation 838803a9-c6a6-48c2-9e49-9783e1b17d8c · outbound

This paper cites Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:15.021174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:15.021174Z digest=sha256:4b51a35780e455f5334b0f669811315caa825ae8c966c1af7e0d5c4d2e4c43e0

Observation 133578b5-d394-491b-b751-3cfc492c3836 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Transformer feed-forward layers are key-value memories

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.386551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.386551Z digest=sha256:dfbfff7fb85ec24b1e9a3902c297a4c8b3e355ecd24cdd313e6ea1370227080b

Observation 4a88acd3-be15-493f-a424-bfa82670cc09 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.751527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.751527Z digest=sha256:01a5f0bfc47ae8b265b9768b6e3767485297e34e4feae880509521f0aa182033

Observation 746873b4-bf13-44b4-8e96-bf83f4bce58d · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Are Your LLMs Capable of Stable Reasoning?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.881237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.881237Z digest=sha256:66b70cc70556747499914a97587b5f47bbc6ca90fa1052e8be7d7b3bf0140238

Observation 81bd8425-f569-4e5f-af51-1dac109baa92 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.699803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.699803Z digest=sha256:f15c169a4e636ea423aee9a034efb5a7cd7e5f78f450b04b0eb6e37ca12eadbd

Observation fea7e4e0-acbc-4464-9978-08e8b41d89d6 · outbound

This paper cites Exploration vs exploitation: Rethinking rlvr through clipping, entropy, and spurious reward.arXiv preprint arXiv:2512.16912,.

Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Exploration vs exploitation: Rethinking rlvr through clipping, entropy, and spurious reward.arXiv preprint arXiv:2512.16912,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:14.325731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:14.325731Z digest=sha256:1f7be2ed98a1f5a5e8bada30aceeeff062898727459ec62606471fbe7113d723

Pith citing papers

Observation e0d494e5-58fd-4ed6-9b78-1b857687d61a · inbound

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training cites this paper.

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:15:16.925411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T10:44:07.426652Z digest=sha256:266077ce08f2f7a9d4577a1fa8940acebdefecb8958e0df1862fee80c4f52b9c

Observation ae66e80a-5c2d-4298-9785-9922735cc5a2 · inbound

VeriGate: Verifier-Gated Step-Level Supervision for GRPO cites this paper.

VeriGate: Verifier-Gated Step-Level Supervision for GRPO Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:16.058852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T08:47:45.666708Z digest=sha256:e4922dbbe551c270a2024531dc0abd08f31fc65e72fa89449c9e9a5e7dca8c49

Observation a07b929c-f40f-4b06-971a-0eb8ac7915e9 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:29:41.280184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:a6e2aa0c7985a36abd89bc7f5b85b52280cab90fccfc86a4b20303b542c58e67

Observation 69d96bbf-4951-4da1-bd98-967b9b092f60 · inbound

Predictable GRPO: A Closed-Form Model of Training Dynamics cites this paper.

Predictable GRPO: A Closed-Form Model of Training Dynamics Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T06:45:29.599112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-01T06:40:47.911021Z digest=sha256:3f9a107d1e196774923f01abc9d64b8c4f718f5ef5c26a6cd0145b2bf4b68191

Observation c26c71ef-6c1b-40f6-9249-eeac1be6c1fb · inbound

Predictable GRPO: A Closed-Form Model of Training Dynamics cites this paper.

Predictable GRPO: A Closed-Form Model of Training Dynamics Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:27:21.698716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T20:26:15.667646Z digest=sha256:e1bc8ce97949975d75c3c3cd6856e03f7af50193992009721453a9c2c9313730

Observation 88b68295-3d5f-44fa-a4ac-f7208502a78d · inbound

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment cites this paper.

Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-08T03:14:31.666850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-08T03:13:07.963343Z digest=sha256:d0a016c2b841b217b50d40cb4e8fb06480f038b4d3a662ea1c1d26dbecc2f015