Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:15.647182Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 6 inbound Pith citation observations for arXiv:2601.11061.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:15.647182Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-08T03:13:07.963343Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T03:14:31.665483Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f5a0e18-e609-4c03-8bcd-d7415847e76a · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77764c6-16dc-4013-baae-24893b47c2f9 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b09f264-9685-4506-981d-1900be54018f · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs The reasoning-memorization interplay in language models is mediated by a single direction
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9ed0df-d29b-4f28-be89-c8d8198ed734 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdf2b441-794e-4350-bc82-1d0146fe3a2f · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ffc9584-3a50-421c-9018-91003611e43e · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3ae321a-ec46-4c73-a723-a87ea545fefd · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs 2 OLMo 2 Furious
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03b36f4-d86f-4452-be41-ff46f16d8cf7 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Spurious Rewards: Rethinking Training Signals in RLVR
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db093bf-05df-4210-883f-92377ccec7e2 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Detecting Memorization in Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa8a444-62f3-40aa-86a1-443620487913 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a6d943-7412-49b1-8e8d-079f10e2ebf3 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Qwen3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dde6f1f-5748-492b-baf2-b2d3e2dd3021 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 838803a9-c6a6-48c2-9e49-9783e1b17d8c · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 133578b5-d394-491b-b751-3cfc492c3836 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Transformer feed-forward layers are key-value memories
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a88acd3-be15-493f-a424-bfa82670cc09 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Solving Quantitative Reasoning Problems with Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 746873b4-bf13-44b4-8e96-bf83f4bce58d · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Are Your LLMs Capable of Stable Reasoning?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bd8425-f569-4e5f-af51-1dac109baa92 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea7e4e0-acbc-4464-9978-08e8b41d89d6 · outbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Exploration vs exploitation: Rethinking rlvr through clipping, entropy, and spurious reward.arXiv preprint arXiv:2512.16912,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d494e5-58fd-4ed6-9b78-1b857687d61a · inbound
Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ae66e80a-5c2d-4298-9785-9922735cc5a2 · inbound
VeriGate: Verifier-Gated Step-Level Supervision for GRPO Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a07b929c-f40f-4b06-971a-0eb8ac7915e9 · inbound
When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 69d96bbf-4951-4da1-bd98-967b9b092f60 · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c26c71ef-6c1b-40f6-9249-eeac1be6c1fb · inbound
Predictable GRPO: A Closed-Form Model of Training Dynamics Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 88b68295-3d5f-44fa-a4ac-f7208502a78d · inbound
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.