Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T13:37:51.086217Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 69 inbound Pith citation observations for arXiv:2506.10947.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T13:37:51.086217Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:52:05.462042Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
21 of 21 outbound references displayed
External citation measurements
8
pith, observed 2026-08-05T02:28:24.338817Z
Observation 4a385ed9-c2c4-4bf4-9ed3-c95874899ef5 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR doi: 10.1038/s41586-025-09422-z
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2dea4de5-be44-48d8-b64c-93c47805eaa6 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR 2 OLMo 2 Furious
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a5e8549-3eef-4533-a2e1-d79e8df51518 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Maximizing Confidence Alone Improves Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9d64740-703c-4c47-b418-e60d75ab1a8f · outbound
Spurious Rewards: Rethinking Training Signals in RLVR ISBN 979-8-89176-288-6
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a1f7f2c2-4234-4f6a-ad6b-3b505e044238 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59b052cd-f9ff-4880-9dd4-617d1ce15ab0 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bde87d3a-33ec-4a7e-95c3-0ba6794ae34c · outbound
Spurious Rewards: Rethinking Training Signals in RLVR clipping bias
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 973f233d-76be-4905-85e2-892b00cd5d91 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 600501ef-5d6f-4b5a-8405-621481040869 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR r={r},θ= {theta}
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a01737a-1ad0-401c-ab0d-6290fa19e4f9 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41602f8a-22e4-4d04-b61a-fbd9a478714f · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a629e801-ab4c-410a-b36b-3827bc20669b · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7076f53d-f1f8-474b-956c-aaddbd3892dc · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da9ee582-868e-45f1-a06d-751106e9c93a · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5deb67a-ce15-48fa-ac4b-50ebb92cbe7d · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b536b42-0532-4243-8aa7-e0cc01c4a81a · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 057e586d-008b-4832-b47e-429b76f7e8f2 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7916eca4-cbc1-4fb4-945c-9b1933db954b · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af755b84-41b2-4e3a-be8a-3ff8821da76b · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a16ecced-68fd-4e4e-8c10-7cb5b99b973e · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6993bc0-5fda-48f1-99e0-57f59dacc218 · outbound
Spurious Rewards: Rethinking Training Signals in RLVR Let’s convert10010 to base six using Python
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41dbb91b-0dd0-4c7b-b90f-4ef002dcd815 · inbound
PRL: Prompts from Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c0b8fcf-9b99-4446-8d7e-3af410306ce9 · inbound
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling Spurious Rewards: Rethinking Training Signals in RLVR
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · inbound
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cef041f-37dd-4886-a7f6-26e174124b9c · inbound
ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context Spurious Rewards: Rethinking Training Signals in RLVR
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7edc7d-f230-4347-9a5c-06a11217689c · inbound
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 349ccc8e-5008-48bf-8d04-fdc34e30e3e7 · inbound
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ebc9ced-6798-441c-9ce9-e5b486c46400 · inbound
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning? Spurious Rewards: Rethinking Training Signals in RLVR
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c615db8a-b98d-4bb0-aa2a-419d1f96996b · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Spurious Rewards: Rethinking Training Signals in RLVR
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d1b1ce-3df5-47ed-81f1-c779200c85d5 · inbound
GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding Spurious Rewards: Rethinking Training Signals in RLVR
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e576f3-303d-4653-9fa9-de1668690b16 · inbound
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adda67e-b61d-4bfa-a97b-07329c450e82 · inbound
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936e5643-2132-4e42-b7ff-7ab06ed3bc6b · inbound
ETTRL: Balancing Exploration and Exploitation in LLM Test-Time Reinforcement Learning Via Entropy Mechanism Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84b70b4-5be6-47c1-8bf3-377ccab63756 · inbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · inbound
Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a43eb2-b4c8-43e4-a591-5430be5eb915 · inbound
The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies Spurious Rewards: Rethinking Training Signals in RLVR
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d69aa6bd-aea1-46e4-914e-affbf4746d30 · inbound
A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining Spurious Rewards: Rethinking Training Signals in RLVR
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28a2971-a4d6-4140-b04f-4f1f68ed129d · inbound
AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfd8145-d278-4c87-bc4a-31f25d7eed51 · inbound
Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74485d24-0a0e-45a5-88f1-77a2265e499a · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 335e52f9-56cd-4d6e-b4f0-1f8cd628784a · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d569e59-5604-43f3-ab7f-8d9ffc7f311f · inbound
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0c33e4-7ff5-4b20-b972-aed121ca57f9 · inbound
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards Spurious Rewards: Rethinking Training Signals in RLVR
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b0b2d8a-1d75-447c-9764-7a64e8d2eed6 · inbound
ThetaEvolve: Test-time Learning on Open Problems Spurious Rewards: Rethinking Training Signals in RLVR
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ec94834-9f43-4502-8714-d84f68b76924 · inbound
What Is Preference Optimization Doing, and Why? Spurious Rewards: Rethinking Training Signals in RLVR
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b03b36f4-d86f-4452-be41-ff46f16d8cf7 · inbound
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs Spurious Rewards: Rethinking Training Signals in RLVR
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3f5c8f-3e25-4b3a-b2a9-fe248c574759 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0359888a-b996-489a-a866-91570c1013e8 · inbound
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics Spurious Rewards: Rethinking Training Signals in RLVR
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33539e0f-63bf-4dd7-99df-98f992a9cac4 · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00440b3d-9da3-4d3c-9dec-6331025970ee · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Spurious Rewards: Rethinking Training Signals in RLVR
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 888f70ca-4c29-4f1e-b9cb-33d0d227567c · inbound
Placing Puzzle Pieces Where They Matter: A Question Augmentation Framework for Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 766f1fee-28d3-414c-892a-3872cab38956 · inbound
Beyond Distribution Sharpening: The Importance of Task Rewards Spurious Rewards: Rethinking Training Signals in RLVR
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa7f8d2c-efbb-4c43-bc3a-a4de58daa56a · inbound
A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions Spurious Rewards: Rethinking Training Signals in RLVR
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4a486c9-bda1-4833-9791-dc67fc342346 · inbound
Characterizing Model-Native Skills Spurious Rewards: Rethinking Training Signals in RLVR
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8597c541-ade7-44b2-990d-16382fbae164 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment Spurious Rewards: Rethinking Training Signals in RLVR
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c955551c-4937-4047-b106-13062012e8e4 · inbound
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d6aea07-dd05-4d32-abd4-019fd8f4337e · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0af7455-69d4-43f4-a332-c2eef49e8601 · inbound
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c23d3da4-be60-49fc-9edb-5f45b380306f · inbound
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c802e97d-8c70-4acd-bf93-69ab6ad87b2a · inbound
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling Spurious Rewards: Rethinking Training Signals in RLVR
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c315a2b-e9eb-49b3-ae7e-cf7e89165677 · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d0b2257-17ee-4361-83c5-0afc8b6c230a · inbound
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50827bcc-6bfd-4371-ba76-e13d04f92767 · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f867728-d6a6-491d-b6aa-2af9f8caa81e · inbound
Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfdb2cf9-250d-4247-80ff-9601d85bf7b4 · inbound
Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 726a0883-aeaf-4dd3-a01c-6345737dee68 · inbound
Holder Policy Optimisation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc8355c6-b564-4f6c-a35c-fce4c1af9507 · inbound
Reward Hacking in Rubric-Based Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a608b4e-ddbc-43ba-8a18-56c3a284fbe7 · inbound
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Spurious Rewards: Rethinking Training Signals in RLVR
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 982276db-f770-4500-a269-4fe318c9a47a · inbound
Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP Spurious Rewards: Rethinking Training Signals in RLVR
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9bcd287-bdb8-46ac-af5f-ae14cf350702 · inbound
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Spurious Rewards: Rethinking Training Signals in RLVR
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25f5bb25-b4c2-4e0b-9f3c-1341ed627632 · inbound
Label-Free Reinforcement Learning via Cross-Model Entropy Spurious Rewards: Rethinking Training Signals in RLVR
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43107580-6423-41fe-8bb9-d97f04f04b5d · inbound
Reasoning with Sampling: Cutting at Decision Points Spurious Rewards: Rethinking Training Signals in RLVR
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb845395-bada-49c1-aeb5-8b25ae86e72a · inbound
Consolidating Rewarded Perturbations for LLM Post-Training Spurious Rewards: Rethinking Training Signals in RLVR
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03f90375-e681-4bc1-915a-4f83010bb428 · inbound
On the Generalization Gap in Self-Evolving Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9a30e50-d8dc-478f-a9ed-d53e814b1198 · inbound
Trust Region On-Policy Distillation Spurious Rewards: Rethinking Training Signals in RLVR
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8950c1b-8edd-464d-bc8d-23c354c61ced · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Spurious Rewards: Rethinking Training Signals in RLVR
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3aadc56-5a37-423e-a0c8-87ff3bbb679c · inbound
A Pre-Registered Causal Partition of Self-Consistency Elicitation and Reward Design in RLVR Spurious Rewards: Rethinking Training Signals in RLVR
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e497ec34-05e1-4055-b5dd-2c784799cdb0 · inbound
RREDCoT: Segment-Level Reward Redistribution for Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ca2855a-fccd-46ad-a9e0-0db11187f3a5 · inbound
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 028afbad-e9eb-470a-b0e9-1f03e1a6cc86 · inbound
It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO Spurious Rewards: Rethinking Training Signals in RLVR
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701020a6-8764-4e8c-886e-c8bbc61bc510 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45b40a9c-d615-4e3b-b516-ed6555117f7a · inbound
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Spurious Rewards: Rethinking Training Signals in RLVR
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 520d274b-f92c-4b51-a10f-b89807016227 · inbound
RLVP: Penalize the Path, Reward the Outcome Spurious Rewards: Rethinking Training Signals in RLVR
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07f73784-95e0-4cbd-8408-e75d5948032c · inbound
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 353c0b54-a1fe-4d06-adee-9740cc3fc99d · inbound
Multimodal Reward Hacking in Reinforcement Learning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fffbe0-93cb-4d5a-8829-f4f83000f389 · inbound
Depth-Entropy Guided Sampling for Training-Free LLM Reasoning Spurious Rewards: Rethinking Training Signals in RLVR
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb1244d4-f004-4d78-b606-b1105598a081 · inbound
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR Spurious Rewards: Rethinking Training Signals in RLVR
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3348cd36-1e17-4155-90d2-f7211246e2ce · inbound
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift Spurious Rewards: Rethinking Training Signals in RLVR
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b24e57-486c-443f-a922-b8654f6369e8 · inbound
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models Spurious Rewards: Rethinking Training Signals in RLVR
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.