Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2508.11800.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:01:22.079256Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T19:47:18.750963Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 0d6f028a-9455-4fd4-8cc9-5ff2f15a4b4c · inbound
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b352ce26-5105-4fc3-b17e-253794193c8c · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2557f5c0-3089-4a60-87eb-5e0a4bd91a42 · inbound
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 81733b23-02e0-4fbd-ae1e-0511fee1aa76 · inbound
OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation db227cbe-956f-41a8-a41f-88ed10404e9b · inbound
Calibration-Aware Policy Optimization for Reasoning LLMs Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3cd94be1-17d4-453e-b9eb-7a9371f59825 · inbound
Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de9c0fdf-6ddf-482f-968d-4fe490c4ae4c · inbound
Verifiable Rewards for Calibrated Probabilistic Forecasting Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d486a7fa-41ef-40f6-972b-059fe2b112e7 · inbound
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b803a01d-d71b-4ad8-9af4-bf3bdf8b73e2 · inbound
Multimodal Reward Hacking in Reinforcement Learning Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88231fcd-a300-4020-ad24-46173b712127 · inbound
The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.