Pith. sign in

Paper Citation Record · LEDGER

Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2508.11800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11800 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:01:22.079256Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:47:18.750963Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0d6f028a-9455-4fd4-8cc9-5ff2f15a4b4c · inbound

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search cites this paper.

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:12:35.934261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T12:12:25.437344Z digest=sha256:c7aac6d0a87a5fc39f8a6c71957ebe2135eb5db3cb87cc245b94699e107a9264

Observation b352ce26-5105-4fc3-b17e-253794193c8c · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:26.480554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:2f1b7baa96ba801624d0ca1467d6c10a4139a1f469ed9547ab91958eb85495b3

Observation 2557f5c0-3089-4a60-87eb-5e0a4bd91a42 · inbound

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards cites this paper.

Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:50:02.363148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T13:49:11.758336Z digest=sha256:0184c387d0ab3ed5cca4074870c55ffb7485c977816944d8a416cf8e4822d491

Observation 81733b23-02e0-4fbd-ae1e-0511fee1aa76 · inbound

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks cites this paper.

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.698585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:09:33.968186Z digest=sha256:143f3ecc3eaff051bbd443cb6d551ec69d2a8ffb95ed73267ec6671c1d96cbfc

Observation db227cbe-956f-41a8-a41f-88ed10404e9b · inbound

Calibration-Aware Policy Optimization for Reasoning LLMs cites this paper.

Calibration-Aware Policy Optimization for Reasoning LLMs Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:01.149053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T15:40:24.396851Z digest=sha256:c090a82919b3278ff2ba49f1dcedf1ab5c87730e0c1009eb5bf9006c210d967d

Observation 3cd94be1-17d4-453e-b9eb-7a9371f59825 · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:29.523551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:d6b3cb5ffafb382a5fabde30bef0a744e1b64f10730453ec0e0c8303e128711e

Observation de9c0fdf-6ddf-482f-968d-4fe490c4ae4c · inbound

Verifiable Rewards for Calibrated Probabilistic Forecasting cites this paper.

Verifiable Rewards for Calibrated Probabilistic Forecasting Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:18.752660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T19:45:23.721172Z digest=sha256:c2852459b0149a6f53cbf6f17f7aa718accfd0dce712d0c5a57d6a692d228727

Observation d486a7fa-41ef-40f6-972b-059fe2b112e7 · inbound

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models cites this paper.

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T20:01:06.623978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:01:06.623978Z digest=sha256:615edccf36c3155b624235445c65ebc2ac43f848d89c8c7035e630778aa0ad65

Observation b803a01d-d71b-4ad8-9af4-bf3bdf8b73e2 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:ce2d9ccce54ac42ac30a73723035877ca55fc24babdc298591435c0fb475da58

Observation 88231fcd-a300-4020-ad24-46173b712127 · inbound

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works cites this paper.

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T08:01:22.079256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:01:22.079256Z digest=sha256:ca95cd94cd9f95e7c98c0582e8343a2f47431f90b6b867b2c9bd4c8273dd31ba