Pith. sign in

Paper Citation Record · LEDGER

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2508.05928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05928 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:10:24.550084Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T22:24:41.760120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:25:22.546587Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef0fa1f5-a63c-423d-94a5-c8b7a24b21d8 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.461471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.461471Z digest=sha256:7fdcb6dd34f940e85c24db10755365f5642ae3c9dcedc4b1cc1d986866c9b3f9

Observation 79ab36ea-c431-47e4-a763-4f9334673b60 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.465211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.465211Z digest=sha256:e78c80996863d3345f2db6370794749fd62b3573ea5d0da6066bfac9efe3e03c

Observation cec1ec34-ebe4-4254-89d9-788162f79dc8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.469058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.469058Z digest=sha256:28450d6fb489518ee1ac0a7dad1edb6849868867d42c29d1911dea32987e1396

Observation 7e1d3b67-3b55-40ea-a6d0-e840fcdd5b6c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Understanding R1-Zero-Like Training: A Critical Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.480610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.480610Z digest=sha256:9bf20720e78ad896ae871f4881ac6673daee0815fbb7b7348f3087821682141f

Observation cb019d9a-d1dc-450b-bf12-b6e2b5bda1ca · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.484196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.484196Z digest=sha256:97eefe8db0351921bdf5622a01e12bcff684dbcd99f6770d2670ac780c36cf39

Observation 66eb2392-1fc5-42a5-8587-b2612640d9de · outbound

This paper cites s1: Simple test-time scaling.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.491071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.491071Z digest=sha256:e346f842389d6820bf4b46bbfbff40b8dacb7fba945ff9bbb9ba0d954c1fd349

Observation a284d844-e747-4df6-81cc-0122ea43d495 · outbound

This paper cites On Symmetric Losses for Robust Policy Optimization with Noisy Preferences.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.494910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.494910Z digest=sha256:ae3738fe7c75995cb06222a500414cff6a7a574e3d61c6674f0f845022a9cabc

Observation 50c28dc4-55a5-4d1d-aa32-733ef375303d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.498499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.498499Z digest=sha256:6ecce34d55389a293e99b5d206eb902c5bd4eb84f5a420d78422efe2819c9a2e

Observation 116f5781-46de-4288-863f-78557b93c382 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.505447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.505447Z digest=sha256:a0f26175485da97581ffd9f4d9c41bac179b65b41de2d9b952b1b377c9620617

Observation ceacfdd0-4f3e-4282-818a-3df7b6b59151 · outbound

This paper cites Long Is More Important Than Difficult for Training Reasoning Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Long Is More Important Than Difficult for Training Reasoning Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:10:24.698575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T23:10:24.509786Z digest=sha256:3130c53a31de65d47a2538ddeafd620f5dd3a8a8f4c8ce04856dda1bf1d12f35

Observation 8405bc8a-abd0-4739-b46f-0993e3b2409e · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.513775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.513775Z digest=sha256:c682da476bad2ecd73bc5328b34e98630667f2116659cdbebb56a18d1612341b

Observation a27d6a7e-75b1-44b2-8bf6-8669645059d0 · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting LLMs cannot find reasoning errors, but can correct them given the error location

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.518161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.518161Z digest=sha256:aff2e53f1c9c2c84e62a82a345f0c88fef9715c0255a7e61a6d793b07412fdd7

Observation 448730c8-885a-4819-9491-e48d9a58ada9 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.525759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.525759Z digest=sha256:e532cc650f13d89d5a6dd2ce07eaed0a1c0d4490bc513394ff7bbca0246a0e93

Observation 7cf9db51-0e08-48f4-90c8-a2c8f098b12a · outbound

This paper cites Bayesian Reward Models for LLM Alignment.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Bayesian Reward Models for LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.529784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.529784Z digest=sha256:6ed78307d88b5b398a132dcf238c0d255bf9a7b5d303af5b8d70d167a30c8abc

Observation 34e2174a-a21a-4f47-bbd6-334613c02e47 · outbound

This paper cites Are Reasoning Models More Prone to Hallucination?.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Are Reasoning Models More Prone to Hallucination?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.533640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.533640Z digest=sha256:266cf6ed2ef0d89165c027fa49ff1d18347f879a5ffd8cd059502914df416f9c

Observation cf79147b-2448-43ca-946e-05d996526292 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.537866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.537866Z digest=sha256:c8f58fa8937820197272f1a919321edc4089a52c9743c265a4594863d179f0b1

Observation 681f7375-76b6-4f50-905d-f111fa72d75c · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.542099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.542099Z digest=sha256:63950dd0ca14371bec1be65581ec4e3877b4edcc8af083a629b32068fc850012

Observation ec5b716e-19b0-4c33-912d-32f541f0a34b · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.546164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.546164Z digest=sha256:49ee1096a2c725d9db9d88c66bfab9b8c91c2590e46b48df31ef51c10c6a098d

Observation 4acf7ad6-eda0-4af7-8c13-6b6d3fab07bd · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting TTRL: Test-Time Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.550084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.550084Z digest=sha256:0dc17e6e1e33a4cdb76729cdfc8107a03bebf8a7f8288cf403f78cf27117b4b1

Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.452911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.452911Z digest=sha256:2cc6b562a0e83f3ab3383fcafae5e495d0546d0d80cab9d0d244be11546696dd

Observation 613c89b2-ed73-4a0b-a26a-bf4af8a9ffa5 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.522055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.522055Z digest=sha256:78a500e5357df56641ec828c0193904ad464838caaedf85887bb48a298035f85

Observation e6e576f3-303d-4653-9fa9-de1668690b16 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.501625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.501625Z digest=sha256:e735f229711c831c6adc89c211c5803496858bb405419aecc9953bd7519f436a

Observation 01562e6f-0fbc-41bd-83a2-ea043bec836c · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.473186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.473186Z digest=sha256:2d1d06d4e057bda3f5bb9fe0a79a2551c292dbc664acf1c6dcea260509df87d0

Observation c3f08595-d6ec-47d1-b174-aa944314cf68 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.477083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.477083Z digest=sha256:bb56823436c7b58fa280e4a4c50ca5b297d27927fdb6610c3709e30395d1ce0f

Observation 9c4e76e6-c273-4b8b-aa6d-237a889f0858 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.487632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.487632Z digest=sha256:3d259a414f2a202a7e1a5664d9c2eb43e0b43b26e3bec63eb23ac0eb029affcb

Observation f0893c3f-5b55-4cad-861a-355f9829a8df · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.456829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.456829Z digest=sha256:5bd7019f43acbb97bcbafd83822d14da5c68a1a2508bb9bdf247d3a546784c7b

Pith citing papers

Observation 863b38eb-7988-4991-b28f-6a14df9dd436 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.548980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:b75ab940a8463f8004f635ebee45f6119cd38102cd84f61d793876fc542ebdcc

Observation 483e9cc9-cc00-47fe-a3bd-7314640d45a7 · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.814962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:7257f19d9ad7c1d040f0286586e6a8f7351f6f98e5b4fb0a9580b07da8129802

Observation e77d17ca-0e9c-42ca-9d00-ceef1daa6a5f · inbound

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization cites this paper.

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:30.130912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:31:42.065093Z digest=sha256:9d1806622a12f43154470aff66d78ef41fbb1f123e03c63f4d268bfa9a9bc9e7