Pith. sign in

Paper Citation Record · LEDGER

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

As of 21 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2508.05928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.05928 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:10:24.550084Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T22:24:41.760120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T22:25:22.546587Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef0fa1f5-a63c-423d-94a5-c8b7a24b21d8 · outbound

This paper cites Reinforcement learning for reasoning in small llms: What works and what doesn’t.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Reinforcement learning for reasoning in small llms: What works and what doesn’t

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.461471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.461471Z digest=sha256:d1f3a878aacbf912d876c8706b80f826f464b6592130fc7144011223239e8a58

Observation 79ab36ea-c431-47e4-a763-4f9334673b60 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.465211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.465211Z digest=sha256:0fa536dd10082f2bc2002a7bc62bc88a1eaa0330062c890a58e93cd74504265f

Observation cec1ec34-ebe4-4254-89d9-788162f79dc8 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.469058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.469058Z digest=sha256:31e9b0b1c9dcfe20ae501bb715ef0b6b5b920221d0974e224371caf63a2a5ffd

Observation 7e1d3b67-3b55-40ea-a6d0-e840fcdd5b6c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Understanding R1-Zero-Like Training: A Critical Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.480610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.480610Z digest=sha256:979fb177d30008990b280cf5ffe56323d916f35f77a4e3cbc8efbe2c68fad872

Observation cb019d9a-d1dc-450b-bf12-b6e2b5bda1ca · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.484196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.484196Z digest=sha256:aaa692a943d024bfac4b4673b468b3d88b626d6e0ca7351187aca847ba99b207

Observation 66eb2392-1fc5-42a5-8587-b2612640d9de · outbound

This paper cites s1: Simple test-time scaling.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.491071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.491071Z digest=sha256:73582107a6c06b30cf511f349a37a6c56eacc8be5ff55f7deeac18a75cf7824a

Observation a284d844-e747-4df6-81cc-0122ea43d495 · outbound

This paper cites On Symmetric Losses for Robust Policy Optimization with Noisy Preferences.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting On Symmetric Losses for Robust Policy Optimization with Noisy Preferences

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.494910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.494910Z digest=sha256:061fd247efec592dd5907d1609e2a4156ca02d8faaa2425834f06ee91b505548

Observation 50c28dc4-55a5-4d1d-aa32-733ef375303d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Proximal Policy Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.498499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.498499Z digest=sha256:36411919897adb63e4c79cc1dba22b27de7a16ba75fe4471ac6d71906ffb9522

Observation 116f5781-46de-4288-863f-78557b93c382 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.505447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.505447Z digest=sha256:8966cf2b254a7edeaace421558621de3d9f291aef21f6aeb7d2c599fddc81735

Observation ceacfdd0-4f3e-4282-818a-3df7b6b59151 · outbound

This paper cites Long Is More Important Than Difficult for Training Reasoning Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Long Is More Important Than Difficult for Training Reasoning Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-05T23:10:24.698575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T23:10:24.509786Z digest=sha256:6f510fef4334d78a532447b3a8fe0c80d57251b3837c0590f2e674edaa2ce4af

Observation 8405bc8a-abd0-4739-b46f-0993e3b2409e · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.513775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.513775Z digest=sha256:0613faa1031ce3871eb5e635286db558a26f1f86d548e0b9b30a95ba8ecabc5e

Observation a27d6a7e-75b1-44b2-8bf6-8669645059d0 · outbound

This paper cites LLMs cannot find reasoning errors, but can correct them given the error location.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting LLMs cannot find reasoning errors, but can correct them given the error location

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.518161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.518161Z digest=sha256:a5b026208658c765b889d0feb6c65a814c5ab00f7736412a13cb24ba7d0cbe11

Observation 448730c8-885a-4819-9491-e48d9a58ada9 · outbound

This paper cites A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.525759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.525759Z digest=sha256:bc532a74efe42d4a297f8b23bba54251e481b7f4255d32c8e5eebeddc52ec50f

Observation 7cf9db51-0e08-48f4-90c8-a2c8f098b12a · outbound

This paper cites Bayesian Reward Models for LLM Alignment.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Bayesian Reward Models for LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.529784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.529784Z digest=sha256:37c14cee8d9bf7fbe7a2c27a3d4b18c29c4f0845b94dcadc4a98f8a53b108974

Observation 34e2174a-a21a-4f47-bbd6-334613c02e47 · outbound

This paper cites Are Reasoning Models More Prone to Hallucination?.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Are Reasoning Models More Prone to Hallucination?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.533640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.533640Z digest=sha256:8ff206c5c56e2a77af80910adb5c65a8c856ab144445ca3576b456a3cdb443aa

Observation cf79147b-2448-43ca-946e-05d996526292 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.537866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.537866Z digest=sha256:aae55ef850ae55b272a05f452b609790b3bb45200e6f9e95529c81895fae9d5b

Observation 681f7375-76b6-4f50-905d-f111fa72d75c · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.542099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.542099Z digest=sha256:8246fff08e0a1c77b9111238a3bfa7eb952f62d0d15d0906e92f4b10ccedd6d7

Observation ec5b716e-19b0-4c33-912d-32f541f0a34b · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.546164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.546164Z digest=sha256:3fd9e19ed1a06590defcbff2efd8b3416ca3b6d3e8187e9bd02dc98c9ec8e1d9

Observation 4acf7ad6-eda0-4af7-8c13-6b6d3fab07bd · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting TTRL: Test-Time Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.550084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.550084Z digest=sha256:b838a482d3cbc372af56a916fa0485a8617a7136bc494a94002fffb24544c4b2

Observation 892fbaba-8018-4613-907c-bbe0cdaaa6cc · outbound

This paper cites SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization

Reference 1988

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.452911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.452911Z digest=sha256:cafa39febb74d4441e80dca2173f4aa4691b5da9ccd8570497e296bcc5cbd27f

Observation 613c89b2-ed73-4a0b-a26a-bf4af8a9ffa5 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.522055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.522055Z digest=sha256:9a4b92c6ded54571da842a15c75ea65684c9a071601a18b68649e5ef428a01d0

Observation e6e576f3-303d-4653-9fa9-de1668690b16 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Spurious Rewards: Rethinking Training Signals in RLVR

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.501625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.501625Z digest=sha256:4164cf388c40a99f19f917946a4e679bf042bbdf6684d5da20b71d89f07370cd

Observation 01562e6f-0fbc-41bd-83a2-ea043bec836c · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.473186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.473186Z digest=sha256:5cfa217c39ece07365661198cb0c9f94e7e6432e7ece49cb17b8a791d18fbcb8

Observation c3f08595-d6ec-47d1-b174-aa944314cf68 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.477083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.477083Z digest=sha256:14218db52cc18cce8432a92b42db3274924ac79a177b1f2d35bbb840ad26d309

Observation 9c4e76e6-c273-4b8b-aa6d-237a889f0858 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.487632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.487632Z digest=sha256:2cd825782ef6599204ec6a7c01bc49bfce6e8b107d04584960be6c93436490e8

Observation f0893c3f-5b55-4cad-861a-355f9829a8df · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.456829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.456829Z digest=sha256:7ebea5635de3638a185fbee6749ff0f3ea6271acf8cc17a0c0f4b1bfa4e66b10

Pith citing papers

Observation 863b38eb-7988-4991-b28f-6a14df9dd436 · inbound

VIDEOP2R: Video Understanding from Perception to Reasoning cites this paper.

VIDEOP2R: Video Understanding from Perception to Reasoning Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:25:22.548980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T22:24:41.760120Z digest=sha256:a8344c70446eeb99b0203a99b38d39789cbf23e5897d802c68bbb03abd40f8f9

Observation 483e9cc9-cc00-47fe-a3bd-7314640d45a7 · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.814962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:ffe393db8b008d779c8d6875af425ed4bee5524a69fea28f0bcec84869513f0e

Observation e77d17ca-0e9c-42ca-9d00-ceef1daa6a5f · inbound

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization cites this paper.

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:30.130912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:31:42.065093Z digest=sha256:aa0c849f4faaf32a01713ddd98a5be208af38d22dfd4f5966166c88953dfc01a