Pith. sign in

Paper Citation Record · LEDGER

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2607.07976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07976 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T14:25:07.006233Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact11
  • verified fuzzy14
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f939aa2-c97f-4ece-a945-caa962951124 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.513959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:02bbf8f7118822eabc797ebc34a9652dfee3c7d0cb06d1b6fd38db1e05ffc5fd

Observation 62f0072d-42cc-49d5-8032-ab5669f3fbfa · outbound

This paper cites O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning O lympiad B ench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 3

Resolution
verified exact
doi, observed 2026-07-10T14:27:07.150173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:c6b61b4ca17a25dabed55f9ed750a6373100467fd62c37d37abe02d765c16835

Observation 889c4786-37d5-48d0-b766-416aac88aafa · outbound

This paper cites The curious case of neural text degeneration.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning The curious case of neural text degeneration

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.522253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:2bf96913bb36e3a6749318dab14cb41abd30901a6c6587ccb7a23d102db4a871

Observation 371fe2a7-ef04-4efb-918f-889151363546 · outbound

This paper cites Large language models cannot self-correct reasoning yet.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Large language models cannot self-correct reasoning yet

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.510160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:7d4adad993b68df3e4f48e3ae6f9c8a8560e436cfba41bd472d5a505c5dfef0d

Observation 5ca50a82-9f01-4344-85cb-904541018532 · outbound

This paper cites URLhttps://openreview.net/forum?id=IkmD3fKBPQ.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning URLhttps://openreview.net/forum?id=IkmD3fKBPQ

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.519445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:826e498766a4651b0707744b3c5a8637f769a166edd573886e0da83ecdc1e28f

Observation d9cacb44-7e90-4b73-813b-888370759ec5 · outbound

This paper cites Execution-grounded credit assignment for grpo in code generation.arXiv preprint arXiv:2603.16158, 2026.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Execution-grounded credit assignment for grpo in code generation.arXiv preprint arXiv:2603.16158, 2026

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.362244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:bb091c40d9445644ca2dc67695b10f857c9f0784522fe42359e5ab50d18e341a

Observation 8f841fa7-bbd9-4c3c-80e9-94cd255008cc · outbound

This paper cites Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.520180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:a33f6b38bb94b499c0e2c6c90330ff6ddf40bab77d5636b7f6af37017cbbde61

Observation c6b3cc15-1af5-43fc-a7ed-70a27c62fdc7 · outbound

This paper cites Let's Verify Step by Step.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Let's Verify Step by Step

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.380941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:2ec9bb20b0f6f36f81b675cb924802f7dcc5df13e49fca4ac428b502bc7f8f17

Observation 3c72e548-1c77-44d7-a5c6-9aee68ec78be · outbound

This paper cites Heterogeneous adaptive policy optimization: Tailoring optimization to every token’s nature,.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous adaptive policy optimization: Tailoring optimization to every token’s nature,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.514418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:8dc606dcdbaa30af4bc7b4daac337a7b37e776ab879ae64fd0975d8158ebfe36

Observation d853adc5-ed7e-40fa-8fcc-0e2ba6282484 · outbound

This paper cites Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's Nature

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.371938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:6b6138bd7336e7c1c799dacf11b9a709a4e6c93c251b748851f637fac1b7a236

Observation d5004d3a-3942-49af-8c16-a1ef2d2c52fc · outbound

This paper cites AMC23: American mathematics competitions 2023 test set.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning AMC23: American mathematics competitions 2023 test set

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.512049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:523b7c33fcfbdd38c4f772a559c73abebc1035a72e732c147bd5790cd425c7b6

Observation bf9449d8-30be-46c4-91ce-3ad63950abf8 · outbound

This paper cites Grpo- λ: Credit assignment improves llm reasoning, 2025.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Grpo- λ: Credit assignment improves llm reasoning, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.518110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:a824dcd0f85469a9753b79d57c53ceed7bbc4de94d1ddca7f0c78de7350a4a8d

Observation 43e0cafc-c7ab-40f5-b264-040dd8c5f203 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.384161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:6c871f2212fe084d60e556397e1b926d3d1e160b80c6700391e19f8b77452d68

Observation 6ab7f1fc-f10e-4e0c-a54e-7d8c23c246df · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.366684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:5699a910c382f6ae0265857e42b72b4345beebc53d7f79eda40bf9ef107e8035

Observation e66dfbd5-cc9b-494f-9305-edaa05cb4812 · outbound

This paper cites Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Gtpo and grpo-s: Token and sequence-level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.360085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:4317c393737de77899035ba7b620169bcb96f2b2f11f262704564aec7264e9a3

Observation c4b2a6d7-bbc5-4613-9c20-d6e97f25b4ac · outbound

This paper cites verl: V olcano engine reinforcement learning for llms.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning verl: V olcano engine reinforcement learning for llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.517642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:8f640255363b95759cbce65112124033fab02938fb0ad02036fa4f40c494662b

Observation a476cc92-eedb-4256-9b16-207b25f08939 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.499701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:4c732e0be915d2bce8f52235d8e134e67f01222722ab5ac17fb182cd50218ac4

Observation dd27a7bb-7cf1-493e-a2b6-7b234622c20e · outbound

This paper cites Neural text generation with unlikelihood training.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Neural text generation with unlikelihood training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.497765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:58e8711136574ebb02e6dcf5d0cc175525d439c2143e874ddb5350ec494a5c7a

Observation 8a490e27-cb6f-4625-a5fb-a8c44d8a8a0d · outbound

This paper cites Self-ensemble: Mitigating confidence distortion for large language models.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Self-ensemble: Mitigating confidence distortion for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.501760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:f6e894932d354214deb3290e045c41dbca1089ec6d705623396bc091645686f7

Observation 226c4206-eba7-4215-ba25-34ba9a82457d · outbound

This paper cites an unresolved cited work.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-07-10T14:27:07.516170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:85f9731702de043079ebce724a4a5d76ccbd8d3a032eb064c1e2738764d63039

Observation e54299af-ac5a-4999-9d10-86526f99c937 · outbound

This paper cites Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Learning at the Right Pace: Adaptive Data Scheduling Improves LLM Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.377916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:42ec29582779f95626e67e2bfc76254ce2d71fb6e8994f066e0c14ce00340153

Observation be331791-00b3-4a40-975e-9490a234aa7a · outbound

This paper cites Qwen3 Technical Report.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Qwen3 Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.365296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:ea6f6fcfbb90c53861f83e57f677b9bc33d0e6b23e4289ce17c5a1515776ee67

Observation 103c5305-bbf0-425a-91ae-708c593246de · outbound

This paper cites Int: Self-proposed interventions enable credit assignment in llm reasoning.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Int: Self-proposed interventions enable credit assignment in llm reasoning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-10T14:27:07.369395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:2259ca8badfc11b23dabf55aa56973043e3f1313d56270f0afa139cfa527a696

Observation de480b0e-a72d-4343-9e30-f44cbd760865 · outbound

This paper cites Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T14:27:07.374977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:67105e3b9727be4907e8d933c455d15322b3c6dd793af96231e1f4c07947864b

Observation 552159d3-a99d-4695-b361-2ce62ababa40 · outbound

This paper cites American invitational mathematics examination (AIME) 2024.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.524036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:4eb7a770cef48e8ea5ae51a914525aabcb4b38210cba7cc86b5a77de72e383da

Observation f0ae9229-216b-4b9c-8ece-02462de4d7f8 · outbound

This paper cites American invitational mathematics examination (AIME) 2025.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning American invitational mathematics examination (AIME) 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T14:27:07.512495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:edbac23d8e0a7a2bdb00a7b68fdaade6c8303c88fda2e166015cda6ff280c03c

Observation d144154f-0194-4292-88c3-3a36c45b58a6 · outbound

This paper cites Group Sequence Policy Optimization.

When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Group Sequence Policy Optimization

Reference 42

Resolution
malformed identifier
local_arxiv, observed 2026-07-10T14:27:07.148343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-10T14:25:07.006233Z digest=sha256:065b9da034aee1fdc1a59e1146a9d35e7d653bc1e311b054020ac16cc13d7108

Pith citing papers

No inbound Pith citation observations are available.