Pith. sign in

Paper Citation Record · LEDGER

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.27888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.27888 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:39:13.949389Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d621b43-55bc-4517-9498-680bc761f185 · outbound

This paper cites InInternational Conference on Learning Representations, volume 2024, pages 21246– 21263,.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning InInternational Conference on Learning Representations, volume 2024, pages 21246– 21263,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.095503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.095503Z digest=sha256:6284da57c10cb204c6bb1287eab63b99518556e7c9960d25b2cb731e7bf82576

Observation edd6685a-cf3e-4f98-a13a-aa63b8682c38 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.337777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.337777Z digest=sha256:c4fe6379a627d29641efd1c8ea89c757faf5ec693053e64d415399bb16883d28

Observation 322fe685-36cd-4ad0-8c23-39f7df92f425 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Reinforcement Learning via Self-Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.428331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.428331Z digest=sha256:f2e2c7a94adaa985073b6368912c29c74539b0f15d98dfe662a46e1ae7c6f280

Observation 8f920095-11f5-4120-ae17-c4e1d2a8c330 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Entropy-Aware On-Policy Distillation of Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.506494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.506494Z digest=sha256:63895045038e23c879d1210d97d64a7ccb9fe461de31e96177868450473524ed

Observation 61fc87b8-3b93-4468-85a1-b746303ea32b · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.667302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.667302Z digest=sha256:e925169ca6be17e7595e916384cd22b00536f7282b4b9285b384bffae56a8af8

Observation c1b745d8-005e-434f-8ad3-6f3bbae76750 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.831350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.831350Z digest=sha256:7c7084c0256a8280a204e18d79a35ee2f7f1e56cb78717e737bc72b5cd55cae9

Observation 16e5ab63-285d-4a4f-9d20-c3d3a8f84e86 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.998170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.998170Z digest=sha256:ddc3bc1ee9a4729efbd7ffcc798e9da9ec154cb87234f01db4971855f008a391

Observation 183c51a9-b916-4a7d-9171-b06bf3c2cfdd · outbound

This paper cites Unifying group-relative and self-distillation policy optimization via sample routing.arXiv preprint arXiv:2604.02288, 2026a.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Unifying group-relative and self-distillation policy optimization via sample routing.arXiv preprint arXiv:2604.02288, 2026a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:12.155460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:12.155460Z digest=sha256:e4f8df674da12a95301f4e136217464c6c3a22caf6e04eb40e9517c6832aaa2b

Observation 8c4addac-5f3d-4248-b34d-7bf335645e1f · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:12.344922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:12.344922Z digest=sha256:de82219b00b22cb2f414b1b1c46f6054239ca11bfcf0c54a620ec1a26c19eebb

Observation eb090669-7c15-4964-ad5d-345e27283d08 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:12.628880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:12.628880Z digest=sha256:741eb3af28007e4aa0778c5446c5fd1dab9ee3132a852ab8eb9fceb806dc42d2

Observation 77ad272d-02cd-4f1c-93bf-3346b51c8861 · outbound

This paper cites Grpo-𝑙𝑎𝑚𝑏𝑑𝑎 : Credit assignment improves llm reasoning.arXiv preprint arXiv:2510.00194,.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Grpo-𝑙𝑎𝑚𝑏𝑑𝑎 : Credit assignment improves llm reasoning.arXiv preprint arXiv:2510.00194,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:12.988092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:12.988092Z digest=sha256:3daedac46086e5796d60e585e09626c2f51e9cf54852712c8fcf7c66b18feea1

Observation b1cba096-2ea0-40f9-8879-457da848da60 · outbound

This paper cites Privileged Information Distillation for Language Models.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Privileged Information Distillation for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.080946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.080946Z digest=sha256:46a3872118df700ab57bbb45b4984fb188b7803b4f9939e1ddb94ac9c97327ce

Observation 6b1bbf5c-41cc-4576-a408-91fe4a20ffee · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.177100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.177100Z digest=sha256:2e1b884cdbd821b2e1624f507e22532401fa9916d30b3586b80d5f3c0e769a18

Observation 31693278-5f59-40cb-9706-7d287d222e25 · outbound

This paper cites Gtpo and grpo-s: Token and sequence- level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349,.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Gtpo and grpo-s: Token and sequence- level reward shaping with policy entropy.arXiv preprint arXiv:2508.04349,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.289349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.289349Z digest=sha256:fc4eeddcb9a85ce3c2337c5866df4449b2407c1faf975e3b7d0800483508c759

Observation e44d51b0-8ae4-4ad0-9230-592978058a1f · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.429089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.429089Z digest=sha256:0f9ae0e176e8a662e216c1bcc2411c25e11ab86e1713069c8b2219fd301bc79b

Observation de69374a-2a35-4ff7-8f60-df09525962c5 · outbound

This paper cites Qwen3 Technical Report.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.580396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.580396Z digest=sha256:41fd8304a75d18c362eb0b39a03ad65780168a9d502b8eaa1abbe8a014856445

Observation ab64159d-5729-4a3d-acd0-d45d501f6fb3 · outbound

This paper cites Self-Distilled RLVR.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Self-Distilled RLVR

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.725332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.725332Z digest=sha256:4ffde18e7e03d40f1461277d0ef605f454167757700eb96d03b81a139aaef96e

Observation 0f42df2f-2808-4c07-af13-d728d899ca50 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.864937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.864937Z digest=sha256:7b0c15139c1e3dd548ac8e3944032190cb56274dd86ba8df1b80147247fb1bb0

Observation a5030ae9-4750-458a-b979-ebf7638713f3 · outbound

This paper cites We consider only the clipped policy-gradient term and omit the KL regularizer.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning We consider only the clipped policy-gradient term and omit the KL regularizer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:13.949389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:13.949389Z digest=sha256:4ab407f2d5065c39ab3910c9508c4fea4994cd491923fb268928b1b0672be1ca

Observation 118f39d8-b5a0-459f-9439-56551cfedf7d · outbound

This paper cites RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:12.802081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:12.802081Z digest=sha256:a73c182ef199302d8e2af8819087f8dfb658402861b587e5ac5c4745647a27ff

Observation ada1cf51-ab70-4f6c-a054-2c96ffe3a358 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.134603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.134603Z digest=sha256:0e8c40742e365d0b90636d3d2c59ba77abd93bfc84dab8bb88da25af48bbae4c

Observation 4c37452a-6277-4cc4-bf04-a4aa8b7de2ee · outbound

This paper cites an unresolved cited work.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-31T23:39:11.249027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:39:11.249027Z digest=sha256:ab408eddd083dec1f479dce10c336001ebca7f2760d0e9724df31ca285b54978

Pith citing papers

No inbound Pith citation observations are available.