Pith. sign in

Paper Citation Record · LEDGER

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2505.15612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15612 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:08.292468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.881659Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8215d2d0-315b-4add-9aa8-a54da7b0598d · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.440468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:e8d90bfe446ec5e1b5909549077534f47765d0c3fc56e53e91e9f06b56b18d35

Observation e2806cf4-e1a3-4291-8d88-34263ede61be · inbound

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It cites this paper.

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:08.292468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:08.292468Z digest=sha256:e1aa4103a99e0749d8f75e6465ad52622ccd4c9874da4360d9cee61e5e999099

Observation 1cb0dbfa-7cc7-4266-9804-628e6c489f8c · inbound

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning cites this paper.

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:37.167029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:14:37.167029Z digest=sha256:c9078b968ed9f53377e5fa8e127dda52beece4382014c78232d37c099a39ead6

Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:4530c762527a6fca9ba7c18a19ad2c73d08f0b4c84ffb0159eb8145f776fab7b

Observation 57e59cd1-dfe9-492f-b79f-d81a8bceeacb · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:7ef9b911bd589cc4f0b4e48a08fffc90cac2c5e56b3264b53f2c5821ad4353c1

Observation d1d3c71b-ce7b-40f6-9879-0ea2168f9d51 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.135373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:1b43dd38bd91dba10d89d7365a9e47ed93acbbfb13de23efda2760332f78f5ee

Observation decffc5e-5dc7-49d2-9828-01f11c1e8ba2 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.969293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:aa433f764f1908057928877234b76f23f706d89a688b97ed256ac4458f154e15

Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:36:00.027949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:32104c924356e748511bb41dafb4078d7a0db05d3fdf004b802242131f8341c4

Observation 1bd3322e-6d0f-47ed-aae6-2d7d2b22e475 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.370255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:7ec9649cf339e2a982948feb52af78c5f57ef4ad8916187278dae1bc4bbaef56

Observation 0b3bda08-724c-48d2-8257-5898f42dd122 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.254204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:46eb3640b8077d13330fa91737ed3d90149fdad71ed18e5368c0f0b7b730f36c

Observation 0f8f4104-ab31-4fd6-a7e2-35d3b24f9e7b · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.776375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:139225f68a57597e912d8ea6753ad9c95a31e766c6317fb99e82d08674061ba4

Observation 9473f81e-b1c1-4bbf-b0f4-d52050b385e9 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.161581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:23a4e9126a79892ab78c18bb7d8797a3f9af5fb6e09607d79195c382f7499874

Observation 5f010621-c287-4485-9c86-3830856bcd32 · inbound

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning cites this paper.

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.250975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T21:29:31.723326Z digest=sha256:7a6b10c13a37d7bf50938c584d33c00b595601d519da8a00c2d4a5bf9cfd70cf

Observation c6993438-c84d-4363-b3cd-d682e4cf7a33 · inbound

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning cites this paper.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.426841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:27:30.783923Z digest=sha256:738123f5865fb75ee7cfeff28a377f96531b9b23657c830ec02764cf014cb41b

Observation d4422fa3-a96f-48cf-aaa0-9b986e39183d · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.421204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:f1dc60e70b3d7fb9fa6216a7125fbf55a5b1b246ce20b9dac95119835ae8cb5c

Observation 5c542a3d-e8a9-416e-a70b-6da1d6e9e7af · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:382df3280e31e4a2e8e7f708fcdca2b009396d87850c25a1366016d3f8e8fb9c