Pith. sign in

Paper Citation Record · LEDGER

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2505.15612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15612 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:04.508009Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:08.292468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.881659Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28e84fc4-fc33-4282-ba75-0814213730e5 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.467282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.467282Z digest=sha256:715530acce8795c7966f50587d2f831a63151a3b96957f5980a1ac5f86b5e2f7

Observation 1cdb23a8-2f9d-4a88-b874-010ff7f73628 · outbound

This paper cites Arora and A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Arora and A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.568317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.568317Z digest=sha256:5a211f9c4b7152ff1dd5e62fddc11f25b85c8f5d034761ecc0782ee5853fcd44

Observation eee27329-6959-4997-b431-f94e55efbdc8 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.667035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.667035Z digest=sha256:3d6cdf7788145bf2714ee80b303d8ddcef35e5a51df7a984f801e73bf85a7abd

Observation ad411e10-4157-4f26-81fe-82d48ecf5652 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.798415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.798415Z digest=sha256:79a0d03bf4ac5be79ea8302862d380aadb5da3a27372fa6d5705eefc65705129

Observation 95b32e92-0252-450b-8771-e9bd405a767c · outbound

This paper cites Gandhi, A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Gandhi, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.683827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:17:59.894964Z digest=sha256:b3c98e4b1f0aca034badd316fae77d675390d080c2f4d2b4fec37f350fb8ceb2

Observation 3d3cf30e-68a5-4ad6-8741-267b5a3249cb · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Training Large Language Models to Reason in a Continuous Latent Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.003763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.003763Z digest=sha256:e558c35f6babf66c29939626e68ecc5302a666261e54ebe5294670c2cf1eb530

Observation 7b40c05d-5a62-4179-9807-fb0ab96c2062 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.102621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.102621Z digest=sha256:615cdaa509770651012aedebc606d7d58521bdad8d76e5225e149f52b3ff84aa

Observation af46f05c-9118-4372-a39c-0345ca05e32a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.199909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.199909Z digest=sha256:0be1bd72e1ea6e0c2cb2f4adba67fc44e527b53b3ef0a49caa37adf68984d4ca

Observation b241c72e-f219-4882-a4c3-24328accca6a · outbound

This paper cites Hendrycks, C.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Hendrycks, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.412863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:00.275045Z digest=sha256:756bc76fb119ca8e1a5120ecc143e743a9a8ad48ea0bab32e9ab6451be846a0b

Observation b913ec90-d3c6-4440-b6c0-178707ad69a4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.357524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.357524Z digest=sha256:a5eed99d3343eab47bbdef90ed9a36e3e50b92d5906c15bff27290f7efe1a94d

Observation fd11b564-bd11-4335-97e6-3cf7d9d1af9d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.496831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.496831Z digest=sha256:8b5488519a952e97ab0c870a19a306d8e658da9f0eb9eac66fa6b0292468d08a

Observation 954cf2c1-8e1d-48fd-b239-df1e25e27ff8 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.995684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:00.595825Z digest=sha256:26f25c75d606fa9d1bd8791a852f3b703145cfb79b4afbedaf07eb8e2a9bb173

Observation 1d8b928e-392b-4957-aeb1-1c5fee32c0aa · outbound

This paper cites s1: Simple test-time scaling.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.725959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.725959Z digest=sha256:5245f4767a65b9b220bd6484778db793ad897c3af74d829de3304d05cf12c953

Observation 288839eb-69ee-4030-a74f-913083a2f4e8 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Self-Training Elicits Concise Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.832579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.832579Z digest=sha256:1eaf9f3db019b890129f699a8712a5726b593000e2ec544f1869db67d940714c

Observation d730ab99-18ba-4bbf-b24a-fe68527707d4 · outbound

This paper cites OpenAI o1 System Card.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.974743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.974743Z digest=sha256:351aa53e523ab81c4f6ef6e42a4367921382ca6d10ca84d0d8c1928ffad92f53

Observation 4c36ddf6-0e6d-4bf7-b3dd-4867468d26a6 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.079263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.079263Z digest=sha256:feb056854414624c5fe67b6c22373a3e4a626fb3cbd28e89b123a94da3fe2f10

Observation 5d08d321-e896-4279-95d7-dd0755e96a78 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.206010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.206010Z digest=sha256:b868ef2aa4714f84cf7eb38df60c2d9757d34f3a7e3dbf3cb1abf9e1bb9b617e

Observation b27b57e0-8994-4bc7-874c-23c3118b2407 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.330227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.330227Z digest=sha256:4429a963fd7ad9c33810942f70502664bd69bc143fa4bea21de875620ba8ac79

Observation fcb88eb2-64be-4b8b-a1f7-82935f1ad521 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.437191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.437191Z digest=sha256:caa8d643949d9ee9e8630af4df1144a1ad3ce770d6e9de61aaf0706e20bf1fa0

Observation 35381598-08c1-420a-bc5c-2d6a5cc66f97 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.525978Z digest=sha256:383e6bc3e3cabef8cff45c713e3d2cfa98696374aab6c2c6e36ce588ea488024

Observation fc3775df-960a-4b0a-baaf-26fe0d688ef5 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping HybridFlow: A Flexible and Efficient RLHF Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.639730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.639730Z digest=sha256:58a0c239c7ca8b2ba944a6bf13291f4b73b82b0ed8414e534e9a19ab3d486768

Observation 474c8158-6c00-4b95-a7e6-e9f2fa52867b · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.733588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.733588Z digest=sha256:273349f23669a3a80d7dc60723a2b227b0ce933f1ca7a870c2966e0791fd3b6f

Observation 2a519d24-8e85-4981-8899-c72fa7b103f2 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.594553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:01.847962Z digest=sha256:a0fcad941ba9c9ccdf14b329e9e8131b35cdcbfa6958d274ecabfc99db45b75a

Observation bcae8ac7-cbd8-45d8-8249-037a2be47198 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.984634Z digest=sha256:7befa55a23d1bf715c1888e94644a16946dc845f1164c68d8ca92fb0c27028e5

Observation 515b6d5c-4d55-4420-a9a5-dfc1e798df75 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.205324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:02.101816Z digest=sha256:7942c5592e6cb566e1f91fd266d5ba95315c906d114376a73d06df6991212c9d

Observation 20148de8-f039-471e-bae1-d6fbfbd23660 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:05.874359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:02.267130Z digest=sha256:e3973708f9d9586d60f13fa84cc82c48ed2b37dcd4578f06bfa68e5dc6fd1a52

Observation 41384598-14da-41dc-ae07-b941ab3cd98d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.954245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.954245Z digest=sha256:bd6b6721e2be9a0280ba19273a3caef05a7d8b8aa7c82929bc6adadf6f1dd3f5

Observation 5dfbb9aa-1fd5-4817-b7d3-6c317d6bb337 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:03.643174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:03.643174Z digest=sha256:e40771885f328dbb317de2fcaae42492521d00f1a821d56e7f16c945aedaf3eb

Observation a01d5355-f674-45c6-923f-75fd472e93ab · outbound

This paper cites <think>...</think>.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping <think>...</think>

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.676753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:04.240695Z digest=sha256:a058c8a8a29d800c1c08e673cec1f8127cd04f60658344a4ed4e5f20a577e7fa

Observation 577e7356-dde9-457a-becc-c5fcc438a479 · outbound

This paper cites Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.438832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:04.346749Z digest=sha256:09b50ec0bb0a2e2fb355efcd975e749bba75dda6d66c3a1ae196575092ec2d59

Observation 01a5be31-685e-4785-a785-45382d5c0368 · outbound

This paper cites Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.150909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:18:04.508009Z digest=sha256:8a631cc6e3d562f9760896d152deadd100739e5bf3ca716c9194913e7cd9f702

Observation e37737d0-0717-4936-91d9-e1195dd1ca07 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.183957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.183957Z digest=sha256:4fff5d69575b1e3981677befd94adbd5dc35780fab75a6a66c83c212d4c2c779

Observation a80ecce7-12be-4b1d-ab64-ae428b6d2493 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.380737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.380737Z digest=sha256:17c6bdaaec36cdbb090b28aa827dadfd1a9174d9209d1c9424aa27ddd955a54b

Pith citing papers

Observation 8215d2d0-315b-4add-9aa8-a54da7b0598d · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.440468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:cf912399379ec738c77e45b9090de27ec233473f95a479d2fb9c8944d9deb0b4

Observation e2806cf4-e1a3-4291-8d88-34263ede61be · inbound

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It cites this paper.

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:08.292468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:08.292468Z digest=sha256:5bca08f9430869704bfd5f938553adcf109aff46e14ac10ff3cbd0316842d3ac

Observation 1cb0dbfa-7cc7-4266-9804-628e6c489f8c · inbound

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning cites this paper.

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:37.167029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:14:37.167029Z digest=sha256:4c06fd02ea388521da63f0107f9da52ef547823981fe38c058bd40645f547091

Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:bd457ced1e0a8099a5da5899549b5b5721e473c2ecdd0b7cc6ba486d0735d9d5

Observation 57e59cd1-dfe9-492f-b79f-d81a8bceeacb · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:1cefc2e8507756352c1a93c39a90cd0a9bbe2d51d24c0e3b18125f79a677294c

Observation d1d3c71b-ce7b-40f6-9879-0ea2168f9d51 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.135373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:415e512d17e24d89cb27eeb98df61a9a4ac2209491a79be9b63f4370131099bf

Observation decffc5e-5dc7-49d2-9828-01f11c1e8ba2 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.969293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:190578e3352907e760aee615fb324a13e1cb0423d94dc1b190bdb23dc043112c

Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:36:00.027949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:3400d43d9b1ed0d8d0e394f9b3e08b06678f9a730e4f5cca118890027ca2e92d

Observation 1bd3322e-6d0f-47ed-aae6-2d7d2b22e475 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.370255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:dc5ebb15270992ebd3d04c9ed56816e105e9b2ceccab9fcfa7e030ae1a122edb

Observation 0b3bda08-724c-48d2-8257-5898f42dd122 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.254204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:bcacf97269227acf662afd66625cd8cc62b71c3a0ec08f1fafccb063d53af61f

Observation 0f8f4104-ab31-4fd6-a7e2-35d3b24f9e7b · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.776375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:ce0c5be0dea95ceb46759bbbc5b0129bc3f83bceef5b4fe31efa81450067739c

Observation 9473f81e-b1c1-4bbf-b0f4-d52050b385e9 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.161581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:5f06031488adcaa9743c6b6ef9d5824e5c3e8dcf584adca8bfabb23c6c47f066

Observation 5f010621-c287-4485-9c86-3830856bcd32 · inbound

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning cites this paper.

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.250975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T21:29:31.723326Z digest=sha256:cb101100ee56ac2d827beb1a443d2a59332dfe284f2b8c9fedaa9cda67d60de3

Observation c6993438-c84d-4363-b3cd-d682e4cf7a33 · inbound

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning cites this paper.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.426841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T22:27:30.783923Z digest=sha256:45de309a192c1e73ab080f03ee9bca71153898ec468c43ed13c2b8b3aef721ef

Observation d4422fa3-a96f-48cf-aaa0-9b986e39183d · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.421204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:5a897ee4a24848950224e657bd5a8ea485f4bfcd712c40bd0413b77620ccbceb

Observation 5c542a3d-e8a9-416e-a70b-6da1d6e9e7af · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:8179605c4d47ac0035fcb2b9d5fc258a5b3811d26c921f883e0f873a66238bbf