Pith. sign in

Paper Citation Record · LEDGER

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 28 inbound Pith citation observations for arXiv:2506.05256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05256 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:31.536924Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:54:17.890896Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.297902Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2e65a2-a26b-454f-a04d-17ba24728abd · outbound

This paper cites Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, and Dong Yu

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.048003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.048003Z digest=sha256:13950957c32ae89fbde98789c9a1ec3cfffd087e99889ec6485afb88da393aaf

Observation 21756724-5271-4197-82a1-57ed4f1d5c21 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.149754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.149754Z digest=sha256:7b54ee2ed1fe2037bcbd1c13063d9a7ad52f2565f5ece8ee91dcb8e7388c107a

Observation d7fd6e85-6f0c-4ed1-88e5-0df5bdedf482 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.270494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.270494Z digest=sha256:abf0ed471a4a88c5e592cef7c1200c30d269286eac3d2c3416e10a6991c2500f

Observation c17b38e5-69a3-4707-bee7-f1d9912d501c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.467637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.467637Z digest=sha256:e069fa3fb5fe58ca814629a65a64884592f2e45568dc26798e88d872bb3f43a7

Observation 506a3c43-7cb5-4e58-9910-0f3a7236edaf · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.568819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.568819Z digest=sha256:58f1ae2506835e6330a588ca1c83f9963cb76f6a9f893e9d3733812ac0684c6b

Observation 6266d389-f09d-486d-90b0-605c95de981e · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.642928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.642928Z digest=sha256:d4ca55d397d0faeb4a6b76a6de8d7bd8edfa5d1557286305b6cd9a294f372417

Observation 3b9d6cb8-82a1-4aef-85bd-e8ee20e8affa · outbound

This paper cites s1: Simple test-time scaling.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning s1: Simple test-time scaling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.721513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.721513Z digest=sha256:3f2a1035c7be354ea9ad39b0fea3e5807c9eda66acb0829dceebe95a5b0accdf

Observation fc219677-9c0c-4254-ba77-9456d1e1314d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.799790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.799790Z digest=sha256:222ff8a6ac14408f55cea7b2a558d951c5604164ecf00f9457a3aefc52cf3b8e

Observation aa4c12c2-f0d5-4d45-a420-c80e7d78f85f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.901007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.901007Z digest=sha256:84b0caa67fc7634d159515f681fc207829f0c71a79a4674e04a0ab6e54779870

Observation 90f3d839-d2ed-47a8-b0f0-16b8717c79e3 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.991793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.991793Z digest=sha256:86f1715e3d01ed1ed14977b800fbc8e938c7fb49f67855379dbf6c891aab74b4

Observation 73b43e58-27e3-456e-afad-2512004f8b5c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.092526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.092526Z digest=sha256:6ca4ec1e6f475bdd795551a1cbb8bdd8dd3ae71597344ce53d0088c88ed8b18f

Observation 49c3822b-ec35-401e-a7dc-565487c5fc36 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.175370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.175370Z digest=sha256:0faa186d08eab90c6172ab4912541a6d3d8593397f5bb39171c06031be7495df

Observation 79876e20-bc1d-4a1e-be92-fa16de0f8d84 · outbound

This paper cites Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.296126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.296126Z digest=sha256:2614c72bee6831f177ed514309b5093e8c09f479fbe55ee6e22ee9f0872ebdd7

Observation f65d5de8-86c7-4087-8689-c8d1d2d00d61 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:31.403312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:31.403312Z digest=sha256:cf9ac7965b460ba19e1626af7e91d8a0c161265aca86b25570673bab5198a810

Observation 58daa4dc-cb8d-4f60-915a-c311ae2b4f9c · outbound

This paper cites Backtracking.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning Backtracking

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:32.170425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:28:31.536924Z digest=sha256:56fa264411aa0148fef10fe249784e817098a51c3b549e3b4ac97bf1ad0e9f6e

Observation 419cdef7-efd5-4303-9654-5891a7af65a2 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:30.367732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:30.367732Z digest=sha256:7241accf0ad0612f0d5e17a4268bef5aaefb9da1ac73cddbc7d6441feddb2297

Observation 99687c21-cddc-4122-ad45-a82abd54e362 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:29.980350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:29.980350Z digest=sha256:3b6443734687e68ff117b5f79d9e3b3b652a91d81d3b2271c31c81f72e205fd9

Pith citing papers

Observation c090e580-0bf6-4349-920f-77f28307ce44 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.890896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.890896Z digest=sha256:1bb00c57408e7bc3ea61eec3d30e34d5f83a99777bd025ef1c3655d74a1d070e

Observation 16630b78-1342-41b9-9ed6-dc4aacfcc365 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:54.120284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:54.120284Z digest=sha256:6b1c76f86371bf02a40cc94330575859376439b27721f24ed4d8c722753e926d

Observation 4e8c6e8f-b5e7-42ac-8320-38c74b5bfdf1 · inbound

Schoenfeld's Anatomy of Mathematical Reasoning by Language Models cites this paper.

Schoenfeld's Anatomy of Mathematical Reasoning by Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:41:15.359772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T20:39:34.858032Z digest=sha256:24c38e06ff54202e9b6377b84b7c50a8c86d5b836769e9a351b3626c4eb824b6

Observation 4200b067-bc3b-474c-96d4-a8b2e435ba8a · inbound

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation cites this paper.

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:02:10.268276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:02:10.268276Z digest=sha256:c86e82f6371f8409bbe5905b51daeb827606cc50fa2ba1815f32a036714bd0a8

Observation 39b95d97-b81b-4681-bedc-cdd7eba02034 · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T13:22:55.264368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:0b01cf953ecd432291119087ca65ea8455c46f88aaef265a8fd154020db7a4af

Observation 713381ab-65ec-40bf-a8ce-7c0fa3ade253 · inbound

On the Optimal Reasoning Length for RL-Trained Language Models cites this paper.

On the Optimal Reasoning Length for RL-Trained Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T02:53:07.791956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:53:07.791956Z digest=sha256:0cd1f99d939bb4a13ab5f295147d7aed07defa1b5df3f717486c1cea372ca8c8

Observation a4a96d6c-7bae-4897-91f2-4885d260985a · inbound

ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning cites this paper.

ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:47:10.890611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:44:18.723713Z digest=sha256:3c85b702bcc46147601dbcbf688a9097cd5f5913c6f2dda6e1da344ce101993d

Observation 886e77fa-05e3-4eae-8cca-de6899a4bc07 · inbound

CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning cites this paper.

CODA: Difficulty-Aware Compute Allocation for Adaptive Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:25:55.461203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:23:00.793443Z digest=sha256:cc05f3bb129b369f804665c0f0ca5844d6d3901f0568a862225085e256a2f283

Observation 84601f54-4850-4575-a90c-73210dcff0f1 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.152468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-07T10:42:27.644514Z digest=sha256:c58ec42e2f2c5488791ce90b0a6e7f5c572eed769ac51ce156a66d6659bbe29b

Observation 344b53f1-e67c-445a-b1a3-61e04d7501c6 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T15:28:13.138448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:28:13.138448Z digest=sha256:ea631e0f64dc7d0c13f2234f5e1049829b35e9a6c24d25f5c0146cb22c62d36b

Observation 724ea3d7-2db4-4f6c-b903-a7e256621619 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 227

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:04.968571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:94262ba7aa2664b94f71aacd73deca69ad8436ef33be9384a77cb15692a0e4ea

Observation 0933af8a-9e79-4afc-921b-197b5fea19cc · inbound

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards cites this paper.

DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:28.454063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:57:11.065744Z digest=sha256:a07465de4912d2c13a1b6103d11847b9e113f67f3eed6676d6da4f2553a7ea27

Observation df06833a-2720-4f97-a639-0fa242ebb1a6 · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.760115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:baf02322911b6d2d8387adc72c30e37e9d17a994f177a35484dd7b64bea480bf

Observation deecccc9-3022-4352-bddd-5df8ab3f24d5 · inbound

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning cites this paper.

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:43:05.955153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T06:40:06.103206Z digest=sha256:d250cb9be02f66d38b858a65941301a131bc523dd1d6ad27303ae2394d320de5

Observation 07878212-db95-4bed-acb5-058a6488d63c · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.348019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:0b75041622e8fc400bf5473dcc9a5d937eaddbe803a81e447b14321e578863d4

Observation b7ff19b3-3818-4609-aad2-73facda3bc15 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 243

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.590565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:60eeb2b4bae797457f715f88b75aa84e4a9a960c03fd4b5bb63fbb06ee0c5c91

Observation bae9feb9-53fe-44eb-96f0-cc4d667df1c3 · inbound

Libra: Efficient Resource Management for Agentic RL Post-Training cites this paper.

Libra: Efficient Resource Management for Agentic RL Post-Training Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:27.163309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:02:00.385932Z digest=sha256:68963b55f3c1fa90838ae936b2c265a6c57365f5ac2e1b9997048ae1db933c44

Observation 57c65ba3-1f45-48aa-acd3-0038a3bd3b17 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:56.069868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:37875da9be8145add4feb85c2c8a0a7108ed5042d96e971cd0c672e8042ec428

Observation 14dfd15c-1b8c-4bd1-bbbb-bea699947083 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:09:36.866760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T16:15:22.543601Z digest=sha256:cf545e8f91162d01302ca5186c162022cdd96b16aad5cb1194ee6778f8c0c744

Observation 3e778285-b681-44bd-9652-4e0890b5ae65 · inbound

Token-Operations-Oriented Inference Optimization Techniques for Large Models cites this paper.

Token-Operations-Oriented Inference Optimization Techniques for Large Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:07.680058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:07.680058Z digest=sha256:d9f3c8ff2d21c7cecdb9c568cf9b6eb339d80c92e963068dc1fec33c5a69b970

Observation a8dac336-9aea-4217-932d-7a2c3366b668 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 231

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.698593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:3923f2053a40f2a87f471888c48e85f7be52080b657e1f5546656eb32293b5e4

Observation 124265fb-a462-4763-ba7b-0bd2a6e40fdf · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.879965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:5ebdafa23cebab8629819f7370984fa0c765756a3f4b5ba428ede5835b921e2b

Observation 3742ff11-eeed-4114-8244-565d834dc2c9 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 268

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T18:40:03.299162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:289d15993bd0937572e556233efa57c6dcac84063192527787734b0da1948a2d

Observation 59b6edb1-7561-488e-8233-1e658a261cf1 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 268

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:15:59.049118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:2e00de63e911e1637e7fb54c540a4f70332c1eb38f7f36422f80c68fa0a257b4

Observation 25ff8107-92f1-4eb6-a063-61c8481c02df · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:e40aabcb6d44d97109fc7a341e9bb61f74e875e4fcfbb13be3e8360dbf5aa452

Observation 72a9e867-bdbe-411f-8fa7-12cecb22928b · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.856964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.856964Z digest=sha256:2c1c2271e58307a2df57616c7548508fa153bd52ba8b2e6ee7e141c074061c88

Observation 1c6f5859-2199-4f0b-a129-13d2f68e69a2 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.198277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.198277Z digest=sha256:d813eb2de247e80b674726dd0ab81f54af6e6b69df1cbced6d54d11ef925bfa8

Observation 22d43118-764f-4480-89c7-77bb223199be · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 237

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:11.853588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:11.853588Z digest=sha256:ac2491c4aa44d205fa9c9d9235e1a4dd167be5ee4ce871777e240490c5fe7085