Pith. sign in

Paper Citation Record · LEDGER

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 58 inbound Pith citation observations for arXiv:2503.04697.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04697 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:19:22.140009Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:18:54.255096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact9
  • verified fuzzy3
  • unresolved1
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fa81d11-7996-46ed-8743-588e7aa90531 · outbound

This paper cites Let‘s sample step by step: Adaptive-consistency for efficient reasoning and coding with LLMs.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Let‘s sample step by step: Adaptive-consistency for efficient reasoning and coding with LLMs

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:19:22.324673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:d3d656304275235694281605efc2a907fdf9278fbc12a83b4d269f61f45c091d

Observation bb9e500a-512d-4bcd-90f5-e9a1159aa4cc · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.761.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning doi: 10.18653/v1/2023.emnlp-main.761

Reference 2

Resolution
verified exact
doi, observed 2026-05-18T00:19:22.166445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:1ddaa61406bfda3acaccad51403fc400ac280af9b47bddd62376a366d7750490

Observation 0175b9e3-d0be-475d-a0ba-0bfcab57c2ae · outbound

This paper cites Training language models to reason efficiently.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Training language models to reason efficiently

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.181628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:3a8a3b4657f210413085fcf6c5fe50b302107aa0a2dc5b67e38d9c1bd72c4b19

Observation 8457cd01-76f5-4779-8d85-06dfbc2b3392 · outbound

This paper cites Precise Length Control in Large Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Precise Length Control in Large Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.190910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:d93254921dcd3bfd9afc768573a4bf0d2ad42398ceac1eb6e053e5d300eb0650

Observation 693e5b07-7cca-48f0-b816-fb3b36404d09 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.196152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:637dde53cb2ccc9d70c6187b683fb634321a200ed33cc5ca8ef5bc36dff6fcfa

Observation 1dd3c76e-bd63-468f-9143-dfa452f4a92c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.201688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:22dc6f3fd000ee9f5fad35da059b2a0070b8a64cab354b9ac8fbdf85ff8c9ccf

Observation 906baf4e-c8d2-41ed-9100-66a0d22d47f7 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.206556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:daaaa9bed7a3272b2e7ddd2b427d8584d94cde9b228010b22f992ce343a6878d

Observation ca3a1670-4346-4d94-9258-7c191e6aa7e9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Measuring Massive Multitask Language Understanding

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.211767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T01:08:06.256034+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:c5b69d1bc8587f0c69f98a5401f3e1b0ae6311e09946dc553ce2c90d1bb1980f

Observation f679074e-85d3-4ca5-9804-e9bb25ce8be0 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Self-Refine: Iterative Refinement with Self-Feedback

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.216675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:24c90773a3781a1b6bdea73f6e3137e921c5747d6ffbc5fd89f38fa623c62b1c

Observation 970e4475-3336-4207-bfad-bb842a42f4d7 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:35:31.553327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:519f53eac9363d3936c6b870a15cf5a17119bc12c34fdc1f4a46aca3d79c58f4

Observation b03cabc5-fb79-45e5-b738-f67594039a59 · outbound

This paper cites Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand`es, and Tatsunori Hashimoto.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand`es, and Tatsunori Hashimoto

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:19:22.309637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:c1b588051987c0ee00e540fb1ebd7daa496a51c19cac313cc24654d8a9b50fac

Observation 6889fe3d-f329-40cc-ad02-9442b8617701 · outbound

This paper cites s1: Simple test-time scaling.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning s1: Simple test-time scaling

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:19:22.227539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:4737cd8c4e20ff6a1b526d52a711d69a7d6ece1000ebbea3172078a4f8f99652

Observation 51bca921-17ac-442e-a70b-b5b802bdf172 · outbound

This paper cites Qwen2.5 Technical Report.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Qwen2.5 Technical Report

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.234153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:1e055fe2338cac6b95edfe551d1fddc92b0101c123f3b83872b8ac1cd59bad38

Observation f6542abc-4698-49bf-908a-6df105a069b8 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.241234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:64f32b0c32fc8eea6faebb28db9ad634a18b792d80ea1a71333528298b98c4ac

Observation ca35c44e-b644-488a-abf7-f99c6e96115a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.247265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:3c3e265776b6e4b01fe2bc133e6d6bd32bffacf70bd5686089460fc21cef591f

Observation b57e7a6a-64d3-4f8a-b014-b0291c7ecc5c · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning A Long Way to Go: Investigating Length Correlations in RLHF

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:f2ea849b3a1bb7eb5e7f011813f759fddce976c744b22204996c3983ee141ee9

Observation b9360124-0598-4a90-8084-6954ed8f2422 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.259367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:3ea41a0d6472291f2d8a90a1ddfd0e29e88812c677552018db81322514e109f7

Observation 36eda330-1dec-4839-b241-50e9e4941b91 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.263660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:2f57f6983696544e14950a34535f009edbdf40427fce9cebffbce2fb638391a6

Observation c1ca1916-90d6-425e-9400-bf4524171c3a · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.271647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:63c7a898ad5152beb21542c80b4d273a5b7944d934220acf80aa98862edbbcba

Observation 293b0fbd-a5c3-4c8b-891c-47ef59a14796 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.277539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T01:08:13.648188+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:5dcb392498196fae0b0300d8556c6114edc721610c923068f0b1d029d595c93f

Observation 758d6a05-6046-44ac-83b5-af68e33c09c1 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:38:37.536547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:b7a57d497d950f04f8d1e7ba9bbf58c338b865029d61d19672f94d2b309ab8c5

Observation 2be1dd12-3a24-4ba3-bc55-51dda81dff4c · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.290253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:bd6eee2209f98d414cd0842ae47bf17bce1a524e3450896db417d23320af6ecf

Observation 12234fc8-e450-4c76-b010-59c7d10c6538 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T00:19:22.300226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T01:08:14.037057+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:5bf7c31e5131e437144a4a8b54738d500b3a8fbcbe95e7caa20939c49d2001fa

Observation 1d7e3f5f-15e0-41f1-bb05-ad519c6b1cbb · outbound

This paper cites Following Length Constraints in Instructions.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Following Length Constraints in Instructions

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.295347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:5c994812cad81e0c7c3793fd65c5e1f22cac9892b4be297d19de577dba12e34b

Observation 16882dc6-a0b9-46f2-89f8-3a2bbe4acf2b · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 25

Resolution
malformed identifier
local_arxiv, observed 2026-05-18T00:19:22.173928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:f9b68a96095f65bac2de192a3a7677dd4061bd56c24d78e07a91090079c7b15d

Observation 9958e3e8-5516-4bad-a013-495a4c099ce4 · outbound

This paper cites an unresolved cited work.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Unresolved cited work

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T00:19:22.313163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:944a87a3c4ff4bf7d5ca066c76486b2d58bc07b4775b291c831ca229b5ee7304

Observation 22f01b3f-12f3-42f8-b33d-5fc46c21dac5 · outbound

This paper cites an unresolved cited work.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-18T00:19:22.317336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:cad94cc3ba406c90bdfac75a392e4b10cf1f1bec7ab5e0187c01b83be42f7df9

Observation 4faff704-0b5c-4030-bbd5-fe6f709c50ef · outbound

This paper cites These results highlight the effectiveness and potential of LCPO to scale to even larger reasoning models.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning These results highlight the effectiveness and potential of LCPO to scale to even larger reasoning models

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-05-18T00:19:22.320897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:7f101e523438a40edfedfe304fa3fa0c87660a0096c60b7b6e71214563bebab1

Observation e9e9129a-40a4-46fb-b373-6b9e9ab6df1b · outbound

This paper cites Thus, the maximum real part is approximately 625.6.

L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning Thus, the maximum real part is approximately 625.6

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T00:19:22.304302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T00:19:22.140009Z digest=sha256:01fb666303189cd80939a08ef79a1a07088a3fecca35807305d5842683ad86e4

Pith citing papers

Observation 2d8dbdd4-e35d-448a-9846-9108b89328fc · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:694b86fa88062da90000de6028c056ee55a949e0b123f99ea0a7159cf4aad989

Observation 2f529aae-b5d1-4972-941d-275476032a54 · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:8a0c4720a041caf988b3c11770fb21d61f652f51199088ca77a66456db907ca1

Observation 1a32c312-b75e-4471-b2cd-b264538e1092 · inbound

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning cites this paper.

UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T11:02:41.335059Z digest=sha256:fca56bfc41f801c01fd186ae43653ff7213149056cf2962ead7588809e4c3423

Observation 817ac0f4-1ee8-47a8-befb-7b032f40701f · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 192

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.109400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:7cff5e910e4ab8f58512f1c5b701364dfbfa895e05a7bb3d30741194e6d87f05

Observation 18446a07-4e0b-493c-94a6-4ccc267d50cc · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:12:02.524975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:c31e88491014e9a03ba5ca17c949402a3d1d7851ab4aec299583ad9cc06003e2

Observation 1fd82ba4-ebcb-4b60-997f-7bf0b3e2a295 · inbound

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization cites this paper.

Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:31:53.193553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T22:26:52.748349Z digest=sha256:6c48d241f91f167954ee574615766ae1f8d5006eafad044cd2cfc004b67dabd7

Observation b2ec4e48-900c-46f7-9d15-d3368ee9c60b · inbound

Self-Aligned Reward: Towards Effective and Efficient Reasoners cites this paper.

Self-Aligned Reward: Towards Effective and Efficient Reasoners L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:31:44.558504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-18T18:27:23.076544Z digest=sha256:291dd36e5b0ea3117666431da783f5656086633ab0394c15b6a93628925f6996

Observation 1107291a-a3b1-42ef-968f-dd80ab2f007a · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:a53a29839ae828512e7121a0f08dd7286a38a6b3580b3ddfde346984e5b630bf

Observation 5aeae402-b595-4ac6-9205-079dbeecfc93 · inbound

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning cites this paper.

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T00:18:54.255096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:18:54.255096Z digest=sha256:fe7d034175e77cd6bb02264369826f38c271dd32e147fa51a6d29a6583ea5429

Observation ff693d59-b172-4e5b-a32a-fe37107681d4 · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:f71e8bad1ceb87a20555f5371798217c367ce05362330baf459aeed88cf634f9

Observation cd61e5c2-b858-4ead-b612-3acf27a2c984 · inbound

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning cites this paper.

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T15:49:26.750021Z digest=sha256:8806c621a65d25337ee9af7793c5930344d846b7f2da6ecbefe58beed1abbb09

Observation c6678ef3-6f58-47d7-84ec-448e66f9b3fe · inbound

Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking cites this paper.

Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T08:09:47.113503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:09:47.113503Z digest=sha256:2854144ba0d39402706d91e59c748b225aba7e5b823650bd886d452169c2dc97

Observation 6924bb75-9a4e-4f58-9547-652c19bc877e · inbound

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure cites this paper.

ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressure L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T05:43:52.654039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:43:52.654039Z digest=sha256:beb7abde7f77aa15cc37ab32060ef8bf47da81cd45c1f881013b69d64c6b6b9a

Observation 1323ee15-722b-4546-b5d3-f3b96a4aae2a · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.149056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.149056Z digest=sha256:bfc6461088a3b2325b45459d6ca0c2a971fe5af27fd4a6801a04f0c644550ad5

Observation 1644e321-7682-4755-bb86-b97c71eed2ad · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:42.897700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:42.897700Z digest=sha256:777e6a947e74433c7cdfdd07dd801efacaeb4cedac74c61e7c0cecb377e819a5

Observation 2f59d64a-cd9f-4bd5-ae3a-7e17876ea507 · inbound

CRISP: Compressed Reasoning via Iterative Self-Policy Distillation cites this paper.

CRISP: Compressed Reasoning via Iterative Self-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T15:43:43.892684Z digest=sha256:8daf38dad2698420f0a7e2974784ea087149441ec3952a4ef8a96cf81796ca6b

Observation 0c860a27-a842-400f-b339-ab2fce5648d3 · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:62fef890004085d5ea7493519b784bfbda30d66c836ab80c682b1b6bf165aef1

Observation d3c2db13-f82f-4ffc-8961-46c6d0502710 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:d35e91fbb2077b4710c77faf457e51556881a22d101cb5a89aa2a08b7cbc6cc8

Observation 39be36f2-4c63-4c89-ba81-63a78e240b29 · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:b2627eb103093aadca90dbcd562661df982ca97202e938bc1f9956e7335a8a90

Observation 40e3c3b6-7117-43ec-ba3a-ec6518019675 · inbound

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression cites this paper.

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T17:16:47.284564Z digest=sha256:bafd93207add5532c6ccea58b26cd266ce23d01a1331fbd0831d0de80a11868e

Observation ae0c2924-0e93-4178-9fae-b430b5453e30 · inbound

AI Achieves a Perfect LSAT Score cites this paper.

AI Achieves a Perfect LSAT Score L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:35:01.079284Z digest=sha256:5c36d069fe7a2d1f33e4a619782869f8cbb107cd289329d61c3828571adeabed

Observation 9e3c924e-b2de-40c9-b9df-3535548e64b9 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:40ce21c5c040238d6c788f70acdf4686fe7d8461793135e946b71f25b4f7a0f0

Observation fffc6d1a-6ad4-448b-88b6-d16ea9b7d2f2 · inbound

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning cites this paper.

NoisyCoconut: Counterfactual Consensus via Latent Space Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T00:51:40.815981Z digest=sha256:b8a3a76b51d4c0176be159f2cfb63dc8274d27656153726d2fb1c10d2dda0818

Observation 803a5ec3-ef33-4874-93f2-2af4b5f69543 · inbound

Reasoning Compression with Mixed-Policy Distillation cites this paper.

Reasoning Compression with Mixed-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T01:29:02.387691Z digest=sha256:56bd6eeba0240527e3b2bca2af3b0f55aae5d12ad1c08754f85974d762c17539

Observation df7b3897-72d3-407c-ae08-636e4bc3c5c5 · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:3632eeca651250297b3df9f16ec3a34382de246de8e713a87e58245cc2a60915

Observation aca0d2ea-0c44-49e1-b212-ddc0ad5e0407 · inbound

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning cites this paper.

Nice Fold or Hero Call: Learning Budget-Efficient Thinking for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T01:31:21.377700Z digest=sha256:b7be4be2e41160dbd37c1e32bcfc7d3f8f60bd5e232a13263ff69eee64e93ced

Observation 5a6d6e28-7551-4c74-8fea-92619d3c1a59 · inbound

STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes cites this paper.

STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T19:25:07.440630Z digest=sha256:7c72c7ac6407d0ec1fab162773ba131bbfa581fb6701468526c06a89392fa285

Observation 023a4cb0-f4b3-4b31-bb7f-d58fe5f2b03a · inbound

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization cites this paper.

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:19:22.325854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T19:24:05.375951Z digest=sha256:a22dfb34da4dd389fc14ec266b3073378026b8d8209a6604a8adf6ab4cadd4fd

Observation d78ba50b-a751-4e15-abc0-ebbc5afdbf05 · inbound

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning cites this paper.

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:43:05.918289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T06:40:06.103206Z digest=sha256:99b67de928407ba2d58e78590c8c0368d096e3d4ed434e76dee4cbd7af595c2b

Observation 312bb44a-2a3e-4037-8b2b-3602306f0a2f · inbound

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate cites this paper.

Mem-$\pi$: Adaptive Memory through Learning When and What to Generate L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:29:34.591072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T04:27:25.041652Z digest=sha256:3e7296e32e45e7b94d823748707c34bce757b2300383ee9f3974ee2455b40fd4

Observation f2b77133-3e36-4d22-a3b9-9d75b5f433a2 · inbound

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning cites this paper.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:34:41.019224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T06:33:36.846345Z digest=sha256:68a6bdd2670b98b1b80e7ec2a367e4e515a821a4ec4b935b52f8e5f4c9fa0ac4

Observation 696136e2-793a-455b-b747-29311ac43f46 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:08.252527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:5f0221b400190dc75534c775f99558ad18753334aff64556e77058131aa7d79d

Observation 0c77cd8e-4cbd-49a9-9019-d8dfe8a8da12 · inbound

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning cites this paper.

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-05T10:20:57.261652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-05T10:18:18.717871Z digest=sha256:d072b3b3a7d72b6d89dd130794aa34dc20c32297c03f6c9e98a1bbb3a1c784e9

Observation f264eb7c-f68a-442e-a821-c30051363312 · inbound

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models cites this paper.

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.390648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:18:10.508034Z digest=sha256:5ad888fd41f411b85d55e9ff44a37393b02c4d60dc1b5701b31bd5e37ec7c285

Observation f89ebc58-fb66-410a-a04a-f1c084358a2c · inbound

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning cites this paper.

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:33:59.247448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T21:29:31.723326Z digest=sha256:1a16b51d831c22b9e2c39698e11b952dfa34a0eed448c67804862728548ac26d

Observation 0fe43f3e-450b-47fe-9999-357dd4ae2610 · inbound

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains cites this paper.

Selective Latent Thinking: Adaptive Compression of LLM Reasoning Chains L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:43:59.086669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T21:42:08.610000Z digest=sha256:75407dae6f8a1d247eae70536dd65918ebdf08a20336d95261be6d396e00169d

Observation bcb92eb5-13f0-4a01-bddd-76825724e788 · inbound

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning cites this paper.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:32:44.430054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:27:30.783923Z digest=sha256:84f04604e054467321b663efbb9cac27a370fc7813abc74be33b1231bbde9578

Observation 1a962ff9-7b82-4e6a-acc6-3009f4b6c7ae · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 245

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:56:13.617303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:4ed9f580a2a3127d4a71fc622ac3fe9dd590d64d2f5d8cabe3a6894220b0d2ec

Observation 7ffc2a5e-d264-46fb-b4b3-cd022fbf7bae · inbound

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete cites this paper.

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:06:16.189252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T15:48:48.046003Z digest=sha256:38159ca0b9b6c8a47c8a7642e720bdb81262e62cf58ed092db3fa676925c610e

Observation 27c7c1cc-dd59-47ac-8d31-8150c27874e9 · inbound

Adaptive Latent Agentic Reasoning cites this paper.

Adaptive Latent Agentic Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:26:22.290730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T14:24:22.855486Z digest=sha256:16e26e8a17c32b4fa5f2f3e946809148dd5f201c76b70c0444d709eae31b2ec1

Observation da594ade-0083-4092-aa11-9d7cec5744ae · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 140

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:56:29.846576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:a1c2fa8ab8253b305c0017bf86c733aee91414aaaf853442e5fa61be18d5ca6e

Observation 6a17d518-577d-485f-93ec-a358ec9531a0 · inbound

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning cites this paper.

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:28.727043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T10:07:16.700499Z digest=sha256:ab5f2f9bf2b0d78e070b0ec2d25402b238a67ac0ba48e7879c990c56d270f593

Observation a9b2f759-910e-4fb0-a6a8-1ff9d0254e77 · inbound

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating cites this paper.

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:08.734755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T22:54:44.329613Z digest=sha256:f4ca7937d5ab38b5718cb2092fa741b77981297b94367b2aca37e576a752fed5

Observation ede45ea7-fef4-4441-b872-83bbe816f0e5 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:59:40.641042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:a555a5c2fbbd1171cc2220ea2833e8834f7620a4115a615c920332f24ce092f2

Observation fbabde79-d719-4723-9e30-c79d3e012503 · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T09:19:43.895953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:47b24c4ea351c0cd92125eae050353db37fee17d8369d69009775a306870d009

Observation e700ee3d-95c6-466e-b68e-4c00cc452bf3 · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:29:51.940709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T05:08:19.504454Z digest=sha256:330167e4e4eb248329aefddf80d3bef42e867f20db79e482b9312e3d51131dea

Observation 88044151-cb10-4bd7-91d0-dd242c5ba572 · inbound

Finding the Time to Think: Learning Planning Budgets in Real-Time RL cites this paper.

Finding the Time to Think: Learning Planning Budgets in Real-Time RL L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:34:34.913753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T09:26:31.944405Z digest=sha256:c6fb82aaaac64f59053150d1368922200e529771dad65d07f84c1a041d660469

Observation 4ea7314e-e216-4389-88cf-15489acf6310 · inbound

LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge cites this paper.

LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:55:43.672011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T03:07:28.653313Z digest=sha256:01505e873fdb2b13505604309b245cd2306e14273e96d66aafde2d9d81d6876d

Observation 3b3ce192-d2c2-4d9e-a1a7-9850e6fcaf07 · inbound

CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models cites this paper.

CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:26:58.146529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-07-02T13:22:27.566432Z digest=sha256:3ca48ea377e6f8da4a3312ad91338bd20650f3d5b240c112f0da55c264537c10

Observation fc8a6eaa-0aec-49cf-8452-22b4f1c2fea2 · inbound

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training cites this paper.

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T10:50:54.419477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T10:50:54.419477Z digest=sha256:4b668dd4e8440b3d4a9053d601f631f0dc8506f0a2bf0b31b6e5ec45bbd85865

Observation f0f61310-7bfc-4e5a-98d5-f0c3f2e26e29 · inbound

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning cites this paper.

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T02:35:53.844868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-09T02:29:12.366018Z digest=sha256:f230b56d3a629ecdb4425b458daf3f6114551b911dbf997ca7573c5ebd21f28b

Observation e2a4b3ad-323c-44bc-a3c7-dcde0069c807 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:b45aa6181d7ca53b13f8dc21a7f97b7812be796f808471300616894b64cf4947

Observation 4c5c03a6-9d1f-48da-8fd4-5589d1c98456 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.752086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.752086Z digest=sha256:8d9d1ca2fb1650d68927aed384587b6ef602508a833ef7fbd8b4449120230396

Observation 4ca6bead-6994-420f-a141-59d3a0c82426 · inbound

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning cites this paper.

Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T22:30:26.382580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:30:26.382580Z digest=sha256:6679d67807a2cc2e2fcdc76af5758a5f2299f52438eb78c7cf4a7e602f85655f

Observation 41008c75-52f7-41f4-80a7-895bdf913391 · inbound

Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment cites this paper.

Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:49:25.932836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:49:25.932836Z digest=sha256:8f91699f85a60ef60af47c080e12ce7a4a004209967343d7e31a7e4a87e25835

Observation 87a3a126-65b9-4054-bf34-fadf898fb3d6 · inbound

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning cites this paper.

LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T12:48:50.195619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T12:48:50.195619Z digest=sha256:2471d491f3c5a59ba8b7c429fdeb166d39b143c94ef62b37b1a50f3cb2933834

Observation d07c7f28-408c-4824-8cbd-92ddd48df36d · inbound

Masked Distillation: Internalizing the Chain-of-Thought in Language Models cites this paper.

Masked Distillation: Internalizing the Chain-of-Thought in Language Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T10:49:22.213737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:49:22.213737Z digest=sha256:5ee37e908f18e2a5c2c8f3f735131a212f7edebb7d0448d12451a76c4fc8ae5f

Observation f23f478a-906b-4dba-8968-13c4034c7211 · inbound

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR cites this paper.

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T00:28:58.772850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:28:58.772850Z digest=sha256:164e9794868786a8f96fb24cf52f39e72c7e17e561e1dcc668861e9aabf03d3f