Pith. sign in

Paper Citation Record · LEDGER

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models

As of 4 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2606.17890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17890 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:19:55.164835Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact31
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19af4c52-33ce-42b9-81b5-cde8d0740526 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T01:20:20.417002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:1b1627a0b6de0545d8ce4d88836d09532de12cb48f8699efc6857c229522ad4c

Observation dad79881-0cd5-4ff8-8ac1-4679e8b2ec71 · outbound

This paper cites Nature , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Nature , volume =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:be7876ad102768fe188dc0edbf828fdafa228db4930fa363fccc5ed0aa21dfda

Observation 4e825fd0-57ff-4e37-925d-a85ebce1b320 · outbound

This paper cites OpenAI o1 System Card.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models OpenAI o1 System Card

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.466847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:beda487b0ff230fdd23ee0e59287a116371762dfb7a1910ebbac7e32f1fd4455

Observation 1a824276-315b-4efa-821d-de3b518bc006 · outbound

This paper cites Qwen3 Technical Report.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Qwen3 Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:28:55.672673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:9be4cb4e3309d2c3caf73bc71d2ef79f102902efb8620ae7e949a42c09fdfbc1

Observation f026f964-c5be-4ae9-a483-802765abe360 · outbound

This paper cites 2025 , month =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models 2025 , month =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:d00eaa3fcbb099a967d8cdf7c5ae6167b7a84a05fa805136ab9ca8ffa2f7e540

Observation 8eb7da15-fae3-4903-99e1-a21da45fda28 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.424956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:195c9daf580d7dbb3106ec764f5997547e05c095d5362bb2070d58aa5c52c5eb

Observation 9cde0d36-7585-4f9a-a4ab-d2b522e4f35a · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.479970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:6a0fce4aff6a9dc29bb4eb67c2d2313d3382a224397f8d7ffaea37a4b496350e

Observation 3baafb09-be1c-4d2a-9faf-7491f3bf1ee3 · outbound

This paper cites Stop Overthinking:.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Stop Overthinking:

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:76684b274e5449872ca230982703bad59da187295c6701a4adba2a48ffe55296

Observation e1e0ba81-caca-453c-8cbb-93274f58aa9d · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T01:20:20.407741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:eca48d3676e7de1e76a715f44e1114d5dcaaa9a5d87bdb78978f041dbe1c6f72

Observation b962f982-3edd-4d66-9356-c796f94b344b · outbound

This paper cites Distilling System 2 into System 1.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Distilling System 2 into System 1

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.448324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:a5d9587d8a8ca9798b1e149577926424f0d33a8aa6d8c4c848144438df605af6

Observation 928c4a76-0e7a-4c4e-9943-a38a3637f8cd · outbound

This paper cites AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models AAAI-25, Sponsored by the Association for the Advancement of Artificial Intelligence, February 25 - March 4, 2025, Philadelphia, PA,

Reference 11

Resolution
verified exact
doi, observed 2026-06-27T01:20:20.419589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:315876378dce65544028b0ffaa873eddd428194069216480b6a0b37d65629387

Observation 7b2357ed-6a3c-403f-aeeb-0b8233a6b20c · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:108abe6754f2d3872c60a9af33195b2aa9db42773adcd54731770919ada2c151

Observation 5368e969-2474-456d-a6cc-5d5eab2a9724 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:810e99adb3dcdad234921ef0001c36a9d1c59b99b3fa8bc281735e470aef216a

Observation b64f57a9-5df5-45d7-8af4-5ffd85f7e1db · outbound

This paper cites Can Language Models Learn to Skip Steps? , booktitle =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Can Language Models Learn to Skip Steps? , booktitle =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:86a0b468750781d379a725bf919e5384dc6851b5c9ba86441dc923667833e3dc

Observation 051c7cd6-0507-4e99-b240-0c40e6847112 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.458337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:fcf965aa910125219332d808d2e7fcf36e29cdc11577793f12c8043e16300c81

Observation 66a84732-3d1d-47bf-abed-772c10a71ce1 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.477310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:f538c3067fc96eedd90afc4114537dc77f0bcdb7b11ee05201d7b4d162e122c4

Observation 10b2fc23-8611-434a-9926-6419f25d6a85 · outbound

This paper cites Second Conference on Language Modeling , year =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Second Conference on Language Modeling , year =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:23a7454a5b22cb915e9b2089835bd17ad9684723abc81ba5373f5e639bce9e4d

Observation fc63ee45-7a8f-497a-9280-54a85dfc5e07 · outbound

This paper cites Weston and Yuandong Tian , title =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Weston and Yuandong Tian , title =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:cc11c9a5909da20c961b26daa1fa751cab7022495b6e6481e5d16d3edec9cf64

Observation 3b74bfb3-6c72-490d-975d-c86454c660e4 · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:8738a47d604b2d56e47e01a68ad83200608e5cf479e4b6c4ac9045b87645fd3d

Observation 34908b32-4c0c-4a7a-a55f-972cf9fe749f · outbound

This paper cites Compressed Chain of Thought: Efficient Reasoning Through Dense Representations.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Compressed Chain of Thought: Efficient Reasoning Through Dense Representations

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.488804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:2b1c08af28b0b9f38ef2044a1a656f04530291eaf9749f4bf8c82be2b6f937e5

Observation 78ceec20-2394-4e8b-8ef6-c20aa9a7a53b · outbound

This paper cites s1: Simple test-time scaling.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models s1: Simple test-time scaling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.472173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:538f21484811c9484273c74cee6f25dac06553d484cfd1a7ba8de5a7c4cef0aa

Observation 8ba32e1b-27d5-401b-9136-70e8aa103061 · outbound

This paper cites Token-budget-aware LLM reasoning.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Token-budget-aware LLM reasoning

Reference 22

Resolution
verified exact
doi, observed 2026-06-27T01:20:20.440299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:caa0260db4a599dd4b7e3a4b4a2bc25c32a65ca283cb61ddfeb4aa492df4896c

Observation 94a043c8-4983-4d23-a67b-97e9b9be399f · outbound

This paper cites How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.433389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:37bd5837257128e35ecc7755a696d6ff79f7f71f187bd5b1cc008c6ab7f1f198

Observation 2208c267-3661-479a-91fa-1ec6c1a8c73c · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Reasoning Models Can Be Effective Without Thinking

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.410691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:4bb083244a7657ca13eed0ab90b1fad2267712cd558f7e4127e46ba76b919a8a

Observation 821f09f2-014f-4f52-908d-b527c50a65ca · outbound

This paper cites ICLR 2025 Workshop on Foundation Models in the Wild , year=.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models ICLR 2025 Workshop on Foundation Models in the Wild , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:98159f6711b58abf4cb693e02110ed16ccfe7bf83619c66bddbd26093fe8e5ef

Observation dd00c702-9405-42d4-9e44-32d499c17dd2 · outbound

This paper cites Seal: Steerable reasoning calibration of large language models for free.arXiv preprint arXiv:2504.07986, 2025a.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Seal: Steerable reasoning calibration of large language models for free.arXiv preprint arXiv:2504.07986, 2025a

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.442980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:cc049a3e7fef7cca9cd0b80ecb197a727c1c1bfa17cb7c00acb9557b49f36464

Observation c69b05d2-50f6-40a5-9a9a-44f704b3b9ca · outbound

This paper cites The Fourteenth International Conference on Learning Representations,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The Fourteenth International Conference on Learning Representations,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:a55d7a7b5e61ef67b8d263842203d5541aee8bf73986c0a51e9ff9322b73efd8

Observation bccb366f-1943-4c80-a936-05b7423a85bf · outbound

This paper cites 2025 , eprint=.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models 2025 , eprint=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:a18ae3f35af28dd5490d701b5bb23f71ca94c3dbba635c672813d6dda1d1331c

Observation 130a2e5d-78a2-4b78-8107-b402bbf2b606 · outbound

This paper cites The Fourteenth International Conference on Learning Representations,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The Fourteenth International Conference on Learning Representations,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:28844eb255173869d19512ff19ea14001223d8088613ec61e134c235025672c4

Observation b938dcb8-f066-4242-aec9-db628bc12bbf · outbound

This paper cites Answer convergence as a signal for early stopping in reasoning.arXiv preprint arXiv:2506.02536.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Answer convergence as a signal for early stopping in reasoning.arXiv preprint arXiv:2506.02536

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.491283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:ffa9b0e87f8ebf1025a1f2708cb2e05b2d3d5087e42be66036f1c25a80c2091e

Observation d83e4d72-44a5-41a2-96ab-9d5af7a230b3 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Advances in Neural Information Processing Systems , volume=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:ba838157a183721c10b6784233348d77c9475bc12993c4c0a99d74d2ec24d8ca

Observation 0effb874-2f0a-469e-b734-8f1a5cd3b589 · outbound

This paper cites Does Your Reasoning Model Implicitly Know When to Stop Thinking?.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Does Your Reasoning Model Implicitly Know When to Stop Thinking?

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-27T01:20:20.485987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:cfb6878e669bcb348daaaf1a4a5f161c46ec20f7478631d37e550a09bfff183d

Observation f5629f6e-05b2-4f83-9aeb-45cf0d3c743d · outbound

This paper cites Proxythinker: Test-time guidance through small visual reasoners.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proxythinker: Test-time guidance through small visual reasoners

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.436395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:3ff88a97f1e4483eb0380485a59f3f7e46240b053d62da65b53873dfb936a202

Observation 122cc56d-829d-458e-a32b-598604d67000 · outbound

This paper cites arXiv preprint arXiv:2505.16142 , year =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models arXiv preprint arXiv:2505.16142 , year =

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.482859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:431006a37ab88141f099780d9b041e61bc8f5408787e9deb160a6454f5563b8d

Observation 9d5f2695-ba0e-461e-abfe-036d6d9d4751 · outbound

This paper cites 2025 , eprint=.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models 2025 , eprint=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:1cd1152c15a541163e07cf8df4d36938447cd706dad76243662e08ff31ced630

Observation a8040bf1-5352-4add-ad47-f6a9a7db81ba · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T01:20:20.474640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:a6720071128129ff73e556b238e021712b1a6e5b5d15be4e1bfbfdbaeb5ae9f5

Observation 30463d4b-51c1-40e7-8dcb-95f70dc2b80d · outbound

This paper cites Forty-second International Conference on Machine Learning,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Forty-second International Conference on Machine Learning,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:0ab0e247f3f1436b5b71c3afa5606186e91ade41913fbf8ea027efd350b47fed

Observation 9f9b1fce-5308-449c-9279-b3faeb261aff · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.427621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:c7a1f41567ebbbb83176401314ab36f5ffcd2aedf56f4104062c1dfe80dd9b3d

Observation 87b9f984-2fe7-4fe4-800d-db18d699125a · outbound

This paper cites Bowman , booktitle =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Bowman , booktitle =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:fcc66a3e148cc35d507f8ac31fe09ae4e652242d34d9029f3a62aa51d0b75c09

Observation 64850e6b-02cc-4e95-9a14-53db735465d2 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The Thirteenth International Conference on Learning Representations,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:c854d24884f4785fe13797b714cfd62b838f914b116c248b51b39713ca601515

Observation 875ab860-dfd0-49c9-bfff-14c4c2a17e71 · outbound

This paper cites CatBoost: gradient boosting with categorical features support.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models CatBoost: gradient boosting with categorical features support

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:28:55.670228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:ea54f0ddef43aaececabeb69b4e108bb0840a4066331669c0998be0fb747241e

Observation 63cdcf28-7f01-4254-ac58-860ae92b80b5 · outbound

This paper cites Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.430752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:79a35e797814bcff390d1f09f4759131d8e33b3c4c74ba0c3198a9bed6d59ab9

Observation 9c75a31b-a1b5-4f76-93b5-bce026bddc92 · outbound

This paper cites Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.414019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:fa863fd639f20ef4abaf959ad78d1b3a543b5f09f1fdc36faf21ab5045770ac9

Observation 915c136d-ffa7-4926-99a0-0c210dab04e4 · outbound

This paper cites Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics, Barcelona, Spain, July 21-26, 2004 - Poster and Demonstration , publisher =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics, Barcelona, Spain, July 21-26, 2004 - Poster and Demonstration , publisher =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:db319f6f7c7a18f29708b09041ce26e4d15b0d0cfcdf6018c9ed634a8ac1249b

Observation d2fcfcb6-f5cd-4aae-803d-1a18f4f4f5dd · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T01:20:20.460749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:561414168a1fdc06a5930ba3b87bc5ac307409ee40e303ceb221c08f1b896da4

Observation f6b8bc6c-820a-491e-a546-4a2bf05e08f7 · outbound

This paper cites Answer convergence as a signal for early stopping in reasoning.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Answer convergence as a signal for early stopping in reasoning

Reference 46

Resolution
verified exact
doi, observed 2026-06-27T01:20:20.451214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:a4a8de76a18038a74dd1ea163f54f7abb87e47b89d6ab5451375b47bac1adc82

Observation fc34335f-3fb4-4612-9a7f-d1a3610b2ccf · outbound

This paper cites Stop when enough: Adaptive early-stopping for chain-of-thought reasoning.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Stop when enough: Adaptive early-stopping for chain-of-thought reasoning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.455211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:1219524a356a728cb334c2a8b9ce2fe0526c9f9001c738c508b8560f95b57db5

Observation 14627de6-5ed9-4683-92de-60bac275179f · outbound

This paper cites THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.464479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:e53f2fbf9595a35e5862114d0cf100688256b3280c0db81b4a000a9c77e8a878

Observation b6be4521-cc61-4f29-ab5d-14c4c3bc7aa0 · outbound

This paper cites Introducing LongCat-flash-thinking: A technical report.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Introducing LongCat-flash-thinking: A technical report

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.445767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:d1eaff2781ba6bc84fdb0dcde481defd7f290cbc65e92fa9bfcb003ef013d0fb

Observation 1a1ccc34-5b04-4df7-adab-e8c470ae2c6c · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Proceedings of the AAAI Conference on Artificial Intelligence , volume =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:3afb6116de35d4a9cbc8144c71b6fd44164edb74266450fca1dba9eb67782b60

Observation 68f461a4-51ca-4079-8372-3c5f2e54d8b7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Advances in Neural Information Processing Systems , volume =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:609b3534d784bd3fba9b1f76f16bd897c01a8f5fceb270971dcf55ad161ede39

Observation 3172a978-5950-4bcf-b0b8-41bff2d9b5b3 · outbound

This paper cites The Fourteenth International Conference on Learning Representations,.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The Fourteenth International Conference on Learning Representations,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:f71c2107422d6f9c470ce5e5be492f24adaa1bda069d88237f56cbfc47b53d06

Observation fd03b8c5-8032-43ac-9176-5da1f969216c · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.404474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:88190a7863f7154a9a539aa232ccf303672e29e27ac79c6d416cac1ef9f38492

Observation fa8043a2-74ba-4037-a898-d80c21ac9bb2 · outbound

This paper cites The evolution of thought: Tracking llm overthinking via reasoning dynamics analysis.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The evolution of thought: Tracking llm overthinking via reasoning dynamics analysis

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.394714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:d6d3dfcabf1a062183e2f4309b1feaa25dd87922833c24f14de7bebc0a305136

Observation 3844f17a-6ff0-4657-b60e-cb50109e83a7 · outbound

This paper cites arXiv preprint arXiv:2509.23024 , year=.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models arXiv preprint arXiv:2509.23024 , year=

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.379323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:87c6289a3e28403926612b92cac01877268a0b3543609e77ad6da7a10d6deede

Observation 2793d116-fdcf-4db7-b403-7eb8c5a445a0 · outbound

This paper cites In NeurIPS 2025 Workshop on Foundations of Reasoning in Language Models (FoRLM).

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models In NeurIPS 2025 Workshop on Foundations of Reasoning in Language Models (FoRLM)

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.388880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:8f2326519c8d6ba105f5880dbf342a2f9fdf3301bf5eacfbac895232c1afd40c

Observation eaa7aeff-abeb-4ed9-b411-a7e4a1cd06d9 · outbound

This paper cites The geometry of reasoning: Flowing logics in representation space.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models The geometry of reasoning: Flowing logics in representation space

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.391753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:25a9b13bd5f2290be89372b2133987d217ab85512615d02c720c58a6eb53a371

Observation 0f47999a-197d-4b56-abef-192d3bd6a512 · outbound

This paper cites Let's Verify Step by Step.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Let's Verify Step by Step

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:28:55.674815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:0665b7f3b7510b492b6b50c5d10b484c859c20b12a8ba3e3db944ab668c18fc4

Observation 0c66043d-ab02-4817-ab44-633d04a7ee4d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T01:20:20.384992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:88cc979a5708781790d0f8c7d5d4d933e19f2b6a0644bd082ea49935f33e69e5

Observation 85d85f9b-1ba0-4ff3-9f70-b9b0eb3345ad · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Advances in Neural Information Processing Systems , volume =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:1a5269e6bce69414f8c32b5df1bb3a9e843bcae475110d8b2fa69700a9b4988e

Observation 80c37040-d1ab-4db5-b4e1-35e7b39b70de · outbound

This paper cites an unresolved cited work.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:33691f73521e5d534c6536018232d192b57ab3280c23474c371a166afa13b8cb

Observation 722a1cb1-2076-40c3-a3f3-663b42a6ab7c · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Advances in Neural Information Processing Systems , volume =

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T01:19:55.164835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:ae1592960f94e078da48d2f9890be313190af63c749bc453755a3ba1c51b8def

Observation 7de60c0a-a784-43eb-8721-f0f0428df11d · outbound

This paper cites CoRR , volume =.

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models CoRR , volume =

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-27T01:20:20.397961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T01:19:55.164835Z digest=sha256:074fdf53903a3b25860a6ecbf10d2e6a5cee34645c3b02087e880a93b348146b

Pith citing papers

No inbound Pith citation observations are available.