Pith. sign in

Paper Citation Record · LEDGER

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2509.04027.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04027 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:31:37.980968Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:02:25.420148Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 653d22f5-53a2-4c5d-bd2d-c921f972abc5 · outbound

This paper cites write newline.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.734744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.734744Z digest=sha256:af25d1ddcb54d16d603e63742c3ba518ad594c689987f9ec5124c483eeebf1b6

Observation e0ece3d5-b6d2-4ecf-b2e1-086f710063dc · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.745578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.745578Z digest=sha256:a997c05fd306f73e912547c6033d6512942666503fd8f0d8f2410003a9d28876

Observation 3a99df52-5d06-4985-9c17-1f152740cc5b · outbound

This paper cites Language models are few-shot learners.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.751394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.751394Z digest=sha256:d7d1fbf2e69c22c393062ac0c4921bfe2ca1a84711e15af74389c139f193fc09

Observation 7eadecf7-832c-4bab-b6de-39a6d38e2ae3 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.755360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.755360Z digest=sha256:0c14df833e01693d47dd09954e2af599e971088f646f98f61d3327806d3ac10e

Observation 1d8dd784-010b-452d-ac80-2bda6dbdfb2a · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.759551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.759551Z digest=sha256:58d305b84a7f525e9b669f2881433d6bbd73d34c5f93cee3e8d69d15604ca7cd

Observation 0f1e3433-c858-49cf-852f-f44ed4fbd56d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.763822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.763822Z digest=sha256:f8f1323b689cef35cddce734df61cea77064f1e0f2b4795a83807f79b5db5d45

Observation 0469f239-9c70-4708-bb4d-0844b17ddcdc · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.768029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.768029Z digest=sha256:24ba0b117d53e14829d7eed7fd5d8bbdead4111bc6722330864677e867be2a75

Observation 022a69a8-5b54-430b-bcf8-f3e022ec8531 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.772049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.772049Z digest=sha256:83f400a25eed5337c0171a6709a55931899e9b53cb89c3ed6d3fe70012f76597

Observation 8b0b2129-bb5a-4ea7-92e5-788a2ffe8717 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.775535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.775535Z digest=sha256:9de12decd8f320bba72fb893a9cbc7c4be52ccdc6789584916c01f250280613e

Observation fa980435-17e7-40da-a5ea-0905e43f7941 · outbound

This paper cites Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.779402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.779402Z digest=sha256:6468e28eaffbe907df53019cb159b7196c574c6087dda37b2649b387a2453004

Observation 9bec0efd-2831-4438-9c99-49d7773451f0 · outbound

This paper cites On distances in uniformly random networks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On distances in uniformly random networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.536573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.783759Z digest=sha256:464b9bda8f86458d4a3d401fdfe8b70af47c0eed2799457c7f37bacf92abc4b0

Observation 758aa007-4123-474a-a24a-2b735e7b363f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.787916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.787916Z digest=sha256:9bd045a6d77bd24eabdd2c023d81f7f2b542b4c898d8e59675e43b7d643d9998

Observation 6dd9e2e9-cabd-40dd-9a55-44c83f72e0ef · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.792123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.792123Z digest=sha256:c805f6d383f3b87c26dd352f926e8ba120e9b275396ccca108ce9da754cc9565

Observation 295b6ad4-dade-4ed1-8e1a-2143e6727bdb · outbound

This paper cites Survey of hallucination in natural language generation.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Survey of hallucination in natural language generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.525058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.796260Z digest=sha256:7de9d0c1d309be50bc26c911a26eaf90fba15119f81a87ad6ae56dc383e67e87

Observation 0a8439d8-5404-4d31-926f-f181fcc433c4 · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.800280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.800280Z digest=sha256:141b6944b68218a65b551db9bc3009f9f4f90fdae5399e7de2d8126e64befd18

Observation ada20603-706a-4c55-be5c-3159391cb56d · outbound

This paper cites On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.804582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.804582Z digest=sha256:968730eff40d1f6cf936b71bf7bafc3e699319227f0eb7d635332169755c4fb6

Observation 0270234c-1444-4e64-a0a9-8a94c501a99b · outbound

This paper cites Decoupled Weight Decay Regularization.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Decoupled Weight Decay Regularization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.808968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.808968Z digest=sha256:3f51d7be9f66c6457392ae90198e367fd5161534188cbf3877729d22b5a44ec6

Observation cc6366ea-b12b-4c5c-807e-a77ef077744a · outbound

This paper cites Some pac-bayesian theorems.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Some pac-bayesian theorems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.514186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.813305Z digest=sha256:fea5e0b9c927c64f6041af18cb8bb7ccd6b86a48c1b53a9413b8bb10eddb6f3d

Observation a7d29796-62ed-4555-b21d-81aadb7bb521 · outbound

This paper cites The Llama 3 Herd of Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.819065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.819065Z digest=sha256:aae212a587b614ebda010e7f195abe2798d6cc43a29016c879ac3f8085f851fa

Observation 99b43959-d9de-4c31-8bb8-631a3602bb7c · outbound

This paper cites Nearest neighbor distance in three-dimensional space.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Nearest neighbor distance in three-dimensional space

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.502802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.823520Z digest=sha256:e3253b219a0b92eb1812bb2699072402199abadb1c62d7300caec9d023a8c0e9

Observation 18994fd2-58ee-43b0-8eaf-e72069020449 · outbound

This paper cites Learning to reason with llms, 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Learning to reason with llms, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.490571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.827688Z digest=sha256:a0fbd821aaeec4a802251fe88a56479e95bce138ceca2869e511b098b77a1553

Observation a947c17b-463c-415e-bf3c-2bae4b76115f · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Introducing openai o3 and o4-mini, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.477629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.831809Z digest=sha256:1043cb4f713bfb60e181260aa74078bd864fd6a7e381565774afbb55375f94fa

Observation a54cd083-fb2f-4c11-8b4a-c49510dbb830 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.464167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.836030Z digest=sha256:3002e8cefcd532f6b3b010a7e879884d538cc20bccb5cb2634eab13e28733ab2

Observation 6e0006be-46fd-4878-b922-e522d5a14d00 · outbound

This paper cites Qwen2.5 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen2.5 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.840224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.840224Z digest=sha256:9a2aa0e676df0e41d6133a45b44dd6d5c08e9ee40915c1c700c9e995cf7b09d3

Observation f30b8c27-7c1d-4049-a9f9-9ec2465ce29f · outbound

This paper cites Qwen3 Technical Report.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.844935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.844935Z digest=sha256:ff488ad7ed058b1dea1e0cc3e121affc6803a06d3bf596a1d3554d6517d956fc

Observation ed7d4204-0de5-4b55-ae19-7960911b9773 · outbound

This paper cites Benchmarking prompt sensitivity in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Benchmarking prompt sensitivity in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.451294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.849362Z digest=sha256:20d0f9ada4f83f40adcd35704a7944f4d9f2155a77c061b7ab8bb22c1aad6d20

Observation 2571ea31-6b53-4441-9afc-b0e244fadcb5 · outbound

This paper cites How much does your data exploration overfit? controlling bias via information usage.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought How much does your data exploration overfit? controlling bias via information usage

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.438143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.854081Z digest=sha256:ec148fab689fc12290df93b5c226fd5b86045920327e8523cc9a91d2440e51e0

Observation e3cad880-d189-4e67-bcf5-def90fd559b8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.858445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.858445Z digest=sha256:c582d0330711f7bcd69d4c3ff1e00a74713f9ef15c591faf5159269e129055c3

Observation c3121a39-00b2-46a5-a1ab-654f921f914f · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.862981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.862981Z digest=sha256:6805b4cb68d34e07d1a04ed5286443c7669e2e09f8c49545469ef869577f10ad

Observation 70d53b55-92e9-418f-8b2e-2d0617b2bff0 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Spurious Rewards: Rethinking Training Signals in RLVR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.869115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.869115Z digest=sha256:ea7f473d1b7d464c5fc4b59b73cefa70ade3a0e26b881881c0987e8ca04e5d0e

Observation ddd1d032-1732-4740-a8a4-bf1983309fac · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.873877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.873877Z digest=sha256:5a7c1e4767e2ad03c55f1c8f53603113fde4f3d1ae54440493016370c7b676b6

Observation a67f407a-8b36-4419-adf6-781640bcc753 · outbound

This paper cites A bayesian perspective on generalization and stochastic gradient descent.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought A bayesian perspective on generalization and stochastic gradient descent

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.879593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.879593Z digest=sha256:3aeec7141b2f558debadfbf3ea83f4fd1a83bac405e138f9b9fad0aed00b521f

Observation 21dc2f48-eff6-498b-8b74-69cc9287bd3b · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.884342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.884342Z digest=sha256:51ab1a4ac86220da495101d2bbb1833d8e6cd868287dd540f8ba8ed6cca31ee5

Observation b671102f-5d38-4cf1-801e-336e7c3bd280 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.889298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.889298Z digest=sha256:3bd131a4e609de5a286039dfa7d5f855d7769ad525458e2ab49fe825508c3f97

Observation b5bd238d-19cf-4f4f-afa0-29c1ecc4ad53 · outbound

This paper cites Understanding Chain-of-Thought in LLMs through Information Theory.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Understanding Chain-of-Thought in LLMs through Information Theory

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.894683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.894683Z digest=sha256:9d194347b78f92dd109bd7a68e343758ab382f4d347354af75b3b0c9cca73935

Observation fcb68915-e810-4f07-9ce4-755df89ed3d5 · outbound

This paper cites Alphazero-like tree-search can guide large language model decoding and training.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Alphazero-like tree-search can guide large language model decoding and training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.418602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.899762Z digest=sha256:7399c4f300472c37b7004e654477df22abcabb451fd34e928a274e6af5ccd1e5

Observation f9dfb13e-8c84-4749-abd2-807892557a2e · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.904517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.904517Z digest=sha256:7e40a232048472e71a51e765b21fa85e78c16f621dcc719d4093872e01895c66

Observation 4f90a21e-332b-47fe-8db8-77670c608173 · outbound

This paper cites Emergent Abilities of Large Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Emergent Abilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.909588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.909588Z digest=sha256:1ffbfff0aa24baf6d18d30cf85591b11c5ca4fe30c25b7c2f9778f263d0dd6f8

Observation 5f0d9a22-a300-49fa-b531-53970135c7ae · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Chain-of-thought prompting elicits reasoning in large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.914448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.914448Z digest=sha256:45486a00fd749971d1dc6ba43c144f284cd6e69f7f0dff23ee8dc7af1072ff07

Observation d155c5de-8c81-4195-a9f0-170c5792b7e6 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.919359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.919359Z digest=sha256:31fc6131ce34dc9d463116fdd7f806620c920a1129be75941d58ef13da647f26

Observation 295e2ff3-2aef-472a-9343-23a1c135d222 · outbound

This paper cites Information-theoretic analysis of generalization capability of learning algorithms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Information-theoretic analysis of generalization capability of learning algorithms

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.398678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.924011Z digest=sha256:0c149dc6d95ef45501e8dc5299392f5a6ea95a123e5ad6e35e32aad1ae017d87

Observation f010506b-53ab-4da0-8a04-855bf8c5800b · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Tree of thoughts: Deliberate problem solving with large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.928512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.928512Z digest=sha256:4f573364fe60105979dbce95f7d9af7a85242c672b89b477e5acded11cfba9e9

Observation 25fff4de-bd51-413e-97d0-775fd7de3d63 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.933115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.933115Z digest=sha256:f20a11e9cf8e652eabc17d210acd46bb30d1923703be22fb239e5b0a1e509746

Observation 5c9667f0-812e-4e67-8d41-0e8f7003c117 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.938067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.938067Z digest=sha256:7436d19125d794b5d49113fcb27496a58f60637668503d125d35270cf4444c36

Observation c7f1001c-2bab-4f60-8465-b46f3d0069de · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.942705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.942705Z digest=sha256:17013ed5785e06899924645ad276f2b34c78ea15ab6c5b4c4633dbc43c7fe9e4

Observation 97bdd282-cc34-4964-b295-0a99b8da6273 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.947268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.947268Z digest=sha256:37446882a3f53f1cd5adb15065e3680ecd146f53fd2375fa9c0cac1703c6c061

Observation 1d10c2c0-d047-4c30-b3ad-a1a73d9121db · outbound

This paper cites Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.951774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.951774Z digest=sha256:5f5c46a5e8833825dc979f2311721f0bb09b56058ab8c7a9752485c9bc9c6fd8

Observation 6174444d-a37b-45bc-95a6-2838c75b832c · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Rest-mcts*: Llm self-training via process reward guided tree search

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.956608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.956608Z digest=sha256:542462ea3d796051561c4c451f33f805e2caf9ca83ba846ca40bf6fee3aa4072

Observation cbe4cbfe-156e-46e5-874d-4a5c2542018f · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.961065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.961065Z digest=sha256:79a9f4b3a2d56a62967e50b7974dbfebbc1727d8f8b30d3ec24159d43b082509

Observation 238330f5-2f6a-4340-aa8c-b4321bac6f44 · outbound

This paper cites Prosa: Assessing and understanding the prompt sensitivity of llms.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Prosa: Assessing and understanding the prompt sensitivity of llms

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:31:38.367597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-05T10:31:37.966101Z digest=sha256:b1b413f96e32828048ffe0534d34688b6c621898f6351ec79b359218a35417e0

Observation fa14debb-98ee-4397-88ac-456b940528a8 · outbound

This paper cites @esa (Ref.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought @esa (Ref

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.971079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.971079Z digest=sha256:356fa95dc0152a3370843c375ca92dc79b3f2477304a862a057f3cecba960c95

Observation 9c1cc768-96eb-40c5-9962-4f4079d4f000 · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.976071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.976071Z digest=sha256:89bd553f4e881021be9e634ade749239f4408734a9ba6a3aa24009794cfaabc0

Observation 2db63415-3ea4-479d-9bbd-00fd28ff3b2b · outbound

This paper cites an unresolved cited work.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.980968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.980968Z digest=sha256:b210878ece3024763125927c4dbcca5d66683f05c7782e31b74b39e899b0513e

Pith citing papers

Observation f8bb9d9c-cda0-4a4f-8558-1a01fd930ed4 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:16:19.319482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:56a6e22d3a1c86c057f493ef04ce712844f0c0f9dcc85c9e8769bd40571193d1