Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:07:16.700499Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2606.03503.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T10:07:16.700499Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T02:43:21.225324Z
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6a17d518-577d-485f-93ec-a358ec9531a0 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f49cf19f-65d6-4322-bcfc-d2f993da2c88 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Training language models to reason efficiently
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b90073b6-9a96-4cb8-80e1-c83b82a53349 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a08c83bf-c2b8-453c-9c63-8599e7c17ba4 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 424624be-bd9d-4c71-9937-eacdacd03bca · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00d49656-5c80-4b7c-a988-356d2dd7e3bd · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5df9ee55-0072-4fd6-bfff-245c931b6251 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b9681ed-768e-4778-b904-b2b54736f487 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 456ea93d-5648-4ef8-b3e9-cf9f254949de · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b1985960-767d-4637-8247-67a50ccada79 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Learning to reason faithfully through step-level faithful- ness maximization.arXiv preprint arXiv:2602.03507, 2026a
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fb7b005-ff50-4cc2-a597-31b394148e56 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 26cf7ce9-43bb-4cb4-a279-853901432f8d · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning How Well do LLMs Compress Their Own Chain-of-Thought? A Token Complexity Approach
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e195617-8c4b-4b9e-9fa8-cf4dc2706ae0 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Drpo: Efficient reasoning via decoupled reward policy optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c66803d-ef90-423c-94bd-8205ca26b35d · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Reward-Guided Speculative Decoding for Efficient LLM Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd6a2328-aa74-4ade-942f-ff01d09b57fc · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 07c7eee8-2297-4021-9873-67667021732c · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9ec2970-a2df-4853-99da-8c2886a0abc0 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b86c0352-790e-429e-90fc-b886c36b1cfc · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Frost: Filtering reasoning outliers with attention for efficient reasoning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66151ee2-93b0-491c-af25-4bdabfcd5357 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 977a99e0-495d-45ce-8a61-167a6847f610 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Reasoning Models Can Be Effective Without Thinking
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc880a4a-ee00-4cfb-bbb6-da8304e9a34d · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Topology of reasoning: Understanding large reasoning models through reasoning graph properties
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ead1920a-fd43-4a17-acb3-f7ceeb88bbd4 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Qwen2.5 Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 58b9f8cf-c810-4e1f-a792-f6977e045ab2 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7747edd6-a26e-4d2c-bc2f-f7bae1c63685 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb18b22d-e0e1-4696-8d94-68b31877574e · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Proximal Policy Optimization Algorithms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a668184-5526-4283-b467-d483156f3efb · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Chain of Thoughtlessness? An Analysis of CoT in Planning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f6dd0e9-963e-4cbf-8c38-76b256353d4f · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7972abe0-66d3-437d-abbe-99fa35460622 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dee13020-45f3-4eb0-a6f7-6ecf490e045e · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Qwen3 Technical Report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d572c47d-b0fd-4fb3-84d6-b4530c9b83af · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Solving math word problems with process- and outcome-based feedback
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 39722990-0ca4-4561-97f6-690734e12278 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae4513e5-c3c9-4c0e-aed9-990bd0152ffa · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Memory Networks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6535bcd3-7f99-4609-b5a4-28c92b11fa9e · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c4100a4-29da-44cd-b19e-b60ca32712d6 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 465725f1-26d7-4222-91bd-8585078bb725 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Distilling System 2 into System 1
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 023cf850-694b-4f31-97dc-469375cd5844 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 277f1f2b-2cac-4fcd-95de-d094d35fc5de · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Efficient rl training for reasoning models via length-aware optimization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdade60b-217c-4cc8-8d78-7d16e40bf6c8 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Promoting Efficient Reasoning with Verifiable Stepwise Reward
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62a23dee-e4dd-4c7c-8039-48fe40462f59 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning STaR: Bootstrapping Reasoning With Reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d9e616b2-281e-4a7c-bae7-49caaba5bf14 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Intern-s1-pro: Scientific multimodal foundation model at trillion scale
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33041e62-0011-49dc-bacd-603141ca1b7e · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Quantile radius analysis We characterize the resulting representation distributions using the Effective Radius
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c42a5d70-495e-4e28-b384-9015b1553330 · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Recent works further improve PRM data construction or stepwise correction for mathematical reasoning (Sun et al., 2025; Wu et al., 2025b)
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e5e486-5ea0-4d3b-8911-aefbb7f4716e · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning Beyond these applications, RL has been widely applied to practical optimization scenarios, such as chip design (Geng et al., 2024; 2025; Wang et al., 2026)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d192b59f-c233-435f-b6d1-5eb2de62510b · outbound
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning These findings support the use of answer-to-reasoning attention as a practical proxy for identifying low-contribution reasoning steps
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cb42830-614a-4c62-bcd0-5bb19caaf82c · inbound
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.