Pith. sign in

Paper Citation Record · LEDGER

How Far Are We from Optimal Reasoning Efficiency?

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 2 inbound Pith citation observations for arXiv:2506.07104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07104 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:40.577890Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:53:48.051608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T13:22:55.270834Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved54
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dcf89b3e-d763-4ed3-af92-b11c20922e51 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.499159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.499159Z digest=sha256:e08d0deb33329bcd8c45dee7fd7dc8045fef6fb675761c75dfbca29a31955ca4

Observation 8db244d9-d272-41db-af6f-7ea1ce392dfc · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

How Far Are We from Optimal Reasoning Efficiency? Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.596778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.596778Z digest=sha256:75146a3c19673850dd68fcdfb3ba10a748001f6119ac99ee8b58c0054f89b760

Observation 6938d49c-3201-49b0-89f0-480e540e7bec · outbound

This paper cites HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs.

How Far Are We from Optimal Reasoning Efficiency? HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:34.696226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:34.696226Z digest=sha256:33837b2bdcdf116856dddc6ccb490adc9ff100de4749f0d1392a8dd6c045c52f

Observation 9769068b-d3e3-4743-a4f4-d51336ab5248 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

How Far Are We from Optimal Reasoning Efficiency? Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.028025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.028025Z digest=sha256:8ef2ead78fe9a162a03604b3a23c5a47292d5f17779187b2eb0c9e3835f5cfbb

Observation 24297eed-6dbe-462b-9d27-58e0c3675a5e · outbound

This paper cites Thinkless: LLM Learns When to Think.

How Far Are We from Optimal Reasoning Efficiency? Thinkless: LLM Learns When to Think

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.129268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.129268Z digest=sha256:5c47de848fd5bf5e50c740fd57e22cff519b002e4200a020526fddfe5499a637

Observation e92c41ce-5ba5-40f2-a2e8-addd6482a401 · outbound

This paper cites Efficiently Scaling LLM Reasoning with Certaindex.

How Far Are We from Optimal Reasoning Efficiency? Efficiently Scaling LLM Reasoning with Certaindex

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.259278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.259278Z digest=sha256:83373e4d8dbec45e925922249239e9579f4b87f4b1c5a045618ad8a894f47807

Observation e3fa284a-d61b-4d75-a819-64c533390610 · outbound

This paper cites Reasoning without self-doubt: More efficient chain-of-thought through certainty probing.

How Far Are We from Optimal Reasoning Efficiency? Reasoning without self-doubt: More efficient chain-of-thought through certainty probing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.358710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.358710Z digest=sha256:d84e9f6bbd9048ebd50e5f29f2cc19c334795456d3d30d7013b0ccb04dc6c2ea

Observation c1ffa9ca-0a8a-4e7c-9c76-86b37ecdcb0b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.468940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.468940Z digest=sha256:e46b57e45f91b10b600adf055d254841b0ceb65de37234380d0b08286194e5df

Observation 640a7e38-dd98-47ee-bef0-a35c7b92c105 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Token-Budget-Aware LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.589682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.589682Z digest=sha256:a347f077e8923169426ee1a26c442f9e41b9a9a52a7624316f1a0fd6800184b2

Observation 9e3c7f12-3266-4a4c-a2fd-a26d3488b72c · outbound

This paper cites Skywork open reasoner series.

How Far Are We from Optimal Reasoning Efficiency? Skywork open reasoner series

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.383342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:35.713968Z digest=sha256:4fbed7e44273f560d6106ecd167721d94ec54f54829730e222a01b287b83f366

Observation 549e5200-69e8-4e89-9a22-a52ccaa7d237 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.815835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.815835Z digest=sha256:72ab1844877d5740bc0d9c03faf1173df90e8101b18bbfc1d4faf130d089b7e4

Observation acdd038c-6254-4ca2-82e8-ba764b5ad559 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

How Far Are We from Optimal Reasoning Efficiency? Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:35.897085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:35.897085Z digest=sha256:072462641b0dff0462c7bfa5c15708282aa2825cc3053480a703a97050825393

Observation 96c1265e-c539-45b8-b150-55341afd6eac · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

How Far Are We from Optimal Reasoning Efficiency? Efficient Test-Time Scaling via Self-Calibration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.050461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.050461Z digest=sha256:d0e4134938dac96e8170ef5c7c1dac1c01e813cc0a720785e1b596edfc9c90f5

Observation 7b95b0bf-bfc5-4c36-a823-b52e471b94d0 · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

How Far Are We from Optimal Reasoning Efficiency? Think Only When You Need with Large Hybrid-Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.120187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.120187Z digest=sha256:78077fa358242b17676680f95183e126cb25fb329327eb5fce03c3ceb1cf8e91

Observation 31a2eb7d-20eb-4a0d-a9e9-925851659732 · outbound

This paper cites C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness.

How Far Are We from Optimal Reasoning Efficiency? C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.187134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.187134Z digest=sha256:25de69491b20cfe1f0db56619064376f6c7e3aa3e98fb016d7b5ebf9335e07cb

Observation d09a825c-5489-4883-9108-1d19d919c400 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

How Far Are We from Optimal Reasoning Efficiency? Solving Quantitative Reasoning Problems with Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.305178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.305178Z digest=sha256:dc499ecea3ec4fb817dd902b8ca09c2c6e3a0e274b1d8f1b5d6830f5a1eccf81

Observation 86b8cbe4-6afc-4a59-89fe-b898730394fb · outbound

This paper cites LIMR: Less is More for RL Scaling.

How Far Are We from Optimal Reasoning Efficiency? LIMR: Less is More for RL Scaling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.467346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.467346Z digest=sha256:6049a6bc5d07ae4aba7d28b86499aba11d8d856f49c243ebb1ee5efa9086b373

Observation 2be6c0ee-0ca8-4e3e-abbc-00a4b245e895 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

How Far Are We from Optimal Reasoning Efficiency? ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.597126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.597126Z digest=sha256:c7d77713c9ef5113e11fbe9c8c851c3c7ba6f69456d7f83ad9c679e09227839e

Observation 1a1662d7-9fd7-4836-a74c-7d216dbed935 · outbound

This paper cites Reward-Guided Speculative Decoding for Efficient LLM Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Reward-Guided Speculative Decoding for Efficient LLM Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.717578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.717578Z digest=sha256:771db27f802838e0f3af2f82c1294995c2a60babf64693698360360f26c270a9

Observation 29cdb302-a971-4b8b-8d35-b197d55327d9 · outbound

This paper cites Can Language Models Learn to Skip Steps?.

How Far Are We from Optimal Reasoning Efficiency? Can Language Models Learn to Skip Steps?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.874631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.874631Z digest=sha256:ab1cbd6517e2401fa30836ac4964c1d660a26b0a5546212c5583a2514d3d3fc8

Observation 602394ee-2587-4421-b319-6b919e1f567d · outbound

This paper cites Fin-r1: A large language model for financial reasoning through reinforcement learning, 2025.

How Far Are We from Optimal Reasoning Efficiency? Fin-r1: A large language model for financial reasoning through reinforcement learning, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:36.987668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:36.987668Z digest=sha256:512d205c9cf2a154e06a7faca33da1aebc3992e6cc20fc3cffb2a8dd8343310f

Observation 58910aff-413b-4669-83d8-436907fc1c4c · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

How Far Are We from Optimal Reasoning Efficiency? AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.109195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.109195Z digest=sha256:6501e873efda809232901444b473c288062a2350cb16bec3b4b6ae8ec1292b15

Observation 1ebbfb26-892d-4476-a44f-5c1083f4cc91 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

How Far Are We from Optimal Reasoning Efficiency? O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.215102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.215102Z digest=sha256:5b253776f6fcaa994c235dd691b544559c581fe05ae543d2f3c0940cd4afb61b

Observation b0453166-e802-4e6b-93b9-798552d70296 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

How Far Are We from Optimal Reasoning Efficiency? Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.236363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:37.308823Z digest=sha256:3e780b73f7ef1a397df94d1b2616df748e68d962465dce82af6bb0f0a9017c17

Observation c377da9b-340c-492d-9b00-f5ca907016bb · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

How Far Are We from Optimal Reasoning Efficiency? Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.401902Z digest=sha256:bc04358213736b2ce7547846a9421e851f51e119efe26048d7d023e13e1c5115

Observation 4597f838-16df-4b7d-875d-2b572aa48740 · outbound

This paper cites CoT-Valve: Length-Compressible Chain-of-Thought Tuning.

How Far Are We from Optimal Reasoning Efficiency? CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.540313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.540313Z digest=sha256:eb91f98d8350c85a2b373a1d315a95454dced10270554a3c08a97908afe7a770

Observation 7f0e3b78-4b8a-4a0e-b1d8-301a544ca41d · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024.

How Far Are We from Optimal Reasoning Efficiency? Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.684760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.684760Z digest=sha256:080a865e98711d4b88d4bac9bcaffbc307fe46c8b4567eccc7cc811bc28b564c

Observation b0bfffb9-fc19-4282-bd12-c6c407233330 · outbound

This paper cites s1: Simple test-time scaling.

How Far Are We from Optimal Reasoning Efficiency? s1: Simple test-time scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.783145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.783145Z digest=sha256:2a89b44975c1c763ebb6fc2fd241813bf73f0473f82213d0bfa5d527797b7cf2

Observation 5fb48e33-8d90-4743-a1d6-bc5beb2fc1aa · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

How Far Are We from Optimal Reasoning Efficiency? Self-Training Elicits Concise Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:37.914205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:37.914205Z digest=sha256:a7460d17eaa9cba05ed63acf6f91f12aa0ce5591f592483d191e169e7217bb52

Observation c885bb26-d3fa-4261-b1ed-9b76950c1eab · outbound

This paper cites Learning to reason with llms.

How Far Are We from Optimal Reasoning Efficiency? Learning to reason with llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:42.041410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:38.058152Z digest=sha256:2a612ed2917faa30e1fff4259a9c899e516b0ba2dd5b491343ca633471796625

Observation 30c26e1a-4e14-4d49-b0b5-ba9f08e06c7d · outbound

This paper cites THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models.

How Far Are We from Optimal Reasoning Efficiency? THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.203391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.203391Z digest=sha256:488bd40fd3cb863ebfae808b0f7d39ecf063a81fa002f39c14999777eb4bb982

Observation 12d3f2f4-1a96-445b-97b0-651bbd5e1875 · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization, 2025.

How Far Are We from Optimal Reasoning Efficiency? Optimizing anytime reasoning via budget relative policy optimization, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.275527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.275527Z digest=sha256:c7210917832cfc28f708fcf8f56c6da9a96c7b29f51650c3cce40faf09d35e65

Observation 3b642e89-0914-4e15-873d-ee3d7fab022d · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

How Far Are We from Optimal Reasoning Efficiency? Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.378369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.378369Z digest=sha256:1c49384948d6603083b2bae219470c69cd101dbe50f93116f783a45b6a093c0c

Observation b79cc64b-1687-40a8-aa04-422e5c900707 · outbound

This paper cites Areal: Ant reasoning rl.

How Far Are We from Optimal Reasoning Efficiency? Areal: Ant reasoning rl

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:41.896795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:38.442171Z digest=sha256:8cea1f02efb2eb1879ffe8bac927beb49813e336462d4385c9b0f395e7c14205

Observation 1a5c1ea8-2935-407c-8dd8-4cc7c720bb0e · outbound

This paper cites Hawkeye:Efficient Reasoning with Model Collaboration.

How Far Are We from Optimal Reasoning Efficiency? Hawkeye:Efficient Reasoning with Model Collaboration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.520326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.520326Z digest=sha256:c55eacee55adc47e4cd0adebfd8e2f4c6a29d4a67b665edb1e6acd2776f69128

Observation 881669e3-6609-42a6-a6ce-41f8cef58788 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

How Far Are We from Optimal Reasoning Efficiency? VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.588973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.588973Z digest=sha256:6aed34e94005d5bfe0ad1a172ab14a244783f6b553153e0d085e209f26c18db1

Observation da7c21a4-0088-49f6-9cba-0772561ccad2 · outbound

This paper cites Dast: Difficulty-adaptive slow-thinking for large reasoning models.

How Far Are We from Optimal Reasoning Efficiency? Dast: Difficulty-adaptive slow-thinking for large reasoning models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.672014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.672014Z digest=sha256:44b516d6414b5539fef6bbb24d176969e85fb1947bbdcbac7d24a3261022ef37

Observation f44f5bda-0b92-4924-abc8-e3dc9bcced9a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

How Far Are We from Optimal Reasoning Efficiency? HybridFlow: A Flexible and Efficient RLHF Framework

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.781973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.781973Z digest=sha256:ef09950f25bb8f230d767f8e84333ddfd75e586ddfd914ae3ff952e201905ee2

Observation 1ea4eec6-f2da-4d76-8c2c-e7ede3698bc7 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

How Far Are We from Optimal Reasoning Efficiency? Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.874482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.874482Z digest=sha256:46b547d15d5bc7e44ce257492b6dfcb19bcad6a636833663d6a544a461ac3c71

Observation 33a60b7a-8803-4660-9fd1-0d1350d14818 · outbound

This paper cites Fast Best-of-N Decoding via Speculative Rejection.

How Far Are We from Optimal Reasoning Efficiency? Fast Best-of-N Decoding via Speculative Rejection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:38.967912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:38.967912Z digest=sha256:5d8c605dc38b0e338d3ca4bb928d8f14dec4917614060cb6571433cf70df717f

Observation cc5e8e5c-9620-4805-8292-b4227d436795 · outbound

This paper cites Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233, 2025.

How Far Are We from Optimal Reasoning Efficiency? Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.034718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.034718Z digest=sha256:51a9a19e52b040cddf26e47238befd10785b58f61d9a851aabe092317d428059

Observation 62567a75-a81b-4646-a49a-eddfb5bd2f03 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

How Far Are We from Optimal Reasoning Efficiency? Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.110881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.110881Z digest=sha256:e2f00f823ee563a4c936d9a0d0ca3b4cc964151c8ecfaefd1a326ec2f0d34cfd

Observation c2579418-fe62-44ed-a47a-5b335be1e611 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl,.

How Far Are We from Optimal Reasoning Efficiency? Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:41.736320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:39.175313Z digest=sha256:dd44b175104ab3f022f9225bfc00030ea9e3ab48d4a6c5036c72a7db140a7ffc

Observation 71bb0eea-045e-40b4-be43-5834fba86634 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

How Far Are We from Optimal Reasoning Efficiency? Self-consistency improves chain of thought reasoning in language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.336776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.336776Z digest=sha256:b763c410d3957e354d3c1871f280e94054c50dbabdf96013a01023aaf5a50383

Observation f6eb5b31-48c2-4dc8-a33c-42347b1f8cf9 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

How Far Are We from Optimal Reasoning Efficiency? Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.427926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.427926Z digest=sha256:ffa4063f6342ee6ae1e0844621872e0ece2f378fe1e634c9ea40cd8169294e96

Observation 50582151-68b1-42f7-aeed-02770b8720a5 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

How Far Are We from Optimal Reasoning Efficiency? Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.495400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.495400Z digest=sha256:2b1ff318103c0fe777d7ba218fe83ec5ceca99d43c1828be90e001af0c326c6a

Observation f1a8cdf3-9093-43e4-b5df-b24053add6fd · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

How Far Are We from Optimal Reasoning Efficiency? Scalable Chain of Thoughts via Elastic Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.585269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.585269Z digest=sha256:2a4f0dd41b1399f02460d880b9b27778cb910448094a80ceaa34c052d82dc4c2

Observation 66b4e1d8-cb54-4e43-bcdc-9bba618053fe · outbound

This paper cites Qwen3 Technical Report.

How Far Are We from Optimal Reasoning Efficiency? Qwen3 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.672165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.672165Z digest=sha256:6749d4cfa8d898b4d9cc08b92f93a6488e7c8fe9554240e36d9288c86edcfc33

Observation 69f3e7db-693f-42a6-82b8-c1ec8f1dc793 · outbound

This paper cites Think When You Need: Self-Adaptive Chain-of-Thought Learning.

How Far Are We from Optimal Reasoning Efficiency? Think When You Need: Self-Adaptive Chain-of-Thought Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.753920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.753920Z digest=sha256:1568d5c92b4e5a9d3d8c6be4ab50162580cded6eae7f38871a4234a58fc72e63

Observation bd9e2959-fb9d-4430-8ac5-4ba499f62c51 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning, 2025.

How Far Are We from Optimal Reasoning Efficiency? Towards thinking-optimal scaling of test-time compute for llm reasoning, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.840510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.840510Z digest=sha256:18712c43299c2660aa787cc9ad11db4bc9a46e3e2f3eec1d33c97772e2017936

Observation 6af1f9e4-a779-45f5-a6d3-8a07b84ae273 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

How Far Are We from Optimal Reasoning Efficiency? Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.909550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.909550Z digest=sha256:b9720c81d2f7db9fff519026170991cca1c17895af94c0857b416279fde3714b

Observation 589595c2-21f7-4948-a720-d00509f74398 · outbound

This paper cites FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training.

How Far Are We from Optimal Reasoning Efficiency? FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.978220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.978220Z digest=sha256:40f59ae6452aa6e03f104146e6c863b86c7e45a2aa04a8533512a544a6df65d6

Observation 3def0778-7059-489c-b93e-13253303564e · outbound

This paper cites Distilling System 2 into System 1.

How Far Are We from Optimal Reasoning Efficiency? Distilling System 2 into System 1

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.051297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.051297Z digest=sha256:9e1ff8907d0659980e988408c5a3bb062d2d2d2bd5930259fc6e59bfd8eeb58d

Observation 6322f138-1f81-47ad-8802-2bdb1a21490b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

How Far Are We from Optimal Reasoning Efficiency? DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.129855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.129855Z digest=sha256:490a3016de9397a72cbb8577b7cb72925b049ce4dd6b21a328dc2770ca41b367

Observation da1c3027-fe12-44b0-9033-217c568dacd6 · outbound

This paper cites Z1: Efficient Test-time Scaling with Code.

How Far Are We from Optimal Reasoning Efficiency? Z1: Efficient Test-time Scaling with Code

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.227034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.227034Z digest=sha256:f38a338280d86cffceb2b30d75cd15bd442137741d850239230ff074f09fa0ed

Observation 487a986d-c299-4028-900a-37dcdba052b3 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

How Far Are We from Optimal Reasoning Efficiency? VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.307416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.307416Z digest=sha256:d4bcf0e8d86ab2b3b175e22820bbe70d8819e5193de9f39feef40a304d3fb6f2

Observation db605667-f43e-46e3-a2bf-6367e64205fb · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

How Far Are We from Optimal Reasoning Efficiency? AdaptThink: Reasoning Models Can Learn When to Think

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.388824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.388824Z digest=sha256:2b11057d73c1bc554fd7206969120656d133f516096d0a02f50073f8c596589a

Observation 2ebf69cc-e2b5-4f41-a2a4-bcc8a30e0fa4 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

How Far Are We from Optimal Reasoning Efficiency? R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:40.486336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.486336Z digest=sha256:7bab2d94cad51b6d57663073697637ca06e4f37f04fa89a0a77c19d48aa444ba

Observation b26ff1b8-67af-4eb3-8ce0-a5385bfacbe9 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

How Far Are We from Optimal Reasoning Efficiency? SGLang: Efficient Execution of Structured Language Model Programs

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:49:40.577890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:40.577890Z digest=sha256:71e2733506a8b9f6f427a2ea412b3afb2a64009b0020da77b8d33786c6ee325d

Observation 29e8489d-c029-4956-befa-11f8158d156d · outbound

This paper cites an unresolved cited work.

How Far Are We from Optimal Reasoning Efficiency? Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:39.249076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:39.249076Z digest=sha256:77f7a3645c4509669c2e3adccb37c4f8d1aaa816fd447f5208c4358892cad6f4

Pith citing papers

Observation c1c4647e-b274-4e3b-8244-f619c76f0c3b · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey How Far Are We from Optimal Reasoning Efficiency?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:48.051608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:48.051608Z digest=sha256:7c4334d52f017bc457bee3acd6ec13d3d310dfd72dcf0c82b189b86bdb52a22c

Observation 8549c248-218b-49cd-b042-681bf3ad756c · inbound

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models cites this paper.

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models How Far Are We from Optimal Reasoning Efficiency?

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:22:55.272653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T13:21:36.606855Z digest=sha256:6700f70ba47f7a8ef2e609a7b662d643e9498a4efd59e682fe5bf45dd1ae2b9c