Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Budget Policy Optimization for Adaptive Reasoning

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.15844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15844 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:31:28.767651Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 282d0983-76a0-4e7d-9269-597ef0fabc77 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.600626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.600626Z digest=sha256:beddd86c4fda8d25c56e7bc5cc2372a1671fe648d7b5c2c537fae66c15fb36ac

Observation bc1f98b9-2226-48e4-a27a-b217e678a2d8 · outbound

This paper cites URL https://doi.org/10.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.692180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.692180Z digest=sha256:0a39582bf846b61ba9d203e717e6692c69437d510498024e62864f89a03d9821

Observation 2a4f7ba3-08d8-4a34-9586-a2519ef98720 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.868293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.868293Z digest=sha256:8b702a10cd2f41cf9b5a0f2e4172247db7cc18fd01adb5acda1811e3cd993dc8

Observation 4fc7dd44-8184-4f22-b1a9-b918be534d1f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.319209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.319209Z digest=sha256:ae571a142797341353ca40e7246bb33ee5e5c6709471facac55f071d0c07ea06

Observation 6c0aaae6-8018-48ee-8579-0d863a965da8 · outbound

This paper cites Thinkless: LLM Learns When to Think.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Thinkless: LLM Learns When to Think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.442287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.442287Z digest=sha256:69cf7cd4de93b3b071b3528ae014556267f19f7d38fd6a444871e109e3ab68a5

Observation 87229f48-e3d9-4439-b38f-18c0427e85b8 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.609818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.609818Z digest=sha256:29f572fe86e866bfb0d1a57ca423ee998ba6163f3d0cccec959e297531aed97e

Observation d8c3f112-3400-466c-8a33-b1f78825ba81 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.748148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.748148Z digest=sha256:7ca42094f819fac5334fd173adc43f7b083314be7e18bdfab06fbcc3a6623d18

Observation c534d5d1-4d77-4b2e-b298-c0cf235de80c · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2505.11225.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2505.11225

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.958656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.958656Z digest=sha256:8ff1e6e51d59a094f020b557e137b3dfc410a125f527cb9f96a372ca126c8416

Observation e0714b9d-f4a4-46c2-9aee-ba6de4f8bcca · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Think Only When You Need with Large Hybrid-Reasoning Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.102020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.102020Z digest=sha256:e998861f91053f3558d710af5ca589dbcfc0b8bd37e861950dab725599db1057

Observation 02bf1114-67cd-4d06-8c9d-f67a7d202e2f · outbound

This paper cites SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.228889Z digest=sha256:4d1eceb1edb829f5e46de82b8b35c9ae50768719757b95d567df8631541b1302

Observation 3754db1e-dab9-47d9-9389-9b7d56f7f4d3 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.312222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.312222Z digest=sha256:20b7fa4fa5a0c6f2cb2923cce22cf7ef2f5d3388a948c4fde42360ef1d5f063d

Observation 06e9ff18-fa4c-457d-ab68-213df15f9805 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.441783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.441783Z digest=sha256:330925508d6459ac19e2933aeba72d7cc552244c347467b4c7f7bc84dc210f10

Observation 8c52a92a-bdbe-41f5-8298-c7fb10f3f786 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.526529Z digest=sha256:ba9999f31ccbc39f59590a306a3e72d676d2e8e24ebba87e7d5a875feff1fc3a

Observation 040d2d23-8868-452a-8f41-466916fccb9b · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.578878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.578878Z digest=sha256:d3ce2e4182b1280688000e1dff76a40aa0088c7387799963f0788c0eb61a7eea

Observation 35312d21-cb31-4f3e-8bda-dc2b7aae9586 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.703711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.703711Z digest=sha256:c77b9aa3d00652506ce18dbc4d4f5f570d03b050cb1bb3d1446b18cf7e5ce693

Observation 5968c0c8-97cd-47c5-a1f7-f0cf736660e6 · outbound

This paper cites s1: Simple test-time scaling.

Hierarchical Budget Policy Optimization for Adaptive Reasoning s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.722892Z digest=sha256:19984b2c0f18541993db55494c23cd750793c00ee890b4a9ec8519c16badbaa3

Observation 08e6a92c-a9df-4c47-970e-d32541be5277 · outbound

This paper cites Accessed: 2025-07-22.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Accessed: 2025-07-22

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.728133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.728133Z digest=sha256:34afd519311e4f63c60d9b5714deaf554a047f12c1b2ed92bf0dba9d3dea49ed

Observation b823124c-5b7f-4289-abb1-10c120f67e88 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.04881.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.04881

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.181641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:31:28.733587Z digest=sha256:b70943b5a7ca4169ae251f077b2e659cf105d4864bbcef517482229ccd760084

Observation a077caf7-ac67-47b6-8c0f-5df45d927f99 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.738590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.738590Z digest=sha256:41d3f1a8e390b09cc761138df2f015258090fec18f0f771eaa77e7b26046cc5e

Observation 5c5a568d-8990-4f95-a77f-857af800e004 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.10832.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.10832

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:31:28.743999Z digest=sha256:55a7d0091bc80f829e210ee291941bd34456d527a35873156e75d4fc4ee4bee4

Observation 3a5e8ee1-c570-490f-991a-e77582d4a762 · outbound

This paper cites PATS: Process-Level Adaptive Thinking Mode Switching.

Hierarchical Budget Policy Optimization for Adaptive Reasoning PATS: Process-Level Adaptive Thinking Mode Switching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.748477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.748477Z digest=sha256:eaf349bff32c94ac2994944ad4592334b73883074c1f747adf521ea9d26ce471

Observation 44edbed9-a921-4666-a661-fa77e446d855 · outbound

This paper cites URL https://doi.org/10.48550/arXiv.2505.20258.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10.48550/arXiv.2505.20258

Reference 26

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.005466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:31:28.753494Z digest=sha256:cf9642388d8299d21e6adf1b26a4f75816aa225e6638f041491020c0f14bbae4

Observation d4571c65-7566-4e78-958d-fde3db65fa6b · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Scalable Chain of Thoughts via Elastic Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.757943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.757943Z digest=sha256:afc24ccf34536616df26bfd22ac6b9537f24eb21fa59cfb497eb40589191027b

Observation 092df8c9-b248-4eca-91ce-1631970e42fb · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2504.15895.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2504.15895

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.763066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.763066Z digest=sha256:bbf3bc9f5425f2ed2c9af0000969b24c494c742bfe80af4d6ee0a5ddc1e6f1d8

Observation a884b321-ad3e-4dd1-8214-a6eead52d5c3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.767651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.767651Z digest=sha256:30d6258e5b9aa567adccd9027fcb8b614b2eda3b4f3456b2e4a8109d51f1138a

Observation 68b378fb-239c-496d-a120-d98241007ce4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.109371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.109371Z digest=sha256:f4e1bedae404ffcb716cfe8ea6e3740b51a238021938d8f57bf2214c6e45528e

Observation 4b63b925-88cd-4732-a3ce-000e546a4658 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.425196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.425196Z digest=sha256:c46adadf41d7ac0bd1b2d47b33c05cd11f7c7440ebcfe4999fb8ce56df0d65f0

Observation c7fbf56c-13cb-4258-ac3f-e4ae51ea39cc · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.956879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.956879Z digest=sha256:df33030c460d8254ac54867d07370a1ab8e23a7e1fc35e869702aa7d8cbadd65

Observation 4af2d2de-57bc-4c3d-b51b-d6bed7cff138 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.625005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.625005Z digest=sha256:aec4a6757b2859f86131a8b04ed5b0a89c059b5f091dffb4f37968aa6d715a5d

Pith citing papers

No inbound Pith citation observations are available.