Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Budget Policy Optimization for Adaptive Reasoning

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.15844.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15844 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:31:28.767651Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 282d0983-76a0-4e7d-9269-597ef0fabc77 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.600626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.600626Z digest=sha256:ebaccd9639830ad1af63231e4a4e96f29b757da921f61c94735dc22ac0edfba1

Observation bc1f98b9-2226-48e4-a27a-b217e678a2d8 · outbound

This paper cites URL https://doi.org/10.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.692180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.692180Z digest=sha256:5606693d4e6f0bf8390a1af3d29fa0dcdd01acb2050c8f6cbace466b8dee6350

Observation 2a4f7ba3-08d8-4a34-9586-a2519ef98720 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.868293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.868293Z digest=sha256:7f863ccaf2f1471f19f79bdd83bcae122a8b752c72d7bcef544fad5bc4b5abe1

Observation 4fc7dd44-8184-4f22-b1a9-b918be534d1f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.319209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.319209Z digest=sha256:3f6b3ddfae70b4a29e1b0a2a26a261193dc5b21838615f3ab77a3c64a4f207a3

Observation 6c0aaae6-8018-48ee-8579-0d863a965da8 · outbound

This paper cites Thinkless: LLM Learns When to Think.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Thinkless: LLM Learns When to Think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.442287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.442287Z digest=sha256:097124dbda5abf9c2e62dc2ab0d2c5617bb4b23c7a441578e2339780c24b7c62

Observation 87229f48-e3d9-4439-b38f-18c0427e85b8 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.609818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.609818Z digest=sha256:d0c83be2d08d96dfac38a02111f274207d39de3d3d4236a10437a0ad1f9cda50

Observation d8c3f112-3400-466c-8a33-b1f78825ba81 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.748148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.748148Z digest=sha256:cbb39f4937c4301db4022e0dedfab126039a5a1233197d17def8e6fee655e42f

Observation c534d5d1-4d77-4b2e-b298-c0cf235de80c · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2505.11225.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2505.11225

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.958656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.958656Z digest=sha256:74c06165824abb2fefaa3b47a1844c5032e0ac1638bda52c29cd6b60b1fc6e12

Observation e0714b9d-f4a4-46c2-9aee-ba6de4f8bcca · outbound

This paper cites Think Only When You Need with Large Hybrid-Reasoning Models.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Think Only When You Need with Large Hybrid-Reasoning Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.102020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.102020Z digest=sha256:5b3ef38448432b5dbc89422877f532eb68fbeecb8eaffb59a6a005c3fb1eaca0

Observation 02bf1114-67cd-4d06-8c9d-f67a7d202e2f · outbound

This paper cites SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.228889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.228889Z digest=sha256:f6933cc5fa55af57dc568709e0292d86010d8b887e57164723e0cd06aa9a7047

Observation 3754db1e-dab9-47d9-9389-9b7d56f7f4d3 · outbound

This paper cites ThinkSwitcher: When to Think Hard, When to Think Fast.

Hierarchical Budget Policy Optimization for Adaptive Reasoning ThinkSwitcher: When to Think Hard, When to Think Fast

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.312222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.312222Z digest=sha256:bc8d33b8ab72c1a87027082cc8f4222ba5d06e1544a62f1a8ed6b0e7a595ba40

Observation 06e9ff18-fa4c-457d-ab68-213df15f9805 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.441783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.441783Z digest=sha256:d1c5da5e219dedc036ebe9dd5ca0146e6cffdea58ceafff33f4f798c93f84f34

Observation 8c52a92a-bdbe-41f5-8298-c7fb10f3f786 · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.526529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.526529Z digest=sha256:9108092cea0ae23da099c8bbf87672a4d21924d281152378e7d383c6e13af464

Observation 040d2d23-8868-452a-8f41-466916fccb9b · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Reasoning Models Can Be Effective Without Thinking

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.578878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.578878Z digest=sha256:b401be73c28b0c73f6537ac74da913d5efa7ecc738d2e9b3a34bbef561004f4f

Observation 35312d21-cb31-4f3e-8bda-dc2b7aae9586 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.703711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.703711Z digest=sha256:9487d65c15e06297a2192aa42c2880113657a0be8665260434e3eeeb37f5e8a6

Observation 5968c0c8-97cd-47c5-a1f7-f0cf736660e6 · outbound

This paper cites s1: Simple test-time scaling.

Hierarchical Budget Policy Optimization for Adaptive Reasoning s1: Simple test-time scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.722892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.722892Z digest=sha256:4a7163927aeb3bd062851c9d3797d014def0ffbde71131b8c33591b94fac359c

Observation 08e6a92c-a9df-4c47-970e-d32541be5277 · outbound

This paper cites Accessed: 2025-07-22.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Accessed: 2025-07-22

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.728133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.728133Z digest=sha256:c59cc905624dda0714aa9993586ef692f41221a7cfbbd84e987140359ed1966d

Observation b823124c-5b7f-4289-abb1-10c120f67e88 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.04881.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.04881

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.181641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:31:28.733587Z digest=sha256:fc99c702cab00f6f0008182b0368e3a5a65e5c6e0dff66ecfbdcf0f57b3fb2a9

Observation a077caf7-ac67-47b6-8c0f-5df45d927f99 · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage RL

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.738590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.738590Z digest=sha256:c8187dc8db065744cb188e84bc4f856ca014d02758fcb8feadf0ee2aa5c7480b

Observation 5c5a568d-8990-4f95-a77f-857af800e004 · outbound

This paper cites URL https://doi.org/ 10.48550/arXiv.2505.10832.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/ 10.48550/arXiv.2505.10832

Reference 24

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:31:28.743999Z digest=sha256:63f5858c823b8e616c4fb38296ddf57028b2712ebb89d4a4b168993156ba9748

Observation 3a5e8ee1-c570-490f-991a-e77582d4a762 · outbound

This paper cites PATS: Process-Level Adaptive Thinking Mode Switching.

Hierarchical Budget Policy Optimization for Adaptive Reasoning PATS: Process-Level Adaptive Thinking Mode Switching

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.748477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.748477Z digest=sha256:a8f9d1666c1287197d7abf74d002b16d7759a8f5d0777a78191706586c96ee74

Observation 44edbed9-a921-4666-a661-fa77e446d855 · outbound

This paper cites URL https://doi.org/10.48550/arXiv.2505.20258.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URL https://doi.org/10.48550/arXiv.2505.20258

Reference 26

Resolution
verified exact
doi, observed 2026-08-06T15:31:29.005466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:31:28.753494Z digest=sha256:34ab78bc1fc04c37485450d93d10c2999357c5cbf7566844b309e41060db561a

Observation d4571c65-7566-4e78-958d-fde3db65fa6b · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Scalable Chain of Thoughts via Elastic Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.757943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.757943Z digest=sha256:78563f028ba6ad90ec953b398f9f44f1cf97084e2af23b733098bbfebb37acd9

Observation 092df8c9-b248-4eca-91ce-1631970e42fb · outbound

This paper cites URLhttps://doi.org/10.48550/arXiv.2504.15895.

Hierarchical Budget Policy Optimization for Adaptive Reasoning URLhttps://doi.org/10.48550/arXiv.2504.15895

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.763066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.763066Z digest=sha256:b776df4f3b70d4bd9db98b59856fa5bef56542c5822972cb8e621dee218e77d5

Observation a884b321-ad3e-4dd1-8214-a6eead52d5c3 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.767651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.767651Z digest=sha256:ca4c48b3ba9eb423face66bf4b97a8aa20371d449e13224cfd81bb55cba74838

Observation 68b378fb-239c-496d-a120-d98241007ce4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:27.109371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:27.109371Z digest=sha256:cb790323abcc66e85c9167c1f977737e237e7bb22cd1aa400302d329b465dac4

Observation 4b63b925-88cd-4732-a3ce-000e546a4658 · outbound

This paper cites AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:28.425196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:28.425196Z digest=sha256:43bd42224aafaed096798070a4cc304421274bffcbe813326eb534af5cea3190

Observation c7fbf56c-13cb-4258-ac3f-e4ae51ea39cc · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Hierarchical Budget Policy Optimization for Adaptive Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.956879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.956879Z digest=sha256:e90966cbc0a28260bb69e15cc08ce53bdb5d4412158ee2995ec0c0b50acff47d

Observation 4af2d2de-57bc-4c3d-b51b-d6bed7cff138 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Hierarchical Budget Policy Optimization for Adaptive Reasoning L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:26.625005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:31:26.625005Z digest=sha256:9eb365c5937496c03f0a14e36906861b8c28482279651e55a80ce51e2c36a463

Pith citing papers

No inbound Pith citation observations are available.