Pith. sign in

Paper Citation Record · LEDGER

Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.14283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14283 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:09:46.348559Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c8c40f3d-f2c3-4a9d-b0fd-941c90ec8c8b · inbound

Interactive Post-Training for Vision-Language-Action Models cites this paper.

Interactive Post-Training for Vision-Language-Action Models Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:25:47.236938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T14:25:47.178714Z digest=sha256:a21fe46d1e268b24e0fb7277d19ee18425b1bf4089621d5d0f20a71cbd4713ab

Observation 22076380-86b1-435c-9ce2-0aebe7c31ed4 · inbound

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs cites this paper.

Let's Revise Step-by-Step: A Unified Local Search Framework for Code Generation with LLMs Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:09:46.348559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:09:46.348559Z digest=sha256:bd823d6caa71b29ad4e6cce23fa14c310d2bb6b79fbf0a772c8be6e3209d9d17

Observation a504796f-b7c5-4b82-8f82-e852a7351783 · inbound

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models cites this paper.

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:42:36.865021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T12:41:48.620040Z digest=sha256:e3065e84a8ad4c0f894118f7be094a8fe7d05fd969a77b17ebbced4da36e429f

Observation 329921e6-7667-47c8-bc7b-37c87a97f4d5 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:26.008209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:b328bb20d315f69f5af28c65a4688fd84d33b59ec370a57bf6925ce72ba67389

Observation ec8c5023-e6f9-4d50-ba88-74a9af2fd8d5 · inbound

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent cites this paper.

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T19:46:43.187057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:46:43.187057Z digest=sha256:9408110267f28c9bd2ca8a275df68194f4d1b4962a50e5330edc9fc7622b33c5

Observation 235fde31-3dba-4229-9f32-8fd0c75b7696 · inbound

C-TRAIL: A Commonsense World Framework for Trajectory Planning in Autonomous Driving cites this paper.

C-TRAIL: A Commonsense World Framework for Trajectory Planning in Autonomous Driving Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:28:26.090582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:27:31.054348Z digest=sha256:2f67291feba86bd1109cc9dd6e70f6130d34c7062cbb0ee0468fe98283758c67

Observation 5cf1ca7e-1e7a-4e6b-98ef-9c377b8cba9f · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:54:45.454674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:a83a3e4d05cbf8c5fee9be7283877efc637f3796bc960d5a0f84b1f966c1790f

Observation cac4d1a0-e808-47e4-bbdd-f5a01cffc835 · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.224331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:7930ef551c58df8ddd7388db667f2416cc31de435a5629c58752562ceb334da3

Observation 6c00c4ca-adc2-43c3-96fe-516764254e4b · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.212080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:65289067b825f4d476dfa4ba7f22412f82f1361ee5149eaae933ce6d8e541278

Observation 0e2bdf8b-6f9c-4412-baff-d254f61d7b41 · inbound

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents cites this paper.

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:59:38.134654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-26T13:56:51.914966Z digest=sha256:576e122a8dae60faa224f84a83b670997d815f242e0e40883105908b336e91db

Observation ceb18a38-89d6-42bf-ba72-1e84b1f82223 · inbound

REAR: Test-time Preference Realignment through Reward Decomposition cites this paper.

REAR: Test-time Preference Realignment through Reward Decomposition Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:24:19.164064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T06:20:36.864229Z digest=sha256:99f8f6b252bcf0d80b2a24a003518bb82ac7838c490ec2647b59cf83656ff618

Observation 690d9359-25a2-45b1-ac17-5f455d648fac · inbound

Engineering Trustworthy Agentic AI for Critical Systems cites this paper.

Engineering Trustworthy Agentic AI for Critical Systems Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:28.704221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:28.704221Z digest=sha256:e875c435e9dacfeb663b00b84481ad7cbcd0c09e2a11579a6499e1af311e5094