Pith. sign in

Paper Citation Record · LEDGER

Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2410.09302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09302 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.735791Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:14:10.972580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 929b64d2-b666-48b7-907c-3b72e02f23cc · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.735791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.735791Z digest=sha256:19fdc3cafcd7d0d42d0a2c8c10aa9fbebe0c18f7e41ebbdbbced2475dfc796df

Observation b41e75c3-aeea-4c77-91fd-e1d67ff4f3eb · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.666597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.666597Z digest=sha256:4fc5a2668691d9279c320877589e44039b1a9fecc073297e0c2c6de79b3baf0b

Observation 5dd6666f-6906-4848-a8cb-0064b34dee83 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.974237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:adbbd590a6635f79813bd1a35b2a1b44e641905f9d0ab713c523dc1c90ed6d28

Observation a65996f4-6add-4eef-9d4f-6b316519e0b4 · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:43.637082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:43.637082Z digest=sha256:ca4140d5606a0e27eeb1117f0a7751bad7987eaab799ca14989b1b778dd910c1

Observation fc0a83a1-0b42-475e-a529-e1dede6ea1fc · inbound

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability cites this paper.

Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:30.965350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:47:14.379908Z digest=sha256:d8d7e9b87cb3e920edb1e03f4051cfe58c38c5ef5144d1bb241d1ddaa9df6c26