Pith. sign in

Paper Citation Record · LEDGER

Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2409.09345.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.09345 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:05:50.935027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T00:26:48.578074Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ac43e7f8-0af6-441c-a216-30d54a835a2d · inbound

Aviary: training language agents on challenging scientific tasks cites this paper.

Aviary: training language agents on challenging scientific tasks Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:07:33.638201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:07:33.638201Z digest=sha256:41d5c776511efa30f44fac532732f8796c707b4c262d11dbe983c0448d1743d1

Observation 055d0052-2503-4156-b4c9-94ca80654ecb · inbound

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training cites this paper.

Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T18:22:02.748089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:22:02.748089Z digest=sha256:f10b29f250927ce022c31123e974eb07a87e008f2cec5e1e1cfd303802c9a524

Observation a2b80174-0589-444b-9c80-df89f3992d60 · inbound

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search cites this paper.

QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:31.398270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:31.398270Z digest=sha256:d5ecc5c72622f6a341a53c8c71ebcbac0da2c666d4cd8f5a5af951f4057ba0f2

Observation 2f110b1a-d3c9-49e0-b9a1-cccdc5bde5d2 · inbound

ToolRL: Reward is All Tool Learning Needs cites this paper.

ToolRL: Reward is All Tool Learning Needs Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:26:48.581281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:26:48.291431Z digest=sha256:cfe2e7f34393040ad87154f713fbb62f503de690167da2414a8f94edad371f90

Observation abbbeb2b-ce29-407a-843f-5a0d3db5bcaf · inbound

DSADF: Thinking Fast and Slow for Decision Making cites this paper.

DSADF: Thinking Fast and Slow for Decision Making Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:05:50.935027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:05:50.935027Z digest=sha256:38a99819f29a2def5be05c6d799b78f8bae80195b7d0d0242f78346ebf76c46c

Observation 72928ad8-48a3-482c-9f98-3429532b762e · inbound

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents cites this paper.

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:45.285843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:36:45.285843Z digest=sha256:df66d80a594c62916d3f1aa4a9278d986d4f99618fb288e9ead6fab6bcca84b6