Pith. sign in

Paper Citation Record · LEDGER

ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.20015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20015 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:48.107868Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T16:53:40.615942Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a162cb5d-ab0e-473a-9ff9-4827dec9ee5e · inbound

Reducing Tool Hallucination via Reliability Alignment cites this paper.

Reducing Tool Hallucination via Reliability Alignment ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:47:17.066252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:47:17.066252Z digest=sha256:f36f12c8bca486e652336f3a86748036aff8e7fce78d202625a06c8eacb1d723

Observation 79e100d3-f173-4ecf-8cad-bceb313c8fa4 · inbound

When2Call: When (not) to Call Tools cites this paper.

When2Call: When (not) to Call Tools ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:48.107868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:48.107868Z digest=sha256:306ff766e3ee810de666c2af67bf48ba7429613a7e836e02a9d315fb47a33749

Observation 7a802829-5b63-4fef-9498-339fe3536c9e · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.768323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:02:55.529271Z digest=sha256:70673c05e6edf9556f35d48ad0b1d20727b73fb842629971d2d5031950517ce9

Observation 29b5b9a9-4885-482b-a389-0d078725eb55 · inbound

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents cites this paper.

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:13.276790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:13.276790Z digest=sha256:3e4f889255abfb474ea5efafc6b3d5b1f5b6fcccd6c7f59f49a3d4db8105f8a8

Observation 6c08a245-735d-493f-8bf9-cc39efbb1b4c · inbound

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception cites this paper.

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:50:52.552687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T03:46:03.228969Z digest=sha256:f20b552f21f1668d259c84ba69f74c3ecbe918e3cbc04ffb9be3553435a9fa6d

Observation 3baf782e-911a-4674-a71f-5abbd8aec7c6 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:35:04.458129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T05:31:54.252816Z digest=sha256:7a01ffa67dfdb500cef8e856f02ff025868fa6a78cb8885100bd126d494626df

Observation a24ddf88-4a77-4ff5-babe-b378d297ab64 · inbound

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use cites this paper.

Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:53:43.543715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T20:52:20.975459Z digest=sha256:55bb56da487319360d7ad12b60db455d4dd6e73b5dabc2b6a13843719d910db4

Observation 27a4c51f-cec4-441a-9980-a3141ab48d2a · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.617632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:e2656131583043bad3a91be4c2db9ca613e472d349f0995b507541289b24cd3b

Observation 39fad04c-9c0d-4f00-8d94-09b7a528e154 · inbound

SAAG: Structured Agent Assessment and Grounding cites this paper.

SAAG: Structured Agent Assessment and Grounding ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T15:11:09.603684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:11:09.603684Z digest=sha256:6b5deb7c16fb6fb1713a9df8ea6eb87784172f58c6e9695f6da5ad276c1180f5