Pith. sign in

Paper Citation Record · LEDGER

HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.07070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.07070 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:56:44.076269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:12:34.960785Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5b693abb-6cfb-44d3-9adf-1f45ad6794b0 · inbound

SCAN: Structured Capability Assessment and Navigation for LLMs cites this paper.

SCAN: Structured Capability Assessment and Navigation for LLMs HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T16:11:45.580626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T16:11:36.613334Z digest=sha256:077a68f9acd1d57d029890ba2246cc8ffb595489f92fb18637b972acc614ee88

Observation 10dd5baf-7564-4263-9296-dde44e02ac82 · inbound

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them cites this paper.

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:06:37.984774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:06:37.984774Z digest=sha256:982999df58d2b151adcec1ec001da643499587ffc5c94ebd106728ff9cabf0a2

Observation ccd89c21-4c34-4ef5-aab7-eda1e320dff1 · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.076269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.076269Z digest=sha256:cf18ed8b4437177d14c30a82ca0848451877ab7f4c7dd03e57e42401a4528960

Observation 155a1c38-1a1d-437b-9801-3808f3d88e84 · inbound

ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking cites this paper.

ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T23:33:21.095017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:33:21.095017Z digest=sha256:a9358edc9a298ab8b26dd9bacf283a557b1d15141d8782a68376727fa9cf30b2

Observation e044825c-c392-463e-a695-4204b8d1d0fa · inbound

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality cites this paper.

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:50:58.583102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:49:33.972347Z digest=sha256:30ee3033f7aa6bc14f52a49905aa56011dd36ccd2cbdadb19bd701774422741d

Observation c8de52f2-f6b6-4157-af4e-212a55228296 · inbound

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue cites this paper.

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.962047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:06:27.710746Z digest=sha256:88e584921dca0462b90ff54bfde294a9d6da87ab55161df9b9e311bdd41900f8