Pith. sign in

Paper Citation Record · LEDGER

ChatBench: From Static Benchmarks to Human-AI Evaluation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.07114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07114 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:38.646723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T05:05:54.821162Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2790fd59-30e4-41f4-81cc-bf53fcf0615d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.390293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:c7e494daa9bafd50bdad7f5f5fd45b2925245757017f90f8e4653d03bc212365

Observation 5989d867-c302-43ac-a69b-f03d7820edff · inbound

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations cites this paper.

ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.646723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:38.646723Z digest=sha256:e35ca8883d2e261ce288f65e90ac1bb600763ff89180f8ec8e4d0da8411d4233

Observation 570370e0-e76e-4430-9284-c1068694b47e · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:29.181666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:29.181666Z digest=sha256:38c473fbc61e06aa2de1de25bd17e27c2311e6579a9749ee29192c593ba4892c

Observation e4479443-39dc-4cf1-96c9-f96965848b79 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:07.543314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:07.543314Z digest=sha256:322b2b5e55fea374dec80b88bfdebf562aeee781fd5e2d2d9be9514644f7a7c7

Observation f6526d21-e822-4690-ae2d-1722ca13d9c4 · inbound

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards cites this paper.

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:05:54.829589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:04:10.226166Z digest=sha256:ca5098c4f2fc046982481f846a54d04de5947c32b4072190f791c092f3738d11

Observation 37b2e37a-518e-46b9-b839-f64d1398afe8 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:5be2840f629b8dfd87267a593ba95e0093bc3dcd9b4cbe1ec280be8906c13c08

Observation 52ea65f1-f84f-45ff-b0fc-ca81773892be · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging ChatBench: From Static Benchmarks to Human-AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:34.591235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:34.591235Z digest=sha256:facf27b88fabe1f40f06540ebb95a50bebb57aea08d773dce178cbb535b4744a