Pith. sign in

Paper Citation Record · LEDGER

State of What Art? A Call for Multi-Prompt LLM Evaluation

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2401.00595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.00595 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:01:31.410067Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9c1c3d5e-eab5-40a0-b9de-b33195ebaef3 · inbound

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models cites this paper.

Revisiting Sentiment Analysis for Software Engineering in the Era of Large Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:13:59.976576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T06:12:38.139005Z digest=sha256:92eadea3b392e249cdf4d9e093111627bee4afa49e95b484644ab7b81455de98

Observation a1f4b398-d6c4-4504-8d34-9ade7a41a5a4 · inbound

CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution cites this paper.

CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:57:16.378396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T20:57:16.327963Z digest=sha256:ecd7ea1cef1d5763e787ad187d2a9d6670971095c689aea09717a209b5513a45

Observation 003918c0-7574-4de8-9d64-a19d8f777951 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:34:43.111481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:46bafb3b2dcb04d0fa2ee8f82b8d4b133984223371c7baa08ce51b69f808fc0d

Observation e7fe15f2-ee82-42fd-9a94-9449eb690554 · inbound

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models cites this paper.

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:08:45.143108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-24T02:06:53.585629Z digest=sha256:cca8dc3f05cb841cf60dbec61a0261e5f76ce1cdf155d36001e91df68c28b831

Observation 004717d1-fbf6-4a2a-b95b-438caa3ec91d · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:44:49.776794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:68f431c7dc285d99a935d178017b4807ef35a22bfd071e3d557a65fca06a0221

Observation 67f2ded3-90f4-4e78-aa56-2b349ed0a4ee · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.571352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.571352Z digest=sha256:f31b58205ab125a3908240d804c237437af22950b882aafe12dfd2847f8e99d6

Observation 9b4d1b1f-c45e-4d0d-bf19-4c528f644464 · inbound

Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative cites this paper.

Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:47.279466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:47.279466Z digest=sha256:ca06f1dbe80b26bd36728e13e7a7e0aa557270160233e1a703688f4d14b7b2ab

Observation 810e7897-981a-4a76-afef-92a24701ff33 · inbound

LCTG Bench: LLM Controlled Text Generation Benchmark cites this paper.

LCTG Bench: LLM Controlled Text Generation Benchmark State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:45.078859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:53:45.078859Z digest=sha256:9313bd3eb9f201afbee1f373a2950db81cf54c2c99357d4ef071e7ce7f3d8621

Observation ebf17e8a-803a-46f9-9a86-1f73a7e839db · inbound

Evalita-LLM: Benchmarking Large Language Models on Italian cites this paper.

Evalita-LLM: Benchmarking Large Language Models on Italian State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:43:45.343766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:43:45.343766Z digest=sha256:9b634c048dc74d9b001180398cf45ea31447439443a99f37ff7767ce853dd8b4

Observation fff97b18-4589-4e30-a1df-0d7823c72029 · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.118860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.118860Z digest=sha256:2d5014c7da6ebaaff5c0e5a539e9d8d140ab9e514e4bb881903476effdd9fc28

Observation 8f9febc8-c580-4f34-857e-c867ebda099f · inbound

Personalizing Education through an Adaptive LMS with Integrated LLMs cites this paper.

Personalizing Education through an Adaptive LMS with Integrated LLMs State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:50:19.286087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:50:19.286087Z digest=sha256:50570cfbf92a4be8e8b82be4ccbccd1ddcbf1967695d4d152df7a8cc7b47d01b

Observation 13e0382f-c17c-4226-8a95-c333b5fe7729 · inbound

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks cites this paper.

MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:01:31.410067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:01:31.410067Z digest=sha256:021a8576c8de09e2d6877275a5f8b43279415f305a28b0c03ba036f2aa565680

Observation 83b4c6f7-85f5-47d4-817c-e8189db2350e · inbound

The Leaderboard Illusion cites this paper.

The Leaderboard Illusion State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T05:22:55.599051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:22:55.599051Z digest=sha256:1a6b49602a2b5d62d89f48327f93ab4805926893ee9a589497c6e10d615a41ca

Observation 853fbbd7-28bc-4a13-adbd-0e86dd15e784 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.072955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:205195afc426f4c102ac1a305c2db19652a0f98011c783b00ca728f0e9427f4d

Observation 7d7e74f0-daba-4236-9275-bd8ee43cafac · inbound

Predicting Performance of Symbolic and Prompt Programs with Examples cites this paper.

Predicting Performance of Symbolic and Prompt Programs with Examples State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T01:10:51.491532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T01:10:01.044650Z digest=sha256:510872f6fba2f08d8cd895ccc45672213324b249cda032c510e93b65aee83d88

Observation 6d86b257-40e5-4b84-81d0-4713d9f46cc5 · inbound

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks cites this paper.

SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:00.986796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T22:50:48.263600Z digest=sha256:0bb5195eaba9f5137d8ef73de59b99b7eed714f60682f6ff8d338fde6066c1f6

Observation 7d68ccd6-9171-44d5-b05f-9d322032e90c · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.257770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:20b01dd072699a1761151a104ea39cc976a238e1fefb8c198b2c1bdf947539e8

Observation 51592a10-a247-47d2-a4b1-a9a8a46924ea · inbound

Latent Confidence Alignment for LLM Self-Assessment cites this paper.

Latent Confidence Alignment for LLM Self-Assessment State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:39:41.782645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T11:21:33.744610Z digest=sha256:0ad01218fede6fb58a5c0f9d1bbca51be926fe44077ce9763d4985440cf52981