Pith. sign in

Paper Citation Record · LEDGER

Toward an Evaluation Science for Generative AI Systems

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.05336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.05336 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:35.042357Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f44943f3-cb76-4bfa-aeeb-bf86bf08a589 · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Toward an Evaluation Science for Generative AI Systems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.035975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.035975Z digest=sha256:0e557b85470860859de27274ca260707fe639bc292ee9c3414a2d527ced8da91

Observation 00bf41d9-f51f-42f5-b628-c437707f753c · inbound

Adultification Bias in LLMs and Text-to-Image Models cites this paper.

Adultification Bias in LLMs and Text-to-Image Models Toward an Evaluation Science for Generative AI Systems

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:53.536534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:53.536534Z digest=sha256:cab467aa6644792d955ce4f36a0798881cb1be1757735b2b347e75cb1534bb92

Observation 67010805-7da1-477f-b792-654e356efed1 · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Toward an Evaluation Science for Generative AI Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.091055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.091055Z digest=sha256:ce9ee269287e2f9cff0e15c74717e3f66c6446b3d10cc1a3ad20bd3f6d0ac506

Observation b2f468dc-1180-48b3-b383-c073f710ffd0 · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.341173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.341173Z digest=sha256:0dcf325a497e5b3a34f098de1f6c6483a9faa8dba0fc7c0423f33de8eb2a3474

Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.454150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.454150Z digest=sha256:c05705deaff622fb74f7d6db0856f0eca9113af8df2b9980a971c9b6f6adb062

Observation ec3598ac-b114-4790-9e78-9e1066a36c88 · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Toward an Evaluation Science for Generative AI Systems

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.745328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:8a8d709ace5de297afbfc0c034788db84fba92f190b8005e6dfe30ed436736ad

Observation 828d7cf6-478b-4f3c-8fac-08bb24bb253c · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Toward an Evaluation Science for Generative AI Systems

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.394750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.394750Z digest=sha256:66f85fa4008ba28f6007af97d4bb66df92a418a4e2bd56055c2340c5becbec21

Observation 30b72dfc-149c-4062-9953-cf756c0ac8fa · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.927757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.927757Z digest=sha256:5ee492b501dfaca294f1878249b1a64853e2241114a500bbc830215614f0342e

Observation d1c9d7d6-75c0-484d-90eb-1f5255204b8b · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI Toward an Evaluation Science for Generative AI Systems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:30.933003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:30.933003Z digest=sha256:232298dc884e73af93364941cd1f7294a45d9936f7125d1ca2a794e9e35fe084

Observation f6a775fc-c6c1-4af5-afe9-ea84f218b4db · inbound

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications cites this paper.

An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications Toward an Evaluation Science for Generative AI Systems

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:12:39.805750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:12:08.776876Z digest=sha256:c5f7539e50e6af4529fe0495e88b6aadab9b6bf71db3a8cd6c09f61417e73074

Observation 40b6930b-71b5-46b5-a009-952795f8cf87 · inbound

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models cites this paper.

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T21:55:20.294670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T21:52:26.944212Z digest=sha256:5428f3fde7ad762d417a174bb831047dd87f2483ec11811c32fffab1142d6fd2

Observation eff0a74f-f146-4145-8d11-e8b26b292e0d · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Toward an Evaluation Science for Generative AI Systems

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:05.167848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:4bbf8ac8c06bdbe4f68861c19d8665ceb377093178d9870cee1505a39192bbae

Observation 2e3f3e52-cd17-4f64-82e5-d5df6f151ba4 · inbound

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench cites this paper.

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench Toward an Evaluation Science for Generative AI Systems

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:25:23.882626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T06:20:32.174538Z digest=sha256:e80d79251809454346088b47184792edaddd9d197eb607ae7ee1115ad103f3d7

Observation ec26a659-b467-44f0-963d-666ff420317c · inbound

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research cites this paper.

Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research Toward an Evaluation Science for Generative AI Systems

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:40:16.529102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:38:00.482850Z digest=sha256:92910c338de581fad221e3b03e384a0f3c2de688a2171b7be7b264018917b08f

Observation c2da2f11-fed6-4085-aeee-ee23ff3e76bf · inbound

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing cites this paper.

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing Toward an Evaluation Science for Generative AI Systems

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:53:23.366771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:52:17.310177Z digest=sha256:f74ca0ecca209fb492f11b82f76aca6e60c5fcf59ebbec7a3fece3707955cd55

Observation d5deee66-5baf-4f67-bdf2-f7cf1ddf2c66 · inbound

BenCSSmark: Making the Social Sciences Count in LLM Research cites this paper.

BenCSSmark: Making the Social Sciences Count in LLM Research Toward an Evaluation Science for Generative AI Systems

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:51:08.685249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:02:15.446874Z digest=sha256:1f20a1114a0c9ff6f38d2ece5b34e4653721d9c214082c62a9237010a055d667

Observation 63ac6204-afa9-429b-b0d0-b83fed439a25 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.016305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:699ca36299b0a7548500987d7528a12574b190bb8a7dd72766a9bf2d07c6df4e

Observation 22ee5cdf-d680-43df-a261-32de1b528991 · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.753523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.753523Z digest=sha256:655c2f8796b935c4569ce0ad7ca27ab64a1276b68b5988b77046ddfe502b0c47

Observation 9d80f729-8b14-47c4-8441-e48c95297b47 · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders Toward an Evaluation Science for Generative AI Systems

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:a23a4ccea5d9f686ad3c701c569739fd34d82886ff7a67c306ac481b763e3af6

Observation 5fdd04f2-1eb7-4d1e-8b2a-1f4fae503f23 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Toward an Evaluation Science for Generative AI Systems

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T02:07:33.265081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:15127b161995a18f4698ac1e220ca85ff0219548a5850e758911ef4dc67fdf8c

Observation 6f756615-b9db-4248-ba42-eb1faca750f8 · inbound

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data cites this paper.

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data Toward an Evaluation Science for Generative AI Systems

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.494494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T05:41:56.435040Z digest=sha256:8a01a6e1a618632acfc8257b72ce653c56f2ea9fa8a4e56dbefdfd1968c84f7c

Observation af1ad875-3399-41f1-a4ac-2653c7c2e5cd · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Toward an Evaluation Science for Generative AI Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.614906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.614906Z digest=sha256:d494f8da2bc5a3abd3a7dd5039d2e3f5c2defe3c3cd79488fbc7adfa65b32323

Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · inbound

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) cites this paper.

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:35.042357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:44:35.042357Z digest=sha256:1d33fbb49622531de8e0084ba8b36246b3ac9163233d5a51b4835ba221920ed8