Pith. sign in

Paper Citation Record · LEDGER

Evaluating Generative AI Systems is a Social Science Measurement Challenge

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2411.10939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10939 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:40:07.216514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T03:41:00.103963Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 71c16f38-00e1-4ca8-9015-6337f292eea6 · inbound

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges cites this paper.

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-05T16:40:07.216514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:40:07.216514Z digest=sha256:4dd8477801773c9794abe8c7d59a34e63458780fcfda8389a2b553b4b5fc57f4

Observation 6f3764f2-2522-4499-89c9-7d0de8f765e4 · inbound

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop cites this paper.

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-05T14:38:55.710426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:38:55.710426Z digest=sha256:b1f0d4a29d1be4b83d6f89cbad7cf828a8901a49502cdbf1311aa52ed3845711

Observation ef6e2618-610c-446f-948e-752f28d4d8e0 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.153507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.153507Z digest=sha256:eae0381d547ca7c47f7f37cc2da7b015e3fd195264e412116e1bb35d4471d088

Observation ea55a5d3-78b6-4562-babf-f5bdb69bda7b · inbound

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation cites this paper.

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:21.097511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-07T16:09:27.944432Z digest=sha256:5d2e1215661d84bff3eac3b032e6d648d90f3808c535952b2b90dfd29c0eb463

Observation 2456370c-6165-4ec4-9876-0f4431a0d06b · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.754116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:724871328a5eb63f1f8f990a583fb32e2cfc43ecbb1752c2206d9bbd7d0cdfaa

Observation e537999c-a49b-449d-b305-8ca37468295c · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.223129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.223129Z digest=sha256:f8c81012591db7cdf3ba59b81dfcbc66380acca8e8ff1622c666be83b2ba2397

Observation 48a4c397-c291-4af2-ab40-b86e1b1d9eab · inbound

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory cites this paper.

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:43.866041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T19:40:59.534448Z digest=sha256:972dda4ec2e9543b80c158fcff9d9078e14baa0236f1274f084fb11460106738

Observation 36e9d93d-e554-4dd3-852f-f020103e3331 · inbound

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions cites this paper.

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:41:00.107255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T03:40:25.093487Z digest=sha256:064e4e5905aa5aa7ace476d591ffbf1fdf390533b216f7d57d818bd50dd28240