Pith. sign in

Paper Citation Record · LEDGER

Evaluating Generative AI Systems is a Social Science Measurement Challenge

As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 11 inbound Pith citation observations for arXiv:2411.10939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10939 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:12:55.209578Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:42:53.504979Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T03:41:00.103963Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c515841-4d64-40e6-9d02-41c64f387539 · outbound

This paper cites YouTube Hate Speech Policy.

Evaluating Generative AI Systems is a Social Science Measurement Challenge YouTube Hate Speech Policy

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-12T19:12:55.571058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.096303Z digest=sha256:6c1569176533fc53498141fa7761d2b1a3af0da677df4f947f3135ccd2835850

Observation 91d834e9-b16e-4393-9a37-040723445a03 · outbound

This paper cites Measurement validity: A shared standard for qualitative and quantitative research.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement validity: A shared standard for qualitative and quantitative research

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.101173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.101173Z digest=sha256:18202f2faa63c7df86c071453c6b5ec79b0d878eee73b3be7e0753a68a07cde8

Observation b064fff7-1647-44f2-88fe-f573527e55fc · outbound

This paper cites Content analysis in communication research, 1952.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Content analysis in communication research, 1952

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.779337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.105802Z digest=sha256:aabcddc41353a1bbcb0f7e61e734ae954605b370aa238e2c601f3c93e5361405

Observation 32835267-c04d-4a63-b5ff-2251fa531963 · outbound

This paper cites Making Intelligence: Ethical Values in IQ and ML Benchmarks.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Making Intelligence: Ethical Values in IQ and ML Benchmarks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.110397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.110397Z digest=sha256:580b4f635493f581bdc827f6015177bb480d479f44b05afa7ce08892ff0e505c

Observation add1c81a-962e-4128-9e3f-b6108121cbff · outbound

This paper cites Sociolinguistically Driven Approaches for Just Natural Language Processing.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Sociolinguistically Driven Approaches for Just Natural Language Processing

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:12:55.763447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.114777Z digest=sha256:7f20e8be327fcce225926595f0201dc0f1253b83451989dd4eea44ac659548fe

Observation 101e6a60-6366-44ab-8b14-386973d310db · outbound

This paper cites Language (technology) is power: A critical survey of ‘bias’ in nlp.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Language (technology) is power: A critical survey of ‘bias’ in nlp

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.748677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.120069Z digest=sha256:5b42ade8f2e94aa2826e4ac8b0bdd191bb1af3b1e800b4382aae6732ecabdca2

Observation 1bc8f043-44c9-4829-a536-0649cd83b47f · outbound

This paper cites Stereotyp- ing norwegian salmon: An inventory of pitfalls in fairness benchmark datasets.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Stereotyp- ing norwegian salmon: An inventory of pitfalls in fairness benchmark datasets

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.733739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.125557Z digest=sha256:70fd73a84b56f6bec5e888a62331b69e6bf1b7cc86ff31f038339e3c4134a4e1

Observation 38afce52-aa80-413e-8b11-abab2ad3eda6 · outbound

This paper cites DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.131216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.131216Z digest=sha256:aa8ce8a4f06982208f9494a8244bea3b5433cf2696af1961c92bd30061204148

Observation 986f4050-05e4-4824-b4b9-a11795e3d20b · outbound

This paper cites Feder Cooper, Ellen Abrams, and NA NA.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Feder Cooper, Ellen Abrams, and NA NA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.136919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.136919Z digest=sha256:244778a408af3e765620f0ed988d687b73ca1e1a3680e908c645ee559b9e99a3

Observation 90b5c918-b4b2-44d1-852e-9292b1a017a6 · outbound

This paper cites Report of the 1st Workshop on Generative AI and Law.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Report of the 1st Workshop on Generative AI and Law

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.142295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.142295Z digest=sha256:b0b133d57b3232357d60729f02d39eafa4b95e85fc7621fdaec315c4781ff81a

Observation 8467d862-bdb5-40a0-9609-f968ebab6d99 · outbound

This paper cites Representational harms through the lens of speech act theory.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Representational harms through the lens of speech act theory

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.147508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.147508Z digest=sha256:15758dcc831e3b75ac151354ccfd2bb9e59f4a13d40aa06a7d07ad9c3cf1e4b8

Observation 75e2a104-d303-4c69-9809-96810cc6ed3b · outbound

This paper cites Construct validity in psychological tests.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Construct validity in psychological tests

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.707824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.151808Z digest=sha256:6a94c6289ce6fd187e938a8d9000ccb6d89788dabb473d50b26097be69c0277a

Observation 5f8e446b-2916-46e3-947c-dbfd020ed0e0 · outbound

This paper cites Measurement and fairness.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement and fairness

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.692482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.156530Z digest=sha256:8ad6e828f1025236830bdfe93df18efda38aa8adb1ee1df719047174fe81e893

Observation d8dd5ef3-fb54-4dd9-a304-c9c7657961a8 · outbound

This paper cites ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation.

Evaluating Generative AI Systems is a Social Science Measurement Challenge ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.160667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.160667Z digest=sha256:2e179e65add21078b2045b4b180e5521951375499a283bdaabea4b9f3c74a83c

Observation 648da852-aae6-476c-b7ce-e6222b1123c2 · outbound

This paper cites Hate speech in public discourse: A pessimistic defense of counterspeech.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Hate speech in public discourse: A pessimistic defense of counterspeech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.676896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.165191Z digest=sha256:d607d6d8b0db47b81f316bb231a86d528d3342e9fd2519932c9ad5718fc7b86e

Observation 22e50b36-cf29-499a-b48c-59a449cd2b04 · outbound

This paper cites Vera Liao, Alexandra Olteanu, and Ziang Xiao.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Vera Liao, Alexandra Olteanu, and Ziang Xiao

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.662480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.169333Z digest=sha256:688f6976aa70dc0403550a1156543ae52c3f2591069a4c113ba87f4567048901

Observation 2f209101-db34-418b-a817-9002c6ae1989 · outbound

This paper cites Validity and washback in language testing.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Validity and washback in language testing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.647368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.173553Z digest=sha256:12bedc5ef8192f7a5bc813f2baf17af859029b6c98d567c0d1e888bc46427e34

Observation 1fe3b165-cfb7-48bb-8116-eddcb5d620e2 · outbound

This paper cites Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.632568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.177711Z digest=sha256:f36b102b7c23fcae207eb9fc702e9289b28a9ca5cadcbb352556f2b9f972ca00

Observation a6ab75ab-7f1b-430b-8c02-5cd9e255dc96 · outbound

This paper cites Mulligan, Joshua A.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Mulligan, Joshua A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.618323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.181927Z digest=sha256:24815f39924e2b087fba85979153d6c5326f3568b04dd1a084636a967fe9a0bb

Observation 7ede09c7-c4b9-41b5-82bf-25ba3bad6834 · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge StereoSet: Measuring stereotypical bias in pretrained language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.185915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.185915Z digest=sha256:d3c02dae6e47a97ea74edbf1732145dd849283f3e702add722f996b492e69373

Observation bf0a1e00-ce68-4a96-b888-926f9f26f5e2 · outbound

This paper cites CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.190662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.190662Z digest=sha256:d7828e0659243f114cb2dbdfb02adb828f40dd011ed1f1035ef819f22468c55d

Observation 0736818e-a22b-4792-b0b0-ccbbdf97d5c3 · outbound

This paper cites Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.602679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.195061Z digest=sha256:27f7e29eca0aad2dce9bcfc8b57f4bf1f43abfad623f48d1a45df1bf6b8c661a

Observation fecc7a41-e79f-4c13-bd96-80764b4160f1 · outbound

This paper cites Red Teaming Language Models with Language Models.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Red Teaming Language Models with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.199612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.199612Z digest=sha256:b2c8b7e933db3ea195d97440ee30a34cee5dd62318355072c8439b79c0ee79c9

Observation fba9a1d0-413e-4320-abf7-c05e07483ebd · outbound

This paper cites Evaluating General-Purpose AI with Psychometrics.

Evaluating Generative AI Systems is a Social Science Measurement Challenge Evaluating General-Purpose AI with Psychometrics

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:12:55.204761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:12:55.204761Z digest=sha256:38fc16c3620654fb2aabd6c12829071d91d9211942f89c2e5e755fb53bf0229f

Observation 51eb7304-e85e-4cce-940b-4126f761dd64 · outbound

This paper cites The nature and origins of mass opinion.

Evaluating Generative AI Systems is a Social Science Measurement Challenge The nature and origins of mass opinion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:12:55.587253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:12:55.209578Z digest=sha256:69967696cd23f7f96e45a59a4ceaa48c91ec429132f69aec04f7712c241d4f44

Pith citing papers

Observation d9a316a6-17c7-4f4c-8413-eac0b080776c · inbound

Adultification Bias in LLMs and Text-to-Image Models cites this paper.

Adultification Bias in LLMs and Text-to-Image Models Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:53.504979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:53.504979Z digest=sha256:ce1a74b61f25e0bd6b3482337fa831a383257dcfa6f84dd573837c240270270f

Observation dc84554f-f5ac-46c7-9ab3-8b4ce16d6715 · inbound

Correlated Errors in Large Language Models cites this paper.

Correlated Errors in Large Language Models Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.330863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.330863Z digest=sha256:5206e29d8cf2da50d3e4b2d1b1072ee767819a8134b7de03cd8b5f522805ebe0

Observation 5d511600-3fcc-4607-90f9-c6c2637fb8cc · inbound

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks cites this paper.

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:11.241009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:11.241009Z digest=sha256:190676d5fead497100c5c90551eb2d1ef16e58dfa522dbfd5229631683f5ce72

Observation 71c16f38-00e1-4ca8-9015-6337f292eea6 · inbound

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges cites this paper.

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-05T16:40:07.216514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:40:07.216514Z digest=sha256:dd1f30a18b0527d2e91cd23c91bee003a233dcada7381295c92be66441965c2b

Observation 6f3764f2-2522-4499-89c9-7d0de8f765e4 · inbound

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop cites this paper.

Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-05T14:38:55.710426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:38:55.710426Z digest=sha256:848be92174c061ed8bbfee342a8c32b8516260a251bdb36605517f9e19d24dd0

Observation ef6e2618-610c-446f-948e-752f28d4d8e0 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:57.153507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:57.153507Z digest=sha256:437918bf1d82fbc5fe3ac07dfdf971b219553fa6cf0a78c6c4d46caa212380ac

Observation ea55a5d3-78b6-4562-babf-f5bdb69bda7b · inbound

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation cites this paper.

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:51:21.097511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-07T16:09:27.944432Z digest=sha256:040bb3942452fe3615410957d7bcc846c2295b68d270b779e53d381120b9cd79

Observation 2456370c-6165-4ec4-9876-0f4431a0d06b · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:27.754116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:34:55.538935Z digest=sha256:eeda594b7d8ceb32e6e046a5f708e6f38cfae20f5d16038f075e8e740faf5004

Observation e537999c-a49b-449d-b305-8ca37468295c · inbound

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World cites this paper.

From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T14:22:26.223129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:22:26.223129Z digest=sha256:6857d291da9691e05956649d891ea13cb27391d43a300b74d43f3f0ec6710f9f

Observation 48a4c397-c291-4af2-ab40-b86e1b1d9eab · inbound

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory cites this paper.

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:43.866041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T19:40:59.534448Z digest=sha256:b1c50ab1839b2060f01f92f87bba04539618c55db7eaa0175459787e785120c0

Observation 36e9d93d-e554-4dd3-852f-f020103e3331 · inbound

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions cites this paper.

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:41:00.107255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T03:40:25.093487Z digest=sha256:0a1f891a79ea4fabc2d6460b3f410da8f0eae396fb4fac2b47383713d4f6eadf