Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:12:55.209578Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 11 inbound Pith citation observations for arXiv:2411.10939.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:12:55.209578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:42:53.504979Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T03:41:00.103963Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9c515841-4d64-40e6-9d02-41c64f387539 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge YouTube Hate Speech Policy
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91d834e9-b16e-4393-9a37-040723445a03 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement validity: A shared standard for qualitative and quantitative research
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b064fff7-1647-44f2-88fe-f573527e55fc · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Content analysis in communication research, 1952
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32835267-c04d-4a63-b5ff-2251fa531963 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Making Intelligence: Ethical Values in IQ and ML Benchmarks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add1c81a-962e-4128-9e3f-b6108121cbff · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Sociolinguistically Driven Approaches for Just Natural Language Processing
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 101e6a60-6366-44ab-8b14-386973d310db · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Language (technology) is power: A critical survey of ‘bias’ in nlp
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bc8f043-44c9-4829-a536-0649cd83b47f · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Stereotyp- ing norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 38afce52-aa80-413e-8b11-abab2ad3eda6 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986f4050-05e4-4824-b4b9-a11795e3d20b · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Feder Cooper, Ellen Abrams, and NA NA
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b5c918-b4b2-44d1-852e-9292b1a017a6 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Report of the 1st Workshop on Generative AI and Law
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8467d862-bdb5-40a0-9609-f968ebab6d99 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Representational harms through the lens of speech act theory
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e2a104-d303-4c69-9809-96810cc6ed3b · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Construct validity in psychological tests
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5f8e446b-2916-46e3-947c-dbfd020ed0e0 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Measurement and fairness
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d8dd5ef3-fb54-4dd9-a304-c9c7657961a8 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648da852-aae6-476c-b7ce-e6222b1123c2 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Hate speech in public discourse: A pessimistic defense of counterspeech
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 22e50b36-cf29-499a-b48c-59a449cd2b04 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Vera Liao, Alexandra Olteanu, and Ziang Xiao
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f209101-db34-418b-a817-9002c6ae1989 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Validity and washback in language testing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1fe3b165-cfb7-48bb-8116-eddcb5d620e2 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Privacy is an essentially contested concept: a multi-dimensional analytic for mapping privacy
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6ab75ab-7f1b-430b-8c02-5cd9e255dc96 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Mulligan, Joshua A
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7ede09c7-c4b9-41b5-82bf-25ba3bad6834 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge StereoSet: Measuring stereotypical bias in pretrained language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0a1e00-ce68-4a96-b888-926f9f26f5e2 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0736818e-a22b-4792-b0b0-ccbbdf97d5c3 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fecc7a41-e79f-4c13-bd96-80764b4160f1 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Red Teaming Language Models with Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba9a1d0-413e-4320-abf7-c05e07483ebd · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge Evaluating General-Purpose AI with Psychometrics
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51eb7304-e85e-4cce-940b-4126f761dd64 · outbound
Evaluating Generative AI Systems is a Social Science Measurement Challenge The nature and origins of mass opinion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d9a316a6-17c7-4f4c-8413-eac0b080776c · inbound
Adultification Bias in LLMs and Text-to-Image Models Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc84554f-f5ac-46c7-9ab3-8b4ce16d6715 · inbound
Correlated Errors in Large Language Models Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d511600-3fcc-4607-90f9-c6c2637fb8cc · inbound
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c16f38-00e1-4ca8-9015-6337f292eea6 · inbound
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f3764f2-2522-4499-89c9-7d0de8f765e4 · inbound
Understanding, Protecting, and Augmenting Human Cognition with Generative AI: A Synthesis of the CHI 2025 Tools for Thought Workshop Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6e2618-610c-446f-948e-752f28d4d8e0 · inbound
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea55a5d3-78b6-4562-babf-f5bdb69bda7b · inbound
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2456370c-6165-4ec4-9876-0f4431a0d06b · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e537999c-a49b-449d-b305-8ca37468295c · inbound
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a4c397-c291-4af2-ab40-b86e1b1d9eab · inbound
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 36e9d93d-e554-4dd3-852f-f020103e3331 · inbound
Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions Evaluating Generative AI Systems is a Social Science Measurement Challenge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.