Pith. sign in

Paper Citation Record · LEDGER

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2502.00561.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00561 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:17:17.719091Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c69343e9-7ea8-4ff1-ba37-df93a101e6d9 · inbound

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents cites this paper.

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:56:38.062223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T16:55:47.922639Z digest=sha256:29a548bcf9c83cfbad877de5d4c83551445e154c0300f622c377cd4c118fef28

Observation 3a33c585-df94-4a14-a6c2-f96c21ae8e7d · inbound

Responsible Evaluation of AI for Mental Health cites this paper.

Responsible Evaluation of AI for Mental Health Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:50:54.931535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:50:33.070425Z digest=sha256:a8430f1a4cc027aef6b39888fe55066259bcd9b899fb4e56ae5e6a15e416ff31

Observation 729587ef-4445-41e6-a3ae-e208359ae64a · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:50:05.159501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:8248399d0fe77e9565815af43fe18818d3cea40df26e4e6762be1b7322a722c2

Observation 70749d48-bbd7-4928-bf76-ad85e797bbcd · inbound

Grounded Chess Reasoning in Language Models via Master Distillation cites this paper.

Grounded Chess Reasoning in Language Models via Master Distillation Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T21:26:43.149095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T21:26:43.149095Z digest=sha256:499241dd7e3f50f29c4fd64ec0c8a7834cebbcf2a7ac537bf051efe04c07bcbd

Observation 0191584d-4385-4212-ae50-19f6e3560e0f · inbound

RLHF May Not Reflect Genuine Preferences cites this paper.

RLHF May Not Reflect Genuine Preferences Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:40:46.318423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:39:08.880486Z digest=sha256:cd8d2d9897be0de78a6be403e97865b6c55cef878c92c86e6c969b9f530a9a0f

Observation e1116046-03f9-40e2-9144-9a0d98caa63a · inbound

From Ground Truth to Measurement: A Statistical Framework for Human Labeling cites this paper.

From Ground Truth to Measurement: A Statistical Framework for Human Labeling Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:30:58.092181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:09:43.892161Z digest=sha256:95ed67d741e8a5acb81c5bb8c772c6f3dc77c9ab3fe9f8f767b55e63a7d3613d

Observation 88589ab5-5cf6-4cf8-a6b4-3a499773718a · inbound

"I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI cites this paper.

"I Just Don't Want My Work Being Fed Into The AI Blender": Queer Artists on Refusing and Resisting Generative AI Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:25:22.668278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T12:21:20.533180Z digest=sha256:36b2e23a6c8f5037deb3a1f753401c6fce49336d94dd1189189e86fa5fd6a8ae

Observation fd24ef72-532d-402a-bfae-f5c653595ac9 · inbound

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios cites this paper.

Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:55:53.066086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T02:54:52.984994Z digest=sha256:41b783249aa7706e3fa1b1bf2be0df9aceecc76272b41bc5ed518fb464ee4070

Observation a1357ba7-1d60-473f-bc96-e9f837330683 · inbound

AI as a Tool for Simulation-Based Experiments in Literary Studies cites this paper.

AI as a Tool for Simulation-Based Experiments in Literary Studies Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T15:02:18.979706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:54:14.909995Z digest=sha256:36dfd2203904ccb9a002d9d97d2f27bbbef0b9f3b616590a79327dce54ba377f

Observation 9562a4fc-d8f1-4cda-ba93-c52ece092f06 · inbound

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer cites this paper.

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:27:40.352337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T13:13:19.690367Z digest=sha256:d85127b1cbdaf17b42a1f172faf92b1a39b6c9aa9d553eb339e2eab5061f4421

Observation 7948401f-e3a0-43b7-9be0-a9cd98d0453b · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:08:58.949519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:32545463e499040b7c9636b33f669cdf700bb36b6b63774a5aa0ecb4ec31ae0f

Observation 2f29de4b-0b44-496a-8c61-ed0646ca3c24 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.560871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.560871Z digest=sha256:b51f4750ef1ccfc61c287181eab753ebc9d3ecf5fd23c7625680208f7ebce1f9

Observation 13e8e9b3-298f-419d-a664-93ccfc27f6f8 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.248723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.248723Z digest=sha256:84147dc16c3eae20785899f691b938738f6c30ab23057d944c0bda2f014ecffc

Observation 077493da-6a1e-4c2a-a038-729646409e34 · inbound

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems cites this paper.

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T01:29:53.636089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:29:53.636089Z digest=sha256:ee8caf6dbfb1438dfeca7251bd4a5ecb78bbb46a4af4a8273e44dfcce1031d2a

Observation eab46b33-0a67-4492-931c-1842b1c368b1 · inbound

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions cites this paper.

Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-04T06:17:17.719091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:17:17.719091Z digest=sha256:64c4b865fc2c4ce964b215722587bbed8f5329682b2ce1c2d9a90a31772551f1