Pith. sign in

Paper Citation Record · LEDGER

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2505.10573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10573 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:32:51.345036Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation daec776e-2a51-4fdc-8afc-a0a952caa826 · inbound

No-Knowledge Alarms for Misaligned LLMs-as-Judges cites this paper.

No-Knowledge Alarms for Misaligned LLMs-as-Judges Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:51.345036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:32:51.345036Z digest=sha256:f990d243df58e95f243f19277a37fd7e00b40d7f2ab242b794f569689d26a695

Observation 08ef653f-f02e-4ba4-8052-2fa179d6da71 · inbound

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents cites this paper.

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:56:38.077538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T16:55:47.922639Z digest=sha256:960dfd2acda3c31a0c8e9a70590de4a537d0ecfad22d04ff29557b5b940b7ca4

Observation 5f164392-0269-4c83-904b-7292b3b99a71 · inbound

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems cites this paper.

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:35.604261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:41:49.137745Z digest=sha256:f6e9fc594f114cbd608c49670d05dac49a5c3c5e169c175d1268bc094eceef5e

Observation 83fd9b99-ed26-4611-b85a-f62bb04639da · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:05.155166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:f7394467c30e8fcb404d66eb8f74d49d79e019f57fb076f0dfa1b01963d0cb2b

Observation 562ca067-ebcf-4cb4-a100-7534dccc8cbb · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:19.961732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:3d11b2e8d91a0f567ec99109126d4eb32901c8544b662286b8f4e497ae21c4b4

Observation 8de174f7-0c80-4bd6-80df-08da94e56963 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.749430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:62ed6e0589e45d15601c313723cbd8af2cc1106ea5063c1adf6e5de185e812d8

Observation 4281925b-e72c-49a9-b7f6-fbbeccab1147 · inbound

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition cites this paper.

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:03.874897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T03:32:02.079410Z digest=sha256:af45745a86fa359eb4be18bd686957beadc9315321ad51494c2ad700b5aae08d

Observation f5a476b8-042f-4952-a799-58c1f61035a1 · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.632388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:3e89b68fdbabe7704e2b1e17eb20b5a19bd1d0ab00f4eb55c00bf98ff183296e

Observation 7af2a256-0a3b-4b0e-a449-f237cd74fd43 · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.475216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:85641dbd6c1a1ad0ce38425943c3abe660adbdaa1af59496661ec04519ad3633

Observation 029fdfae-f60a-4870-994f-e268d60cc0be · inbound

Quality Is Not a Safety Proxy Under Quantization cites this paper.

Quality Is Not a Safety Proxy Under Quantization Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:28.804330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T17:23:08.935056Z digest=sha256:c1806060a42dc680c3865f164e10f519ae10947f35d6a1719c7de5050a749c06

Observation b3ddd3a3-d757-49a8-b4be-7684a6319362 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.567081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:40fd800b6f5cfc048dc7a58cd6ea228964fc585af2596d55232cd19d786dcb0d

Observation 8be2027d-c037-43ae-a307-9ef24e6b0059 · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-13T06:19:53.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:19:53.291826Z digest=sha256:8ba53bb771093780175df2923bb1746ac5055761ca608a6c4ce4976a8fca58fe

Observation 30aa7dbf-48ba-4767-8b93-71d7c1898e5d · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T07:47:36.943427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:47:36.943427Z digest=sha256:9ccb9c3ea226a33029cfa97c1f5f5c220e28994cb38aadf0ee477ac7d8f6377e

Observation 0e40aeb1-1589-4e3d-a311-c25994277cd6 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.428402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.428402Z digest=sha256:65afaa2a07610e5ce79dae61c1e509f5e8fdaa01b048982bd56e04d504c2ba04

Observation b46796ac-f749-410b-a7b8-98a6f8f55403 · inbound

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems cites this paper.

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T01:29:52.661309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:29:52.661309Z digest=sha256:4704e4554e3c112283aeb98666fc5a60ae99b9477acdca4b5afa3a46a7baf357