Pith. sign in

Paper Citation Record · LEDGER

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 17 inbound Pith citation observations for arXiv:2505.10573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10573 v4

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:09.243560Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.792055Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b5bad860-1940-46a0-ba56-9177e3dcbb68 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.909820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.042124Z digest=sha256:ff989b9bad6629f1f2430e762897bd8d33212511ad91ef9b6e7a0adc8bde8031

Observation 6c6e53c8-ae81-4a26-8916-659f9d5f533d · outbound

This paper cites Concurrent validity refers to a test’s agreement with a validated measure applied at the same time under the same conditions.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Concurrent validity refers to a test’s agreement with a validated measure applied at the same time under the same conditions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.894802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.047951Z digest=sha256:f85828c27d1c1a396249bedabedcf2d8918a0daa5f9c153090d51d3fbe490577

Observation bc0293f4-a926-402b-8bab-eca02ff9836f · outbound

This paper cites Beyond these core types, external validity refers to the extent to which a study’s findings can be generalized beyond its specific conditions.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Beyond these core types, external validity refers to the extent to which a study’s findings can be generalized beyond its specific conditions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.878640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.052646Z digest=sha256:996ecbf46f21d6e480e96ad99bbc7c82439ffe55797968d28ee6aff6c5b596d9

Observation 83e5ebff-6913-4f21-89fd-5451ad4a574e · outbound

This paper cites GPQA includes diverse topics within its disciplines.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation GPQA includes diverse topics within its disciplines

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.752420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.095464Z digest=sha256:1a2893b360a4e3a1243c61a0860b01c380668959c1712956a1ac1fd325f52a22

Observation 0dd2dab9-e32d-45b0-a2c1-7a70568ef710 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.846513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.064483Z digest=sha256:fc459733f7347f92d1c0f44a5176297893053bf06643299da5d73881ce59f245

Observation 8244b055-cbb7-4896-91b5-18b7f872ac56 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.830532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.070215Z digest=sha256:5562978edb1048201406885e73efa79de4c053a50bf889d33eddf8ab0e19f49d

Observation 4c0b3d3e-a9b9-4aa1-a0a6-feff8b86de02 · outbound

This paper cites Description of dataset.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Description of dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.815245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.075100Z digest=sha256:d6cf2f9d88b5688219f6b94ee937b95e4e3f1da39cc761d59b394b350e1e2317

Observation 0a31adea-0c59-4d9e-a898-6e5d388244d6 · outbound

This paper cites The performance gap between experts and non-experts confirms the questions assess specialized knowledge.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation The performance gap between experts and non-experts confirms the questions assess specialized knowledge

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.799335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.080418Z digest=sha256:4d688980b3da0f72630eadb9a197a86b59e35e838072d47d498ad36b22939c01

Observation 9828cb73-539e-4347-9dc8-fd9687cc8dd5 · outbound

This paper cites • Weakness: Criterion validity could be stronger with comparisons to other specialized science Q/A benchmarks.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Criterion validity could be stronger with comparisons to other specialized science Q/A benchmarks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.783185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.086156Z digest=sha256:0d5e3b6d8a158b38c3da57a28a7834ab2cf90fb77ee1d8cbe4f04ef48dd5d2d9

Observation a5138293-b7fe-4318-ad4d-b26ddd1efe8e · outbound

This paper cites However, models have quickly improved in this benchmark 5.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation However, models have quickly improved in this benchmark 5

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.736963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.099921Z digest=sha256:784f8a937a23773d59d2868a2af57ff7cd8f65bbffce5383d47a0f9e195ea548

Observation a3f3a6e7-cce7-43f0-99bd-4dbf9fa4aaa1 · outbound

This paper cites Non-expert performance gap supports specialization.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Non-expert performance gap supports specialization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.720866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.104310Z digest=sha256:9ad978ed913b92486c48e014c48f5a4d8eb694504b9d6c971331667c4557d428

Observation d6562135-4dbc-4048-b51b-4b00c6736247 · outbound

This paper cites AI-expert performance gap reinforces benchmark credibility.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation AI-expert performance gap reinforces benchmark credibility

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.703482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.108828Z digest=sha256:6cbe7be00ca0477b439beee3be33d05df35f77231af133b39b5614fa409bd185

Observation 929e0e84-e2a4-4e71-ad94-b2911b23f10a · outbound

This paper cites specialized scientific knowledge,.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation specialized scientific knowledge,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.686612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.112823Z digest=sha256:e4d964df2644a896453ec72d6966946d3f641735103b052b6b9208c0711fbc03

Observation 9a63cc4a-52ac-41cb-8eb2-c3d35e56363c · outbound

This paper cites Coverage across multiple subfields increases generalization within biology, physics, and chemistry.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Coverage across multiple subfields increases generalization within biology, physics, and chemistry

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.670831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.117799Z digest=sha256:9c2ff23bc3246d1b2ff0952fcbdba4e6c831fceff62720bad95c83d6e0830c33

Observation a2ed36c2-9f0a-446a-aaee-e6c9af2bb0a5 · outbound

This paper cites • Weakness: Risk of overgeneralization—high scores may be misinterpreted as broad scientific expertise beyond tested domains.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Risk of overgeneralization—high scores may be misinterpreted as broad scientific expertise beyond tested domains

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.655646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.122316Z digest=sha256:1f13a815731741ed95b478a94e6ffe323acd5713754811a7e6814b6eea9086ca

Observation 2836a1c7-22b2-44b4-b9a2-a99b2737b63b · outbound

This paper cites • Weakness: Multiple-choice format limits assessment of forms of reasoning like logical deduction, or abstract problem-solving.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Multiple-choice format limits assessment of forms of reasoning like logical deduction, or abstract problem-solving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.640805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.127364Z digest=sha256:9e4cda10f4558dc43c11523ae103a264092b805aa7cf01b5cfc66678a2fed46a

Observation 25e56cae-ba12-45a4-b0dc-884c5f15a0fc · outbound

This paper cites • Weakness: GPQA tests factual and applied knowledge rather than abstract reasoning skills.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: GPQA tests factual and applied knowledge rather than abstract reasoning skills

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.625293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.132066Z digest=sha256:e4ff4bff9e719582b542ac3b6eb7d2653d6dcefbb29fa1d19bfaf775acba7f7c

Observation 857b39e9-3377-40e3-ac2d-333ab09c7b41 · outbound

This paper cites Additionally, the dataset can distinguish between human experts and non-experts.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Additionally, the dataset can distinguish between human experts and non-experts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.608547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.136495Z digest=sha256:b35870ca485fcb5b09bb77229c1089544c6ec9f4a535f14cbb12de15f5cb5ea0

Observation e212228e-88f6-4821-800d-e25cebbb0ab3 · outbound

This paper cites • Weakness: Reasoning should generalize across domains, but GPQA only includes three scientific fields.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Reasoning should generalize across domains, but GPQA only includes three scientific fields

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.592310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.142141Z digest=sha256:16e29f5b3d61b9c6c32135dce3f0756290d9609877073b092e0eea18e3570077

Observation 8eee1c0d-a64f-46e5-8a27-58db244e3300 · outbound

This paper cites reasonable.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation reasonable

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.576820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.147827Z digest=sha256:56e7bc9364cefc13436e07bcbbeee3562b8c312410b4701ee9afbe54f737c46e

Observation c8d45239-9f9f-4e5b-9b9d-416a08fafc9b · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.561200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.152988Z digest=sha256:0c5f9e638d02f1e051aeeeba72a68881063cfc602336f17d7c1e7c86370c6eb1

Observation 98d92b07-d1a7-4b7b-a0eb-93d007f90e04 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.545479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.158163Z digest=sha256:3e959e61bbba93a64d0901c2dad7521eded8d45e71825268ae93a906b59e9462

Observation defe7196-1ec0-4fec-8898-98ea0c82be6c · outbound

This paper cites Description of dataset.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Description of dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.530163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.162957Z digest=sha256:ec46ea33bc867e199683c46cecb2f1ff8566f42b381b59d2c10556988e2e2577

Observation 3435c0a1-4354-464b-9864-10167e30267e · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.514188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.168052Z digest=sha256:d49455a510615df1deaaff37ee429427e5c417baaedb5623a61b22b1c52a1742

Observation 314e82e5-7ec6-4e59-9974-609555249442 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.496870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.174221Z digest=sha256:4d210646871ce2a4789029ba9f408a61a2d076b72cdd29db685ce760338779c4

Observation 374a72bb-ee27-4041-824a-a01016770d2c · outbound

This paper cites Note, this is not about trained model performance (e.g., Recht et al.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Note, this is not about trained model performance (e.g., Recht et al

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.481088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.181872Z digest=sha256:47e3399111d97662264dc235ab9ab6332fe3dba6279752c0bfb9d5fd0db5693e

Observation 21c4556e-5932-40d8-a431-35f7e403c5b9 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.768035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.187157Z digest=sha256:231da65edab4d85c546fec90ec209b39caf9a1871a6b2ac88d37b5e11149348d

Observation 38c7cba0-3a9c-4a91-b54b-17bf7a1d5bb5 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.464323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.191970Z digest=sha256:5c313a049cd2266b756d9e9dfa423c5be7cf3e57e1338382ea370715ecdbdad2

Observation d63c8b8c-094a-4fa2-a226-cc1a65b02cd1 · outbound

This paper cites • Weakness: It may not comprehensively represent features present in non-natural or synthetic environments, nor fully capture abstract contextual cues.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: It may not comprehensively represent features present in non-natural or synthetic environments, nor fully capture abstract contextual cues

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.448402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.197238Z digest=sha256:97798b0061b332bfc1a6d5cab2da86c28a31215772cd125ce74c6f1b20f75640

Observation c4224d2f-48a6-4596-a71e-77320d685369 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.432030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.202113Z digest=sha256:92c7eb47d4bc8e5e8d21f1689e05048af3da7699c1cb30207d79a5139fc48d54

Observation 0bbcd34a-112f-429f-90de-40cda033317f · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.413825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.208939Z digest=sha256:915119522761206855c16f7583cd1cd77d54e85a9563e4de73cb572fd82a3265

Observation 36d13892-7162-4b15-a966-f65264cd1d9b · outbound

This paper cites • Weakness: The degree of generalizability across the span of domains (e.g., synthetic or non-natural images) remains to be fully validated.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: The degree of generalizability across the span of domains (e.g., synthetic or non-natural images) remains to be fully validated

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.397120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.215052Z digest=sha256:533c0dc2f973b99fa716d424211dbc8969e9ed2257b6e65b5bb9633d0e3e7e79

Observation 664748a4-c781-4548-b7e3-ea9d71d1a539 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.378158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.219790Z digest=sha256:b1f0e37f04228576b88af075577afcc173492145cef6ce09c34328daad2ebae6

Observation aa0bbba2-7fc1-480d-9532-5e2fed1a82d2 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.358038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.224246Z digest=sha256:ab24a285d140d28eebf61fbcfebe11efec275024a24f81611f5b85c7732bd31a

Observation cb75f443-81d8-4693-bbe5-c05a2b03ea3b · outbound

This paper cites • Weakness: There is limited evidence that high performance on this narrow task reliably predicts the broader and deeper aspects of overall visual understanding.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: There is limited evidence that high performance on this narrow task reliably predicts the broader and deeper aspects of overall visual understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.339885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.229196Z digest=sha256:3bba0455f3ae6d28a9ee0606dd77459820947a86a6db6ed9b4a33f735248a01c

Observation 10dcecfa-05b2-42d2-8706-8d2e1f3ae016 · outbound

This paper cites an unresolved cited work.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:49:09.322957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.233602Z digest=sha256:cbce982085d3f380cd4a2516643f92ff5b10024f9729b21b74e746de11903b35

Observation 3b6366ec-350a-4409-9dc1-dcb48df856a5 · outbound

This paper cites • Weakness: Its ability to generalize to tasks requiring integrated reasoning, spatial aware- ness, and contextual interpretation remains unconfirmed.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Its ability to generalize to tasks requiring integrated reasoning, spatial aware- ness, and contextual interpretation remains unconfirmed

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.303974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.238859Z digest=sha256:e982309bc544c28560ceaf3f97cb454f831e440c1a877c02370fb8237b41dbf4

Observation 87eea91f-40c2-4388-96a2-61e11d7dcc78 · outbound

This paper cites • Weakness: High classification accuracy might be erroneously interpreted as evidence of complete visual understanding, potentially misleading real-world applications.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: High classification accuracy might be erroneously interpreted as evidence of complete visual understanding, potentially misleading real-world applications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.285385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.243560Z digest=sha256:d4a4cdb7aaed733dfd607c254643df1accd1bb2c698d29d0fa51b8e79584bbd4

Observation 3c7ed114-8a5e-403c-aefb-8842c53b4342 · outbound

This paper cites blocks world.

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation blocks world

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:49:09.862251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T21:49:09.057629Z digest=sha256:38123f4dc75497474e5521379d9228aaff708e27e5bafae7057dd157b2b566c4

Pith citing papers

Observation 326f22e5-6d1b-4715-85f1-3cc44efa40da · inbound

UQ: Assessing Language Models on Unsolved Questions cites this paper.

UQ: Assessing Language Models on Unsolved Questions Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T17:07:19.792055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:07:19.792055Z digest=sha256:c2d7f238d3d5e26ba7e0d4ee196164b4b979226be3be9d71e942a7bfd8ca19a0

Observation daec776e-2a51-4fdc-8afc-a0a952caa826 · inbound

No-Knowledge Alarms for Misaligned LLMs-as-Judges cites this paper.

No-Knowledge Alarms for Misaligned LLMs-as-Judges Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:51.345036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:32:51.345036Z digest=sha256:ff6eb92c29628302f29dd0af69f0a205b89e9fa8727d1bb5dfebf21750237aee

Observation 08ef653f-f02e-4ba4-8052-2fa179d6da71 · inbound

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents cites this paper.

Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:56:38.077538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T16:55:47.922639Z digest=sha256:ea7094a746180074522dd4b4a0886472852b646b76792e86b626627800c0a6ae

Observation 5f164392-0269-4c83-904b-7292b3b99a71 · inbound

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems cites this paper.

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:46:35.604261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:41:49.137745Z digest=sha256:d3a770e950ae4f1f3edb755fb4d63b704572c21809cf18d095af92accdae3a34

Observation 83fd9b99-ed26-4611-b85a-f62bb04639da · inbound

Making AI Evaluation Deployment Relevant Through Context Specification cites this paper.

Making AI Evaluation Deployment Relevant Through Context Specification Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:05.155166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T14:46:09.944168Z digest=sha256:8b158865dd1a763b8796d09919802d2476233d7c2846b5ce9240f98457604c4f

Observation 562ca067-ebcf-4cb4-a100-7534dccc8cbb · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:19.961732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:cf404ec0a08cce27d29ca3645e10348daacbd38ab04e09b1b5dc19e330d50ddd

Observation 8de174f7-0c80-4bd6-80df-08da94e56963 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.749430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:1d0b170e9b9276c9a51437fbf9aa485c25756618c175b5f50b0fa8a996bc57e4

Observation 4281925b-e72c-49a9-b7f6-fbbeccab1147 · inbound

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition cites this paper.

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:03.874897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T03:32:02.079410Z digest=sha256:c277b323e211491c05d498bcb31ef7cb26584ea86ff9f15c699ce041b1d92f8e

Observation f5a476b8-042f-4952-a799-58c1f61035a1 · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.632388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:a35a7235599907dd0b47bb80ae6a0712237273995596ef0d3a71c662bcff8fcb

Observation 7af2a256-0a3b-4b0e-a449-f237cd74fd43 · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.475216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:daa09565001035eda2b4c40570d2f3e6abc5209e1049ddce1233ad8227d97daf

Observation 029fdfae-f60a-4870-994f-e268d60cc0be · inbound

Quality Is Not a Safety Proxy Under Quantization cites this paper.

Quality Is Not a Safety Proxy Under Quantization Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:28.804330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T17:23:08.935056Z digest=sha256:d0d88d7f969c44a501d386dcf3ae3a57d7b6a7d4a3d81eba73d23dec8b4d5dcc

Observation b3ddd3a3-d757-49a8-b4be-7684a6319362 · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.567081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:7ad16874ee0d13c493588710c343c84aa8f02f89e8c9d4cefd546b32bd7ae905

Observation 8be2027d-c037-43ae-a307-9ef24e6b0059 · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-13T06:19:53.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T06:19:53.291826Z digest=sha256:6687a5145a432ba2c304d700a921789436aae009352091ff6189c382bef4c48f

Observation 30aa7dbf-48ba-4767-8b93-71d7c1898e5d · inbound

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education cites this paper.

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-02T07:47:36.943427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:47:36.943427Z digest=sha256:f1d7a933590eb388e84a656a21fffcc206577b164b52977161136bdc27efa322

Observation 0e40aeb1-1589-4e3d-a311-c25994277cd6 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:03.428402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:03.428402Z digest=sha256:62d780ecd551fb27f078ff4548e24446b5959cc78a6858f67085c19dd1178076

Observation b46796ac-f749-410b-a7b8-98a6f8f55403 · inbound

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems cites this paper.

On the Convergent Validity of Offline Evaluation Designs for Recommender Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T01:29:52.661309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:29:52.661309Z digest=sha256:32b39fedd400f7e69c1fd6ad72ed313e3e06d36b7e265fa0d48b14fb48927b83

Observation 387b1c4c-1805-4ae3-83f0-3715b414ce0d · inbound

Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation cites this paper.

Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T13:42:02.902928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:42:02.902928Z digest=sha256:d83f1c02bf954a18b59e88a28f783ca1033a116e1f391713ceb20ad4a13f698d