Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:09.243560Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 17 inbound Pith citation observations for arXiv:2505.10573.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:49:09.243560Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:07:19.792055Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
39 of 39 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation b5bad860-1940-46a0-ba56-9177e3dcbb68 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c6e53c8-ae81-4a26-8916-659f9d5f533d · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Concurrent validity refers to a test’s agreement with a validated measure applied at the same time under the same conditions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc0293f4-a926-402b-8bab-eca02ff9836f · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Beyond these core types, external validity refers to the extent to which a study’s findings can be generalized beyond its specific conditions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 83e5ebff-6913-4f21-89fd-5451ad4a574e · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation GPQA includes diverse topics within its disciplines
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0dd2dab9-e32d-45b0-a2c1-7a70568ef710 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8244b055-cbb7-4896-91b5-18b7f872ac56 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4c0b3d3e-a9b9-4aa1-a0a6-feff8b86de02 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Description of dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0a31adea-0c59-4d9e-a898-6e5d388244d6 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation The performance gap between experts and non-experts confirms the questions assess specialized knowledge
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9828cb73-539e-4347-9dc8-fd9687cc8dd5 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Criterion validity could be stronger with comparisons to other specialized science Q/A benchmarks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a5138293-b7fe-4318-ad4d-b26ddd1efe8e · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation However, models have quickly improved in this benchmark 5
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3f3a6e7-cce7-43f0-99bd-4dbf9fa4aaa1 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Non-expert performance gap supports specialization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d6562135-4dbc-4048-b51b-4b00c6736247 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation AI-expert performance gap reinforces benchmark credibility
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 929e0e84-e2a4-4e71-ad94-b2911b23f10a · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation specialized scientific knowledge,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9a63cc4a-52ac-41cb-8eb2-c3d35e56363c · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Coverage across multiple subfields increases generalization within biology, physics, and chemistry
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2ed36c2-9f0a-446a-aaee-e6c9af2bb0a5 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Risk of overgeneralization—high scores may be misinterpreted as broad scientific expertise beyond tested domains
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2836a1c7-22b2-44b4-b9a2-a99b2737b63b · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Multiple-choice format limits assessment of forms of reasoning like logical deduction, or abstract problem-solving
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 25e56cae-ba12-45a4-b0dc-884c5f15a0fc · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: GPQA tests factual and applied knowledge rather than abstract reasoning skills
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 857b39e9-3377-40e3-ac2d-333ab09c7b41 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Additionally, the dataset can distinguish between human experts and non-experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e212228e-88f6-4821-800d-e25cebbb0ab3 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Reasoning should generalize across domains, but GPQA only includes three scientific fields
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8eee1c0d-a64f-46e5-8a27-58db244e3300 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation reasonable
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8d45239-9f9f-4e5b-9b9d-416a08fafc9b · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98d92b07-d1a7-4b7b-a0eb-93d007f90e04 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation defe7196-1ec0-4fec-8898-98ea0c82be6c · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Description of dataset
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3435c0a1-4354-464b-9864-10167e30267e · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 314e82e5-7ec6-4e59-9974-609555249442 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 374a72bb-ee27-4041-824a-a01016770d2c · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Note, this is not about trained model performance (e.g., Recht et al
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 21c4556e-5932-40d8-a431-35f7e403c5b9 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38c7cba0-3a9c-4a91-b54b-17bf7a1d5bb5 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d63c8b8c-094a-4fa2-a226-cc1a65b02cd1 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: It may not comprehensively represent features present in non-natural or synthetic environments, nor fully capture abstract contextual cues
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4224d2f-48a6-4596-a71e-77320d685369 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0bbcd34a-112f-429f-90de-40cda033317f · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 36d13892-7162-4b15-a966-f65264cd1d9b · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: The degree of generalizability across the span of domains (e.g., synthetic or non-natural images) remains to be fully validated
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 664748a4-c781-4548-b7e3-ea9d71d1a539 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation aa0bbba2-7fc1-480d-9532-5e2fed1a82d2 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cb75f443-81d8-4693-bbe5-c05a2b03ea3b · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: There is limited evidence that high performance on this narrow task reliably predicts the broader and deeper aspects of overall visual understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 10dcecfa-05b2-42d2-8706-8d2e1f3ae016 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b6366ec-350a-4409-9dc1-dcb48df856a5 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: Its ability to generalize to tasks requiring integrated reasoning, spatial aware- ness, and contextual interpretation remains unconfirmed
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 87eea91f-40c2-4388-96a2-61e11d7dcc78 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation • Weakness: High classification accuracy might be erroneously interpreted as evidence of complete visual understanding, potentially misleading real-world applications
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c7ed114-8a5e-403c-aefb-8842c53b4342 · outbound
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation blocks world
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 326f22e5-6d1b-4715-85f1-3cc44efa40da · inbound
UQ: Assessing Language Models on Unsolved Questions Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daec776e-2a51-4fdc-8afc-a0a952caa826 · inbound
No-Knowledge Alarms for Misaligned LLMs-as-Judges Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ef653f-f02e-4ba4-8052-2fa179d6da71 · inbound
Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5f164392-0269-4c83-904b-7292b3b99a71 · inbound
The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 83fd9b99-ed26-4611-b85a-f62bb04639da · inbound
Making AI Evaluation Deployment Relevant Through Context Specification Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 562ca067-ebcf-4cb4-a100-7534dccc8cbb · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8de174f7-0c80-4bd6-80df-08da94e56963 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4281925b-e72c-49a9-b7f6-fbbeccab1147 · inbound
FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f5a476b8-042f-4952-a799-58c1f61035a1 · inbound
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7af2a256-0a3b-4b0e-a449-f237cd74fd43 · inbound
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 029fdfae-f60a-4870-994f-e268d60cc0be · inbound
Quality Is Not a Safety Proxy Under Quantization Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b3ddd3a3-d757-49a8-b4be-7684a6319362 · inbound
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8be2027d-c037-43ae-a307-9ef24e6b0059 · inbound
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30aa7dbf-48ba-4767-8b93-71d7c1898e5d · inbound
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e40aeb1-1589-4e3d-a311-c25994277cd6 · inbound
Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46796ac-f749-410b-a7b8-98a6f8f55403 · inbound
On the Convergent Validity of Offline Evaluation Designs for Recommender Systems Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387b1c4c-1805-4ae3-83f0-3715b414ce0d · inbound
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.