Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:25:52.911449Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 1 inbound Pith citation observation for arXiv:2506.18213.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:25:52.911449Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-15T04:54:26.888562Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-15T04:55:03.356720Z
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 56a46f12-71ee-4976-a6fb-c9092c2692c2 · outbound
A Conceptual Framework for AI Capability Evaluations Early insights from developing question-answer evaluations for frontier AI , 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac43f7a-de55-441e-8b8a-f3db4df5f146 · outbound
A Conceptual Framework for AI Capability Evaluations Benchmarking foundation models with language-model-as-an-examiner
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aec37df-5ee1-4c78-bc3d-08e2ba2adcac · outbound
A Conceptual Framework for AI Capability Evaluations Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a571063c-b4e5-48bf-9264-cd696e6d2955 · outbound
A Conceptual Framework for AI Capability Evaluations A quantitative study of nlp approaches to question difficulty estimation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b95a89-88bd-4c01-9327-54fbe29ccf2a · outbound
A Conceptual Framework for AI Capability Evaluations Evaluating AI for Law: Bridging the Gap with Open-Source Solutions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e98749-5e33-43ee-bfb2-f52008d9eecf · outbound
A Conceptual Framework for AI Capability Evaluations F., Ammanamanchi, P
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ac09d81-d2ad-44df-b2e6-9c2291cff01f · outbound
A Conceptual Framework for AI Capability Evaluations R., Steunebrink, B
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6b3a82-9480-45a9-a5f2-8775b2a04c42 · outbound
A Conceptual Framework for AI Capability Evaluations T., Li, Y., Lundberg, S., et al
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e05663-89b0-4d63-afd2-e5be8e5a1bde · outbound
A Conceptual Framework for AI Capability Evaluations Evaluating AI Evaluation: Perils and Prospects
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e55a7c09-46e6-4749-9b66-984a37501275 · outbound
A Conceptual Framework for AI Capability Evaluations Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3973a4fa-89d9-446d-a73e-f298d673980b · outbound
A Conceptual Framework for AI Capability Evaluations Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c604f2-32fe-4fd8-8786-6a2409fa4af1 · outbound
A Conceptual Framework for AI Capability Evaluations L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072e7f91-91c7-413b-861d-bd6feb78d302 · outbound
A Conceptual Framework for AI Capability Evaluations Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18bfa59-ee0d-4948-a88e-14d35f42f66e · outbound
A Conceptual Framework for AI Capability Evaluations N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3807f2cd-3c9a-450e-8c1c-5b7852b3bea5 · outbound
A Conceptual Framework for AI Capability Evaluations On the limitations of reference-free evaluations of generated text
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc3836f-5c87-4296-a62b-0f2741beb6ce · outbound
A Conceptual Framework for AI Capability Evaluations R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75646943-5fa9-44f0-8e46-1ff23591c1a8 · outbound
A Conceptual Framework for AI Capability Evaluations Generalization or memorization: Data contamination and trustworthy evaluation for large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eadbae44-69dc-4b5f-9fd4-33b86c0d57c3 · outbound
A Conceptual Framework for AI Capability Evaluations W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd68057e-87f1-4259-aafc-0a5e2ca81e00 · outbound
A Conceptual Framework for AI Capability Evaluations Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34127b5f-062a-4d44-87b3-254d0096ec13 · outbound
A Conceptual Framework for AI Capability Evaluations Second draft of the general purpose AI code of practice, April 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fe028a-aa7d-46f4-9a9f-035d54b4ab7a · outbound
A Conceptual Framework for AI Capability Evaluations Issue brief: Early best practices for frontier AI safety evaluations, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b78b36f-b168-48c2-8ac5-c73310969e8d · outbound
A Conceptual Framework for AI Capability Evaluations Llm-based nlg evaluation: Current status and challenges
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e1fce5-805f-480c-b180-0b148e7989ba · outbound
A Conceptual Framework for AI Capability Evaluations A case for better evaluation standards in nlg
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13ba138-d762-47b2-a3a6-d8d958f6c77b · outbound
A Conceptual Framework for AI Capability Evaluations Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c51e64f6-6ba2-4066-945f-a0fbdd1330c7 · outbound
A Conceptual Framework for AI Capability Evaluations Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe5c778a-c3f4-4eef-b481-74e220bd9707 · outbound
A Conceptual Framework for AI Capability Evaluations R., Hullman, J., and Subramonyam, H
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d66b178-9203-416a-b59f-cba00e3474cb · outbound
A Conceptual Framework for AI Capability Evaluations Deception abilities emerged in large language models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab00757-dd10-479e-bb7c-d091c9ea6e6e · outbound
A Conceptual Framework for AI Capability Evaluations Machine Psychology
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97ec3b4-39a2-481f-b43a-458a2123d40c · outbound
A Conceptual Framework for AI Capability Evaluations a m \"a l \
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e54ed9c-31cb-4272-984b-aa9ada6b3ea8 · outbound
A Conceptual Framework for AI Capability Evaluations and Sharadin, N
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee802a70-7bb5-4a23-ba11-2a5ebf8ea74f · outbound
A Conceptual Framework for AI Capability Evaluations Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 710eef11-3f7c-4be0-82b0-5082899058cc · outbound
A Conceptual Framework for AI Capability Evaluations R., Srivastava, A., and Agrawal, P
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dba6565-0142-4763-9a95-b2de90f8aebd · outbound
A Conceptual Framework for AI Capability Evaluations Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · outbound
A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3b399a-3ee0-4ed2-a902-c0981b5311d0 · outbound
A Conceptual Framework for AI Capability Evaluations M ath P rompter: Mathematical reasoning using large language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1da4d38-4607-4083-9419-d342f740bf94 · outbound
A Conceptual Framework for AI Capability Evaluations Reference-free Evaluation Metrics for Text Generation: A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54925f13-7666-4392-99fa-5193e14fba91 · outbound
A Conceptual Framework for AI Capability Evaluations Toward best research practices in AI Psychology
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55acc792-899f-4c18-8a8f-b133f93409ff · outbound
A Conceptual Framework for AI Capability Evaluations Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45f20b3-e3e9-4e0c-ab3f-ce0a28917286 · outbound
A Conceptual Framework for AI Capability Evaluations Cladder: assessing causal reasoning in language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a9df19-b712-4d06-904b-49ee352b09af · outbound
A Conceptual Framework for AI Capability Evaluations T., and Sch \"o lkopf, B
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d01022a7-f799-4646-a912-8e1f3a2d13f0 · outbound
A Conceptual Framework for AI Capability Evaluations R., Rockt \"a schel, T., and Perez, E
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a592d7bd-4f17-45c0-9458-f5c56aa4505e · outbound
A Conceptual Framework for AI Capability Evaluations Causal reasoning and large language models: Opening a new frontier for causality
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51d4cdcb-d43d-4b85-81ac-bf4ae402e3d6 · outbound
A Conceptual Framework for AI Capability Evaluations AI Agent Governance: A Field Guide
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc670b74-da28-4b39-a8df-8559da724f6f · outbound
A Conceptual Framework for AI Capability Evaluations Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0d73691-b2dd-42cd-820a-3f58668ca67f · outbound
A Conceptual Framework for AI Capability Evaluations P., Wu, H., and Yu, H
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e4129ac-86c3-4faf-9444-68359e50e825 · outbound
A Conceptual Framework for AI Capability Evaluations Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce71553f-637a-4f2d-ae50-a049c57634e1 · outbound
A Conceptual Framework for AI Capability Evaluations J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 004986a1-1769-4978-b256-faebbe908244 · outbound
A Conceptual Framework for AI Capability Evaluations Leveraging large language models for nlg evaluation: Advances and challenges
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6baae05-f4e8-4bc5-a9e8-b8c3293035cf · outbound
A Conceptual Framework for AI Capability Evaluations D., Re, C., Acosta-Navas, D., Hudson, D
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cb3e554-1745-4e84-98c6-229d9424ab14 · outbound
A Conceptual Framework for AI Capability Evaluations Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0b0b00a-9c26-45e8-8fab-8f335df450b4 · outbound
A Conceptual Framework for AI Capability Evaluations D., and Schmidt, L
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a9429d4-ee19-4328-b541-4fb2a90973ca · outbound
A Conceptual Framework for AI Capability Evaluations Against the achilles' heel: A survey on red teaming for generative models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 77577b03-da1b-45f3-806a-ebeccf02e3a6 · outbound
A Conceptual Framework for AI Capability Evaluations Datasets for Large Language Models: A Comprehensive Survey
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac7fea2d-36bc-4c4a-96dd-4c3d6f2e3949 · outbound
A Conceptual Framework for AI Capability Evaluations Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963f8f01-a4d8-433f-b4dc-94a9d4fe8b73 · outbound
A Conceptual Framework for AI Capability Evaluations Ablation Studies in Artificial Neural Networks
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6998e5e9-1914-483d-ae1e-92ed1e729faf · outbound
A Conceptual Framework for AI Capability Evaluations Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a602db4-2242-48e2-95bc-a683cd0dcb7b · outbound
A Conceptual Framework for AI Capability Evaluations Auditing large language models: a three-layered approach
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4c72738-21da-4f60-9f61-36aa311caba7 · outbound
A Conceptual Framework for AI Capability Evaluations Evaluating the performance of large language models via debates
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fd25a23-0496-49c1-83f2-fb582686609a · outbound
A Conceptual Framework for AI Capability Evaluations and Kapoor, S
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9c72da9-038e-43af-a136-36975fa07e9d · outbound
A Conceptual Framework for AI Capability Evaluations Oecd framework for the classification of ai systems
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5ec3c96-8051-49f4-8dac-5ee39d976327 · outbound
A Conceptual Framework for AI Capability Evaluations and Kang, E
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adcbe876-acc1-4905-90ce-ce33c6e81ce5 · outbound
A Conceptual Framework for AI Capability Evaluations Llm evaluators recognize and favor their own generations
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd39a5de-c451-41cb-81e9-2af393f33d57 · outbound
A Conceptual Framework for AI Capability Evaluations T., and Soder, L
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 506118b0-2baf-4524-9b19-e0dbac91c2dd · outbound
A Conceptual Framework for AI Capability Evaluations Preliminary suggestions for rigorous gpai model evaluations
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4123736-a938-4812-b18a-d77283e1e050 · outbound
A Conceptual Framework for AI Capability Evaluations Discovering language model behaviors with model-written evaluations
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4697ef7-3a5d-4011-8d66-007b206b5763 · outbound
A Conceptual Framework for AI Capability Evaluations Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84ad029e-504d-43b8-b4b9-5782528e75fe · outbound
A Conceptual Framework for AI Capability Evaluations The roots search tool: Data transparency for llms
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc042dc2-0ba5-4fe6-ae32-322039ba04fb · outbound
A Conceptual Framework for AI Capability Evaluations D., Denton, E., Bender, E
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26bacb55-2e61-4dc4-80f4-3ccf8a124bcc · outbound
A Conceptual Framework for AI Capability Evaluations Large language model evaluation via multi ai agents: Preliminary results
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a6a02373-3a9c-496b-941d-1f75ed53b553 · outbound
A Conceptual Framework for AI Capability Evaluations A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e4c626e1-592a-4ede-a8c0-b2d74d041af9 · outbound
A Conceptual Framework for AI Capability Evaluations Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8911e041-adc9-4f40-ae33-939659ae9531 · outbound
A Conceptual Framework for AI Capability Evaluations Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be10e650-3bda-45c4-81d5-af80f392c598 · outbound
A Conceptual Framework for AI Capability Evaluations Open Problems in Technical AI Governance
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ffcbf6-190b-4473-9af9-39fdcc71e123 · outbound
A Conceptual Framework for AI Capability Evaluations Better than random: reliable nlg human evaluation with constrained active sampling
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 285b97cd-d038-4788-a98a-28ad185ed40d · outbound
A Conceptual Framework for AI Capability Evaluations A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f09258-be5a-4152-94b6-e678c1758d10 · outbound
A Conceptual Framework for AI Capability Evaluations L., and Agirre, E
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15611d99-e4c8-47f7-9a13-65f0a6cd6eec · outbound
A Conceptual Framework for AI Capability Evaluations Targeting the benchmark: On methodology in current natural language processing research
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ab2b176f-b0c1-4383-8f99-317bb1790866 · outbound
A Conceptual Framework for AI Capability Evaluations The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a29cb56-42af-44b9-bbe1-b5960eeb87d5 · outbound
A Conceptual Framework for AI Capability Evaluations Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7fa0bd3-765b-4d2d-828e-b39219f73bbf · outbound
A Conceptual Framework for AI Capability Evaluations Model evaluation for extreme risks
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118d65c2-2e53-4ad1-b92e-aa2243f528a4 · outbound
A Conceptual Framework for AI Capability Evaluations CHOPS : CH at with customer profile systems for customer service with LLM s
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · outbound
A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b18374-ce72-4fed-aaac-82f17e8759ca · outbound
A Conceptual Framework for AI Capability Evaluations K., Grundy, E
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 185e5305-7c7c-41c5-ace3-3f8266739e57 · outbound
A Conceptual Framework for AI Capability Evaluations A study of translation edit rate with targeted human annotation
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a49e061-705b-4e26-8d90-dee0a0afa5df · outbound
A Conceptual Framework for AI Capability Evaluations Audit Cards: Contextualizing AI Evaluations
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae1515b-33ea-46ea-ab73-2c5b181161eb · outbound
A Conceptual Framework for AI Capability Evaluations Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14000850-7019-4bca-9691-1f96adcad220 · outbound
A Conceptual Framework for AI Capability Evaluations Measuring data science automation: A survey of evaluation tools for ai assistants and agents
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be7fc519-6150-4e21-8327-ac6ee73cc049 · outbound
A Conceptual Framework for AI Capability Evaluations Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0fd337e5-aac8-4913-a99d-2c9805687893 · outbound
A Conceptual Framework for AI Capability Evaluations Best practices for the human evaluation of automatically generated text
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d86eae37-7fc0-4006-979d-4bc464c7168d · outbound
A Conceptual Framework for AI Capability Evaluations M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c09b3736-65ef-49aa-8733-d6afc68e0bef · outbound
A Conceptual Framework for AI Capability Evaluations Mint: Evaluating llms in multi-turn interaction with tools and language feedback
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89be56e1-8322-4bbc-b116-dbb0f5f3bae1 · outbound
A Conceptual Framework for AI Capability Evaluations Sociotechnical Safety Evaluation of Generative AI Systems
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · outbound
A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2f12fe-501e-44e3-9f2d-6f33a124818d · outbound
A Conceptual Framework for AI Capability Evaluations An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36c38833-b24c-4681-8198-f3d3749122ce · outbound
A Conceptual Framework for AI Capability Evaluations A critical review of causal inference benchmarks for large language models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1360cb66-5ad5-4bb0-8ac5-32e483f7ece0 · outbound
A Conceptual Framework for AI Capability Evaluations Evaluatology: The science and engineering of evaluation
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d1310d-4790-46ef-997a-f330a0adec6a · outbound
A Conceptual Framework for AI Capability Evaluations Language model developers should report train-test overlap
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00149752-744b-4812-a5f1-e70f9edab429 · outbound
A Conceptual Framework for AI Capability Evaluations Q., Shaw, R., Anthis, J
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec954761-76ee-44a6-a158-ef68b503ea85 · outbound
A Conceptual Framework for AI Capability Evaluations Pacost: Paired confidence significance testing for benchmark contamination detection in large language models
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 739097ec-7486-4a77-88f7-9f66a3f77e30 · outbound
A Conceptual Framework for AI Capability Evaluations and Kanayet, F
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 452b6faa-c3ac-43c7-8e0f-e9fd5cbe5e8c · inbound
Unsteady Metrics and Benchmarking Cultures of AI Model Builders A Conceptual Framework for AI Capability Evaluations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.