Pith. sign in

Paper Citation Record · LEDGER

A Conceptual Framework for AI Capability Evaluations

As of 7 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 1 inbound Pith citation observation for arXiv:2506.18213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18213 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:25:52.911449Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T04:54:26.888562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T04:55:03.356720Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact8
  • verified fuzzy37
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56a46f12-71ee-4976-a6fb-c9092c2692c2 · outbound

This paper cites Early insights from developing question-answer evaluations for frontier AI , 2024.

A Conceptual Framework for AI Capability Evaluations Early insights from developing question-answer evaluations for frontier AI , 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.551846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.551846Z digest=sha256:9b6f1ac613085fc93f5185ca4ecc6d0c5dd069fdfce5795b5e0e1cdcd3cd0380

Observation 0ac43f7a-de55-441e-8b8a-f3db4df5f146 · outbound

This paper cites Benchmarking foundation models with language-model-as-an-examiner.

A Conceptual Framework for AI Capability Evaluations Benchmarking foundation models with language-model-as-an-examiner

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.636681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.636681Z digest=sha256:458cb6e6e23fd7318a7061fe4a1a9fc9a6be5dcdf0decf816156ae2958635062

Observation 2aec37df-5ee1-4c78-bc3d-08e2ba2adcac · outbound

This paper cites Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation.

A Conceptual Framework for AI Capability Evaluations Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:55.126842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:45.709164Z digest=sha256:1385c76319aa2588544c81ca166a7aeeca5a52e9324c902055daa110d537b00f

Observation a571063c-b4e5-48bf-9264-cd696e6d2955 · outbound

This paper cites A quantitative study of nlp approaches to question difficulty estimation.

A Conceptual Framework for AI Capability Evaluations A quantitative study of nlp approaches to question difficulty estimation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.801503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.801503Z digest=sha256:4f6ca2ca9b45525442e72c5acba2f76cd5e40f5c922b6b056e6185722c856143

Observation b9b95a89-88bd-4c01-9327-54fbe29ccf2a · outbound

This paper cites Evaluating AI for Law: Bridging the Gap with Open-Source Solutions.

A Conceptual Framework for AI Capability Evaluations Evaluating AI for Law: Bridging the Gap with Open-Source Solutions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.871791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.871791Z digest=sha256:9a0f03341aa0485714c2ee973c84b688602db80510ace47e95527aad2dce6ceb

Observation 82e98749-5e33-43ee-bfb2-f52008d9eecf · outbound

This paper cites F., Ammanamanchi, P.

A Conceptual Framework for AI Capability Evaluations F., Ammanamanchi, P

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.907864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.907864Z digest=sha256:5c3c00fb1fcec3d1a01e60b3941e34dcac92b215b87edfd51af1d904080946cb

Observation 5ac09d81-d2ad-44df-b2e6-9c2291cff01f · outbound

This paper cites R., Steunebrink, B.

A Conceptual Framework for AI Capability Evaluations R., Steunebrink, B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:45.928285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:45.928285Z digest=sha256:f88fa3ac14661d798400cf24c076856bb166b3add4b3fa2d4a6d06fe7fae4524

Observation 5a6b3a82-9480-45a9-a5f2-8775b2a04c42 · outbound

This paper cites T., Li, Y., Lundberg, S., et al.

A Conceptual Framework for AI Capability Evaluations T., Li, Y., Lundberg, S., et al

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.009313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.009313Z digest=sha256:5898be64d588b61ab8ad3bc943413700e2dea703d7af8c03ec7e82a2a8661033

Observation 33e05663-89b0-4d63-afd2-e5be8e5a1bde · outbound

This paper cites Evaluating AI Evaluation: Perils and Prospects.

A Conceptual Framework for AI Capability Evaluations Evaluating AI Evaluation: Perils and Prospects

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:54.924747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:46.120217Z digest=sha256:3027c851584e05232ac96879ca80613a5f9202f98ff8f9b5d7583853b7f358c6

Observation e55a7c09-46e6-4749-9b66-984a37501275 · outbound

This paper cites Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture.

A Conceptual Framework for AI Capability Evaluations Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.193957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.193957Z digest=sha256:9740648c8c0992c4de13f620b1b24d9f5f5512044eff40e598499cb1ac430617

Observation 3973a4fa-89d9-446d-a73e-f298d673980b · outbound

This paper cites Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility.

A Conceptual Framework for AI Capability Evaluations Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.325101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.325101Z digest=sha256:5adc940408b566ba8020ab0149c790eeeee727ab7ae015227adbddb0681371fd

Observation 59c604f2-32fe-4fd8-8786-6a2409fa4af1 · outbound

This paper cites L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al.

A Conceptual Framework for AI Capability Evaluations L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.408648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.408648Z digest=sha256:596a668272a000564f1be17191b5b91067d7d00ef5f88c610e38a55c17b2c6b3

Observation 072e7f91-91c7-413b-861d-bd6feb78d302 · outbound

This paper cites Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks.

A Conceptual Framework for AI Capability Evaluations Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.504076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.504076Z digest=sha256:f3967abfc8e6d44a3083d3481f125cf7cd7c65d68b4b02b8bc5eeae2c08781d0

Observation b18bfa59-ee0d-4948-a88e-14d35f42f66e · outbound

This paper cites N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J.

A Conceptual Framework for AI Capability Evaluations N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.548943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.548943Z digest=sha256:02dfe4792d580c718c75ca63f4818667b139d642c873cdd6b2f85d21ef8715fb

Observation 3807f2cd-3c9a-450e-8c1c-5b7852b3bea5 · outbound

This paper cites On the limitations of reference-free evaluations of generated text.

A Conceptual Framework for AI Capability Evaluations On the limitations of reference-free evaluations of generated text

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.659204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.659204Z digest=sha256:405f758604af2884782916d725be76e11bed2756f4a9775e3eec77e778f1fec8

Observation adc3836f-5c87-4296-a62b-0f2741beb6ce · outbound

This paper cites R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M.

A Conceptual Framework for AI Capability Evaluations R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.709874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.709874Z digest=sha256:1cdbec51360d68b5e5ff817497cc446bed6917d1c9f3acd76866712bd1adb843

Observation 75646943-5fa9-44f0-8e46-1ff23591c1a8 · outbound

This paper cites Generalization or memorization: Data contamination and trustworthy evaluation for large language models.

A Conceptual Framework for AI Capability Evaluations Generalization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.762683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.762683Z digest=sha256:ca6f6ccec2b9b100716eaec02bd43f336891da8bfb3c2fafb83c666844432b7b

Observation eadbae44-69dc-4b5f-9fd4-33b86c0d57c3 · outbound

This paper cites W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H.

A Conceptual Framework for AI Capability Evaluations W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.873694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.873694Z digest=sha256:5587f01d648004a3b7e53087702b8d90a0e433da652dafd1ff58adaf2292ce5d

Observation bd68057e-87f1-4259-aafc-0a5e2ca81e00 · outbound

This paper cites Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation.

A Conceptual Framework for AI Capability Evaluations Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.923418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.923418Z digest=sha256:303c396595b223198a5b54c9bc7eb190e7c4e769ba3bbf8c4699a5011e6dfe32

Observation 34127b5f-062a-4d44-87b3-254d0096ec13 · outbound

This paper cites Second draft of the general purpose AI code of practice, April 2024.

A Conceptual Framework for AI Capability Evaluations Second draft of the general purpose AI code of practice, April 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.040113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.040113Z digest=sha256:18ef2cceef0e9e8040bc27e716eefb02639c8eed6c67148fa10d176c0889632d

Observation b1fe028a-aa7d-46f4-9a9f-035d54b4ab7a · outbound

This paper cites Issue brief: Early best practices for frontier AI safety evaluations, 2024.

A Conceptual Framework for AI Capability Evaluations Issue brief: Early best practices for frontier AI safety evaluations, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.131398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.131398Z digest=sha256:405a6eb0e9a80d2b23c794bd3a7ca7352a0ed0dc8b7836da85de33c07dafb5a1

Observation 3b78b36f-b168-48c2-8ac5-c73310969e8d · outbound

This paper cites Llm-based nlg evaluation: Current status and challenges.

A Conceptual Framework for AI Capability Evaluations Llm-based nlg evaluation: Current status and challenges

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.228222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.228222Z digest=sha256:4bca061c435f6431cc73e20a44fad76ca6dace2a6a6343a8b6cbdadc68b78f43

Observation 16e1fce5-805f-480c-b180-0b148e7989ba · outbound

This paper cites A case for better evaluation standards in nlg.

A Conceptual Framework for AI Capability Evaluations A case for better evaluation standards in nlg

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.276468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.276468Z digest=sha256:d04ded4fe9cbad0cb780ce4ee8eebbfe948fe344268b589cd8f2440703a94259

Observation c13ba138-d762-47b2-a3a6-d8d958f6c77b · outbound

This paper cites Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text.

A Conceptual Framework for AI Capability Evaluations Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.313523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.313523Z digest=sha256:a3474b488bc52cb7c8ec97ace54fc255e7a4b7f6ee235966a1b3b34da634d956

Observation c51e64f6-6ba2-4066-945f-a0fbdd1330c7 · outbound

This paper cites Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models.

A Conceptual Framework for AI Capability Evaluations Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.379142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.379142Z digest=sha256:c8fd1871c4734955036f5dac1e637ac520cfab0d16ea0544ae05c41a94b567bd

Observation fe5c778a-c3f4-4eef-b481-74e220bd9707 · outbound

This paper cites R., Hullman, J., and Subramonyam, H.

A Conceptual Framework for AI Capability Evaluations R., Hullman, J., and Subramonyam, H

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.435980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.435980Z digest=sha256:e45be5a521506084247a3398c27cf055a8d8ba2864fad1bde663b79a811015ad

Observation 9d66b178-9203-416a-b59f-cba00e3474cb · outbound

This paper cites Deception abilities emerged in large language models.

A Conceptual Framework for AI Capability Evaluations Deception abilities emerged in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.478675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.478675Z digest=sha256:a9b47a8b15b5ec9dd8464d41ea54bc176713b5cc91d82129682ccb0b392dfa3f

Observation 4ab00757-dd10-479e-bb7c-d091c9ea6e6e · outbound

This paper cites Machine Psychology.

A Conceptual Framework for AI Capability Evaluations Machine Psychology

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.543373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.543373Z digest=sha256:f9f53108e14bba3237b89428a17111d9ba2ba0882b5e9557a3795475310385e3

Observation a97ec3b4-39a2-481f-b43a-458a2123d40c · outbound

This paper cites a m \"a l \.

A Conceptual Framework for AI Capability Evaluations a m \"a l \

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.585201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.585201Z digest=sha256:8f7ddf5b362427659c93a56f08b72663f413d70a1672fa5cd1d894c3d257fe23

Observation 9e54ed9c-31cb-4272-984b-aa9ada6b3ea8 · outbound

This paper cites and Sharadin, N.

A Conceptual Framework for AI Capability Evaluations and Sharadin, N

Reference 30

Resolution
verified exact
doi, observed 2026-08-06T23:25:53.659452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:47.621675Z digest=sha256:cdcb0e9b8226f79d5bc228a505c402b5d086e99aff0e2e4c3f9255c5eefc2a5e

Observation ee802a70-7bb5-4a23-ba11-2a5ebf8ea74f · outbound

This paper cites an unresolved cited work.

A Conceptual Framework for AI Capability Evaluations Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.677631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.677631Z digest=sha256:497a79007bfbf46bc1908807912c72a2ee49a3059027332dd7dfaa7a993a3fca

Observation 710eef11-3f7c-4be0-82b0-5082899058cc · outbound

This paper cites R., Srivastava, A., and Agrawal, P.

A Conceptual Framework for AI Capability Evaluations R., Srivastava, A., and Agrawal, P

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.733085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.733085Z digest=sha256:7e35fd49f3bc48ca217e91f943cf6b36d77dcf327f6cebbea110a87c26c771cf

Observation 6dba6565-0142-4763-9a95-b2de90f8aebd · outbound

This paper cites Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions.

A Conceptual Framework for AI Capability Evaluations Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.781422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.781422Z digest=sha256:7ba9421cbeb73fb8b6f646248107887c0a89c75a095e85df1f67abbd9008b220

Observation 35245c9b-6ac1-4d25-abec-e0eb5ed53118 · outbound

This paper cites An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4.

A Conceptual Framework for AI Capability Evaluations An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.829032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.829032Z digest=sha256:3400ebd4343d75529d7abb93ac8b1dce333ea3bed6df0def992790f0eb02737a

Observation aa3b399a-3ee0-4ed2-a902-c0981b5311d0 · outbound

This paper cites M ath P rompter: Mathematical reasoning using large language models.

A Conceptual Framework for AI Capability Evaluations M ath P rompter: Mathematical reasoning using large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:47.914166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:47.914166Z digest=sha256:1f2ac9a1e60421f1454cdeeb62ffe7c1040b30f1fc610080c829fbf49866586e

Observation c1da4d38-4607-4083-9419-d342f740bf94 · outbound

This paper cites Reference-free Evaluation Metrics for Text Generation: A Survey.

A Conceptual Framework for AI Capability Evaluations Reference-free Evaluation Metrics for Text Generation: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:54.728552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:47.974612Z digest=sha256:8de950e7663363a9e63395716ef0c4d8941c64d7d566e763ae6b4e534fa3ebf5

Observation 54925f13-7666-4392-99fa-5193e14fba91 · outbound

This paper cites Toward best research practices in AI Psychology.

A Conceptual Framework for AI Capability Evaluations Toward best research practices in AI Psychology

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:54.582700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.034859Z digest=sha256:f86c0ef1a0656ff2bb28f080c4997d7895ebfd64e19b1dab211a58f81ea7aba4

Observation 55acc792-899f-4c18-8a8f-b133f93409ff · outbound

This paper cites Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks.

A Conceptual Framework for AI Capability Evaluations Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:48.087777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:48.087777Z digest=sha256:953b997b61a8c63502330da84e374058363e59e551194edecfef58bf2a320d71

Observation d45f20b3-e3e9-4e0c-ab3f-ce0a28917286 · outbound

This paper cites Cladder: assessing causal reasoning in language models.

A Conceptual Framework for AI Capability Evaluations Cladder: assessing causal reasoning in language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:48.125765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:48.125765Z digest=sha256:7db295e4fd49c1f04e649d43ac6ee36d7fa901dda2608cc7cff02835e2c5d669

Observation 00a9df19-b712-4d06-904b-49ee352b09af · outbound

This paper cites T., and Sch \"o lkopf, B.

A Conceptual Framework for AI Capability Evaluations T., and Sch \"o lkopf, B

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:01.298570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.218739Z digest=sha256:99c53a953353dcacb2680878836f0fd863f8b566fc0d0894f354ec6401de1097

Observation d01022a7-f799-4646-a912-8e1f3a2d13f0 · outbound

This paper cites R., Rockt \"a schel, T., and Perez, E.

A Conceptual Framework for AI Capability Evaluations R., Rockt \"a schel, T., and Perez, E

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:01.127649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.251258Z digest=sha256:0093272f7e0108eb030b401c38b7813be244f86367a03bc4a68e65ec2d3a85e8

Observation a592d7bd-4f17-45c0-9458-f5c56aa4505e · outbound

This paper cites Causal reasoning and large language models: Opening a new frontier for causality.

A Conceptual Framework for AI Capability Evaluations Causal reasoning and large language models: Opening a new frontier for causality

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:00.994216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.331076Z digest=sha256:1fdf65289849f2961de06ef233a4ee936f02a9126c9580a95fefd7a4979c5847

Observation 51d4cdcb-d43d-4b85-81ac-bf4ae402e3d6 · outbound

This paper cites AI Agent Governance: A Field Guide.

A Conceptual Framework for AI Capability Evaluations AI Agent Governance: A Field Guide

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:48.401240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:48.401240Z digest=sha256:16e30e364c7f2bc109157695b5ce183fe09469ec73207ace1de8759390484089

Observation bc670b74-da28-4b39-a8df-8559da724f6f · outbound

This paper cites Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation.

A Conceptual Framework for AI Capability Evaluations Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:00.883960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.447578Z digest=sha256:2b59e21817fe357e8867cc19b4421c555bb643c9475edf7510fb84fbac183ded

Observation b0d73691-b2dd-42cd-820a-3f58668ca67f · outbound

This paper cites P., Wu, H., and Yu, H.

A Conceptual Framework for AI Capability Evaluations P., Wu, H., and Yu, H

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:00.734787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.539508Z digest=sha256:28b7d4b9e04625c1c1cd5d8476e23adce131a71f34bc134460e992736659f84b

Observation 5e4129ac-86c3-4faf-9444-68359e50e825 · outbound

This paper cites an unresolved cited work.

A Conceptual Framework for AI Capability Evaluations Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:26:00.584030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.584691Z digest=sha256:beaf1a2d854ef8d72f915f750720487b6d6cd536558d1c293de9b0853b9944ec

Observation ce71553f-637a-4f2d-ae50-a049c57634e1 · outbound

This paper cites J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M.

A Conceptual Framework for AI Capability Evaluations J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:00.401461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.664770Z digest=sha256:c89f83bd0893e602033d8b2c514d203b2923b675502d34fa7af14d1cbd23dc26

Observation 004986a1-1769-4978-b256-faebbe908244 · outbound

This paper cites Leveraging large language models for nlg evaluation: Advances and challenges.

A Conceptual Framework for AI Capability Evaluations Leveraging large language models for nlg evaluation: Advances and challenges

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:26:00.186452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.732715Z digest=sha256:bb6d99935f3b48bc100d286d6ecdfc3fe678e84b7dc54f789f7f38cee7f2d456

Observation d6baae05-f4e8-4bc5-a9e8-b8c3293035cf · outbound

This paper cites D., Re, C., Acosta-Navas, D., Hudson, D.

A Conceptual Framework for AI Capability Evaluations D., Re, C., Acosta-Navas, D., Hudson, D

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.995800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.811403Z digest=sha256:1d923ec1761136b7167c205ca25b5974d6f997ef06c51bcf7cd5c1b5a2bc09c2

Observation 8cb3e554-1745-4e84-98c6-229d9424ab14 · outbound

This paper cites Rethinking Model Evaluation as Narrowing the Socio-Technical Gap.

A Conceptual Framework for AI Capability Evaluations Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:48.867513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:48.867513Z digest=sha256:42f8e477eb76a607185e0f03042573e7ccbecf715ec13a2d49ef35b41c565c28

Observation b0b0b00a-9c26-45e8-8fab-8f335df450b4 · outbound

This paper cites D., and Schmidt, L.

A Conceptual Framework for AI Capability Evaluations D., and Schmidt, L

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.828688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.914946Z digest=sha256:dd30498faadc111a3b58876288cf864a4165ed32f6e5ced7459af49adfd74a32

Observation 4a9429d4-ee19-4328-b541-4fb2a90973ca · outbound

This paper cites Against the achilles' heel: A survey on red teaming for generative models.

A Conceptual Framework for AI Capability Evaluations Against the achilles' heel: A survey on red teaming for generative models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.680442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:48.977933Z digest=sha256:bff5d3a872ab2b819a1826a894ef81ba9cb2e38d562af1bfed242bc57f7b59c9

Observation 77577b03-da1b-45f3-806a-ebeccf02e3a6 · outbound

This paper cites Datasets for Large Language Models: A Comprehensive Survey.

A Conceptual Framework for AI Capability Evaluations Datasets for Large Language Models: A Comprehensive Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.043762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.043762Z digest=sha256:f3293a3d51d10056302fac394d3a9f0ee3d62c327978b70e68bd2515ac175025

Observation ac7fea2d-36bc-4c4a-96dd-4c3d6f2e3949 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

A Conceptual Framework for AI Capability Evaluations Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.091271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.091271Z digest=sha256:f88c76a04a6a83db43e8af988c70e1fe3259b849151d308cc85b25cbd09a5437

Observation 963f8f01-a4d8-433f-b4dc-94a9d4fe8b73 · outbound

This paper cites Ablation Studies in Artificial Neural Networks.

A Conceptual Framework for AI Capability Evaluations Ablation Studies in Artificial Neural Networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.152147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.152147Z digest=sha256:ba962c9d300e4715bb1c90ffb0f3fb89cb25f8f62d434a4c9fc882a163fb287a

Observation 6998e5e9-1914-483d-ae1e-92ed1e729faf · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

A Conceptual Framework for AI Capability Evaluations Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.236771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.236771Z digest=sha256:dabd238928e2ee9e2de8c5150ca496f936359b63d982d58ffc0cae6015c96efd

Observation 4a602db4-2242-48e2-95bc-a683cd0dcb7b · outbound

This paper cites Auditing large language models: a three-layered approach.

A Conceptual Framework for AI Capability Evaluations Auditing large language models: a three-layered approach

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.470894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.298801Z digest=sha256:c6ba70ee039c60aed46af7595eb834fba0108426f9eaf88d2f59efb412dfcc6c

Observation a4c72738-21da-4f60-9f61-36aa311caba7 · outbound

This paper cites Evaluating the performance of large language models via debates.

A Conceptual Framework for AI Capability Evaluations Evaluating the performance of large language models via debates

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.307346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.379623Z digest=sha256:6fa361a6217efbed2e21bb56beed55987d17d0c92f3f1998c441150bcb907a2d

Observation 4fd25a23-0496-49c1-83f2-fb582686609a · outbound

This paper cites and Kapoor, S.

A Conceptual Framework for AI Capability Evaluations and Kapoor, S

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.162486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.434665Z digest=sha256:0cc3efe368bddcbc650fca97f41207d2f5025ec86fe9435edb7245a913278e69

Observation e9c72da9-038e-43af-a136-36975fa07e9d · outbound

This paper cites Oecd framework for the classification of ai systems.

A Conceptual Framework for AI Capability Evaluations Oecd framework for the classification of ai systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:59.068464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.515858Z digest=sha256:5ec9462d563b1811518000b969acb61d20ca30bd76bb2bf6f210eeacfa210954

Observation c5ec3c96-8051-49f4-8dac-5ee39d976327 · outbound

This paper cites and Kang, E.

A Conceptual Framework for AI Capability Evaluations and Kang, E

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.945843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.575099Z digest=sha256:80f88fa4718d5c2184980dd366693e9ec64e269be656d2f4af0d018128e5613e

Observation adcbe876-acc1-4905-90ce-ce33c6e81ce5 · outbound

This paper cites Llm evaluators recognize and favor their own generations.

A Conceptual Framework for AI Capability Evaluations Llm evaluators recognize and favor their own generations

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.679281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.679281Z digest=sha256:9eb159b6b2e9fa574222588481787ec2e9b04a6584cd771bbbc764c1f915181c

Observation fd39a5de-c451-41cb-81e9-2af393f33d57 · outbound

This paper cites T., and Soder, L.

A Conceptual Framework for AI Capability Evaluations T., and Soder, L

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.796939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.735360Z digest=sha256:9bd490f73dcf862c3ca0f280abcdf0627f8d7e6d34de38692e1dd8352344bc0a

Observation 506118b0-2baf-4524-9b19-e0dbac91c2dd · outbound

This paper cites Preliminary suggestions for rigorous gpai model evaluations.

A Conceptual Framework for AI Capability Evaluations Preliminary suggestions for rigorous gpai model evaluations

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.632455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.802752Z digest=sha256:c479f0f42331877ea336e996d5facf34b1e5884e6225631536c5c83a653d38f2

Observation c4123736-a938-4812-b18a-d77283e1e050 · outbound

This paper cites Discovering language model behaviors with model-written evaluations.

A Conceptual Framework for AI Capability Evaluations Discovering language model behaviors with model-written evaluations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.881801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.881801Z digest=sha256:5eeca86a6684f97b2a792cb1fd18881790c22cf37a9b46ffb5b8d3b79e0376af

Observation e4697ef7-3a5d-4011-8d66-007b206b5763 · outbound

This paper cites Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI.

A Conceptual Framework for AI Capability Evaluations Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:54.385746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:49.962510Z digest=sha256:e285aa3318c57d13aa9d32a0d017d936f67ac7717ee471b9427c37014d598469

Observation 84ad029e-504d-43b8-b4b9-5782528e75fe · outbound

This paper cites The roots search tool: Data transparency for llms.

A Conceptual Framework for AI Capability Evaluations The roots search tool: Data transparency for llms

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.501486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.099913Z digest=sha256:f93e17f85a760b3daf280289d6230282b2c33068a30050359057e72546f443a3

Observation bc042dc2-0ba5-4fe6-ae32-322039ba04fb · outbound

This paper cites D., Denton, E., Bender, E.

A Conceptual Framework for AI Capability Evaluations D., Denton, E., Bender, E

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.364595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.202095Z digest=sha256:a4062c3ae5e2c9cbd53bfab17499cfaef1816c32ba176cb1bbffe009a3a71086

Observation 26bacb55-2e61-4dc4-80f4-3ccf8a124bcc · outbound

This paper cites Large language model evaluation via multi ai agents: Preliminary results.

A Conceptual Framework for AI Capability Evaluations Large language model evaluation via multi ai agents: Preliminary results

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.238597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.322225Z digest=sha256:fe7a97904447494d588773be0b5d33091edd7195fb907a2661afb8e0f46f6289

Observation a6a02373-3a9c-496b-941d-1f75ed53b553 · outbound

This paper cites A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al.

A Conceptual Framework for AI Capability Evaluations A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:58.105856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.433994Z digest=sha256:0a0ecb4c98db755277fa3d1d1522f74a6f9d3e3eb789f06534025cc877df309f

Observation e4c626e1-592a-4ede-a8c0-b2d74d041af9 · outbound

This paper cites Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices.

A Conceptual Framework for AI Capability Evaluations Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.941298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.530574Z digest=sha256:a1819e19973395fc013efafb090f68c6e8db7444c403be8d278e1fa64578aebd

Observation 8911e041-adc9-4f40-ae33-939659ae9531 · outbound

This paper cites an unresolved cited work.

A Conceptual Framework for AI Capability Evaluations Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:25:57.803711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.669286Z digest=sha256:d7fc6586d29f742b71d9e9048067fa874516e021be3819a2a4e6af510370bf7e

Observation be10e650-3bda-45c4-81d5-af80f392c598 · outbound

This paper cites Open Problems in Technical AI Governance.

A Conceptual Framework for AI Capability Evaluations Open Problems in Technical AI Governance

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:50.774056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:50.774056Z digest=sha256:761b4cc595c0173607854275484794fdec67cb81e3673e37cb9832bf74d0f8f7

Observation 67ffcbf6-190b-4473-9af9-39fdcc71e123 · outbound

This paper cites Better than random: reliable nlg human evaluation with constrained active sampling.

A Conceptual Framework for AI Capability Evaluations Better than random: reliable nlg human evaluation with constrained active sampling

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.636905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:50.872859Z digest=sha256:799bb7ea19a4d9c63ee0220bc46fa745146419efa4a590650deefc164a7f7548

Observation 285b97cd-d038-4788-a98a-28ad185ed40d · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

A Conceptual Framework for AI Capability Evaluations A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:50.970160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:50.970160Z digest=sha256:2f9edd026ad0dbf837b5cbe03412552f813743cc8f7c106bbb00732dfb23304e

Observation 06f09258-be5a-4152-94b6-e678c1758d10 · outbound

This paper cites L., and Agirre, E.

A Conceptual Framework for AI Capability Evaluations L., and Agirre, E

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.523765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.077904Z digest=sha256:8f202a02e0630146a2054309ed3c5e74f629ebb86417fde53bf8e225e70594cb

Observation 15611d99-e4c8-47f7-9a13-65f0a6cd6eec · outbound

This paper cites Targeting the benchmark: On methodology in current natural language processing research.

A Conceptual Framework for AI Capability Evaluations Targeting the benchmark: On methodology in current natural language processing research

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.394177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.254431Z digest=sha256:f4495f8f53edbd6d76d6db870cedb9d8a1c7fe73fe4cae1ad85c96c140bd9389

Observation ab2b176f-b0c1-4383-8f99-317bb1790866 · outbound

This paper cites The Prompt Report: A Systematic Survey of Prompt Engineering Techniques.

A Conceptual Framework for AI Capability Evaluations The Prompt Report: A Systematic Survey of Prompt Engineering Techniques

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.375160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.375160Z digest=sha256:f4aba098446d10676a313ed80b2e38f8a10490a295c1a50fad277082acddf68a

Observation 8a29cb56-42af-44b9-bbe1-b5960eeb87d5 · outbound

This paper cites Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

A Conceptual Framework for AI Capability Evaluations Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.223251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.485501Z digest=sha256:f58608f8267a8cef004e6502960287d5aa14f0bf56768fd8c3f0d9d59b7f2e2c

Observation f7fa0bd3-765b-4d2d-828e-b39219f73bbf · outbound

This paper cites Model evaluation for extreme risks.

A Conceptual Framework for AI Capability Evaluations Model evaluation for extreme risks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.618603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.618603Z digest=sha256:ed4485a220142e77430276568b88137866b78f1ed70036ef4613e5bb791f17eb

Observation 118d65c2-2e53-4ad1-b92e-aa2243f528a4 · outbound

This paper cites CHOPS : CH at with customer profile systems for customer service with LLM s.

A Conceptual Framework for AI Capability Evaluations CHOPS : CH at with customer profile systems for customer service with LLM s

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:57.126694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.763437Z digest=sha256:8da35bb9c06e70cd7c5420a5ec03356034ba4b7a3580fa32a98c1c8881cb966f

Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · outbound

This paper cites MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs.

A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.856959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.856959Z digest=sha256:45e4c4a6bc006b417f6fed83d55b50debc0710c1f42fc2f68b4ca2a6b8ad4079

Observation c1b18374-ce72-4fed-aaac-82f17e8759ca · outbound

This paper cites K., Grundy, E.

A Conceptual Framework for AI Capability Evaluations K., Grundy, E

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.971466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.907553Z digest=sha256:214ed0923f4dbdbe8b35837f589201f654236d914621c74b1159af51b7625784

Observation 185e5305-7c7c-41c5-ace3-3f8266739e57 · outbound

This paper cites A study of translation edit rate with targeted human annotation.

A Conceptual Framework for AI Capability Evaluations A study of translation edit rate with targeted human annotation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.857437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:51.938242Z digest=sha256:0f82916ab5ffcee75eda9b0968ec1ed2bcb9fc93b78aa547e1879987bba4be6a

Observation 2a49e061-705b-4e26-8d90-dee0a0afa5df · outbound

This paper cites Audit Cards: Contextualizing AI Evaluations.

A Conceptual Framework for AI Capability Evaluations Audit Cards: Contextualizing AI Evaluations

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.963205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.963205Z digest=sha256:1482a4df3de4a482c3ede7f679c8e4c3f1076f44e0c03897e41b776b19a44cbd

Observation 2ae1515b-33ea-46ea-ab73-2c5b181161eb · outbound

This paper cites Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach.

A Conceptual Framework for AI Capability Evaluations Comprehensive Reassessment of Large-Scale Evaluation Outcomes in LLMs: A Multifaceted Statistical Approach

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:25:53.477142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.051239Z digest=sha256:9736f89582745f781682c5d69c8490df7d713d429538cc32f29d343b4a47e6e7

Observation 14000850-7019-4bca-9691-1f96adcad220 · outbound

This paper cites Measuring data science automation: A survey of evaluation tools for ai assistants and agents.

A Conceptual Framework for AI Capability Evaluations Measuring data science automation: A survey of evaluation tools for ai assistants and agents

Reference 87

Resolution
verified exact
raw_fallback, observed 2026-08-06T23:25:54.212966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.117063Z digest=sha256:3dec876135ef19110a9d50abdd53ad210952b519b4bd7c063a71f9a91c95d1cb

Observation be7fc519-6150-4e21-8327-ac6ee73cc049 · outbound

This paper cites an unresolved cited work.

A Conceptual Framework for AI Capability Evaluations Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:25:56.699837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.187327Z digest=sha256:88b5ccd75b007dec7a2449b33b03469a0c7e6bea27f6913c1a7efbc08dd67725

Observation 0fd337e5-aac8-4913-a99d-2c9805687893 · outbound

This paper cites Best practices for the human evaluation of automatically generated text.

A Conceptual Framework for AI Capability Evaluations Best practices for the human evaluation of automatically generated text

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.585602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.257240Z digest=sha256:247d2ffbe791467a89ed0f0a1bc8eac31bfba5f67ab8fa0d3f595011939ef42f

Observation d86eae37-7fc0-4006-979d-4bc464c7168d · outbound

This paper cites M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S.

A Conceptual Framework for AI Capability Evaluations M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.435860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.284798Z digest=sha256:a0e26ba4d63c5b715416d0c94e611afe20bc2afaeb2e533b716a07c1b9eabdcf

Observation c09b3736-65ef-49aa-8733-d6afc68e0bef · outbound

This paper cites Mint: Evaluating llms in multi-turn interaction with tools and language feedback.

A Conceptual Framework for AI Capability Evaluations Mint: Evaluating llms in multi-turn interaction with tools and language feedback

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.273371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.336057Z digest=sha256:3e72e8b8f606509b6670fd0d82146e0897d7f84a9235c44a97700c86cefc5ab4

Observation 89be56e1-8322-4bbc-b116-dbb0f5f3bae1 · outbound

This paper cites Sociotechnical Safety Evaluation of Generative AI Systems.

A Conceptual Framework for AI Capability Evaluations Sociotechnical Safety Evaluation of Generative AI Systems

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.408995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.408995Z digest=sha256:14af71c4ca28131f995147e233fe8992618c2e7217db9f0238928714d212c235

Observation 7ea0c3d7-609e-468a-b979-7bd66ba0adf6 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

A Conceptual Framework for AI Capability Evaluations Toward an Evaluation Science for Generative AI Systems

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.454150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.454150Z digest=sha256:c05705deaff622fb74f7d6db0856f0eca9113af8df2b9980a971c9b6f6adb062

Observation 2f2f12fe-501e-44e3-9f2d-6f33a124818d · outbound

This paper cites An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping.

A Conceptual Framework for AI Capability Evaluations An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:56.134856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.511076Z digest=sha256:7b56ef8567df8030e24ea19ce3a89ebf6522c34af5dd9a55bfdc7343bf18c2c0

Observation 36c38833-b24c-4681-8198-f3d3749122ce · outbound

This paper cites A critical review of causal inference benchmarks for large language models.

A Conceptual Framework for AI Capability Evaluations A critical review of causal inference benchmarks for large language models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:55.938837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.582138Z digest=sha256:2c8011ee14626a17830a77645b5bf19a76bec03377f40b2f81bde311729b9116

Observation 1360cb66-5ad5-4bb0-8ac5-32e483f7ece0 · outbound

This paper cites Evaluatology: The science and engineering of evaluation.

A Conceptual Framework for AI Capability Evaluations Evaluatology: The science and engineering of evaluation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:52.632974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:52.632974Z digest=sha256:5bdf5daa3be8ce6ad95f67dfc2500a8c1e4ee984a1ed6c2401c4557abfca4dcc

Observation b7d1310d-4790-46ef-997a-f330a0adec6a · outbound

This paper cites Language model developers should report train-test overlap.

A Conceptual Framework for AI Capability Evaluations Language model developers should report train-test overlap

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T23:25:53.854157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.690485Z digest=sha256:26c0b10650ae272231eb1b26e3c349f0097cc71207f3b1febc14c9ee1e13038d

Observation 00149752-744b-4812-a5f1-e70f9edab429 · outbound

This paper cites Q., Shaw, R., Anthis, J.

A Conceptual Framework for AI Capability Evaluations Q., Shaw, R., Anthis, J

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:55.777337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.779117Z digest=sha256:ef675ca6a1485921d58f1560eda3040f592244419a65f77f22d5eb8c3a9ce86d

Observation ec954761-76ee-44a6-a158-ef68b503ea85 · outbound

This paper cites Pacost: Paired confidence significance testing for benchmark contamination detection in large language models.

A Conceptual Framework for AI Capability Evaluations Pacost: Paired confidence significance testing for benchmark contamination detection in large language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:55.639990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.835619Z digest=sha256:59c057d7d03cc42942b3aacc51783f7f5ae2d3c574f3212c43eff22e3f949136

Observation 739097ec-7486-4a77-88f7-9f66a3f77e30 · outbound

This paper cites and Kanayet, F.

A Conceptual Framework for AI Capability Evaluations and Kanayet, F

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:25:55.470953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T23:25:52.911449Z digest=sha256:288df478939189355c59e07bddcf79eae614fa9f90887f6f1b3ccecfb5a4947b

Pith citing papers

Observation 452b6faa-c3ac-43c7-8e0f-e9fd5cbe5e8c · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders A Conceptual Framework for AI Capability Evaluations

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.360332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:340d98d81ac3e1cf442e81964c2220984130806cb3365af8172d21af22fa508c