Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:56:21.102873Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2412.08653.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T11:56:21.102873Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T09:01:07.210356Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
29 of 29 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 7468784d-5e67-48ad-aee4-9c0c36b67029 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c07cfaa-96fd-4d42-82f1-afecfaf2b215 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Safety Cases: How to Justify the Safety of Advanced AI Systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc3b5182-caf2-481a-8daa-223e31c20c3e · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Safety case template for frontier AI: A cyber inability argument
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19967bf-392a-4674-a7af-9dbf3aee01c5 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Anthropic’s Responsible Scaling Policy Version 1.0, 2023
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 817b2439-fa05-4b40-b1e5-e808212ae3f6 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Preparedness Framework (Beta), 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e81ae6be-27ad-4d7b-9c40-76a12851e49a · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Frontier Safety Framework, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7d3fc15-ed60-4b0e-984f-dca0c53661e7 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Evaluating Frontier Models for Dangerous Capabilities
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba8390aa-6224-4285-90d1-2c691ef4fdf2 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do LLM Agents can Autonomously Hack Websites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a0ca4c-7e19-4ec5-af0f-07ede74845fd · outbound
What AI evaluations for preventing catastrophic risks can and cannot do LLM Agents can Autonomously Exploit One-day Vulnerabilities
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8630ca99-e29e-44ea-960c-c5679d8e56af · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Teams of LLM Agents can Exploit Zero-Day Vulnerabilities
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3b758c-01ec-4f00-8584-2d1dbeda2f49 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ac8c7fa-0f30-479f-8f82-fc30170e4a39 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 12a40c67-8ce5-441d-8cc3-e7c4a83cb7fd · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Black-Box Access is Insufficient for Rigorous AI Audits
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 300b0777-7082-439e-ae89-b4811143885e · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Number 1
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 26f9a1da-deb7-4799-9710-10215a022e79 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Mistral CEO confirms ‘leak’ of new open source AI model nearing GPT-4 performance
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b269c1b4-c725-4567-b27c-298ddd69fc68 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Coordinated pausing: An evaluation-based coordination scheme for frontier AI developers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c574b4-ab9c-4204-806c-30cee30ae79d · outbound
What AI evaluations for preventing catastrophic risks can and cannot do CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b452253-cdb7-4371-bea3-85ac17dd387d · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Project Naptime: Evaluating Offensive Security Capabili- ties of Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2cdbc31e-6fd0-40f6-9683-586b08b5b7df · outbound
What AI evaluations for preventing catastrophic risks can and cannot do AI capabilities can be significantly improved without expensive retraining
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1e790e-94f8-468b-85ec-2729c0a77599 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do SWE-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0be6524a-f8b9-4584-9f82-8045a94b9f74 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do SWE-bench leaderboard
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6b0b18a3-90b8-43a3-bb7f-733cfd2ffecf · outbound
What AI evaluations for preventing catastrophic risks can and cannot do GAIA Leaderboard - a Hugging Face Space by gaia-benchmark
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4975690c-760e-4c98-ae24-983f7d988b7e · outbound
What AI evaluations for preventing catastrophic risks can and cannot do HumanEval Benchmark (Code Generation)
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 39d36d8a-4b55-4336-922f-1c3e625b3269 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do A survey on in-context learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7345fa4-c59e-4de9-8973-045ab5437f15 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Anthropic’s Responsible Scaling Policy, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f59e5af6-936f-44b7-9435-3081475a4b5f · outbound
What AI evaluations for preventing catastrophic risks can and cannot do We need a Science of Evals
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c19e7bd-3cd7-4c9b-bfa1-4a682fc8b45e · outbound
What AI evaluations for preventing catastrophic risks can and cannot do AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d36241c-6419-43a1-a959-e4f407471c27 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Stress-Testing Capability Elicitation With Password-Locked Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2a2164-8318-49f5-bb5a-3d9fa2172135 · outbound
What AI evaluations for preventing catastrophic risks can and cannot do Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f23a12e-520b-4ff7-b021-a0d8eb82b270 · inbound
From Disclosure to Self-Referential Opacity: Six Dimensions of Strain in Current AI Governance What AI evaluations for preventing catastrophic risks can and cannot do
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c9e9d58b-a850-459d-9472-b294307149e7 · inbound
Scaffold Effects on GAIA: A Controlled Comparison What AI evaluations for preventing catastrophic risks can and cannot do
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7629eaf4-63ee-48b1-8620-7afde37db9c6 · inbound
Verifying Restrictions on Frontier AI Research What AI evaluations for preventing catastrophic risks can and cannot do
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.