Pith. sign in

Paper Citation Record · LEDGER

Model evaluation for extreme risks

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 63 inbound Pith citation observations for arXiv:2305.15324.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15324 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:19:10.576400Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

58
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 604417e3-3264-4da5-b336-4c161674a669 · inbound

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction cites this paper.

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction Model evaluation for extreme risks

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:49:13.910243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-24T08:47:23.231930Z digest=sha256:f2d25b5f9f1a205eece2d9e023912060b4c0487818c278a30625714038819f6e

Observation 322984e0-ed5c-4f59-bc4a-fdd194faa95f · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Model evaluation for extreme risks

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.499819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:3f85b0e0cfa1d6b82564c108715692c7856bb7e14902e826bacd4f07c274c2db

Observation e567a0de-1275-4d07-9569-096c39ef0453 · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Model evaluation for extreme risks

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.559994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:d4d0fea4c1b54ee10ce05cc3b798c25bded514a17514d27619536eaf5f8adf14

Observation 5c16d358-cca0-46ce-8ba2-7bc8d01dac6a · inbound

Gemma 2: Improving Open Language Models at a Practical Size cites this paper.

Gemma 2: Improving Open Language Models at a Practical Size Model evaluation for extreme risks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:11:16.458916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T12:11:16.326752Z digest=sha256:4f7c9d408e83d994d92787658a8f3182fc74c5f8ad6a36fcf87d53b2e5f1581f

Observation 434cf725-7ae5-4725-aee3-fb6d3c70322a · inbound

A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems cites this paper.

A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems Model evaluation for extreme risks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:38:40.858770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:38:40.858770Z digest=sha256:659621b00e2fa06398aa0516af4df0669d4d958e924f299bed0976178eeed045

Observation 1ae69b21-1166-4fa4-9983-581f22e66eb2 · inbound

Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation cites this paper.

Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation Model evaluation for extreme risks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:57.220891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:57.220891Z digest=sha256:47b5b36d64282e45d5637011d9ebe79697c5cc2174e98c7ae43886395f1bed08

Observation a83dfb92-ed1a-4acf-aaa7-f6b5959138fa · inbound

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance cites this paper.

GPAI Evaluations Standards Taskforce: Towards Effective AI Governance Model evaluation for extreme risks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:54:33.245364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:54:33.245364Z digest=sha256:d732ada610b227b2313fb6bab748c9ce4c7c6437428296e8fd45528a9c938169

Observation a4db7fa9-1212-4148-b23a-0100283cd38f · inbound

Predicting Emergent Capabilities by Finetuning cites this paper.

Predicting Emergent Capabilities by Finetuning Model evaluation for extreme risks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:41:46.186969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:41:46.186969Z digest=sha256:e8f836779b882558672f014a0b70a29776f76357c305c34ce5147a7d840c7efb

Observation 9719a814-22c5-47cf-b1ad-ab859268a7d5 · inbound

The Superalignment of Superhuman Intelligence with Large Language Models cites this paper.

The Superalignment of Superhuman Intelligence with Large Language Models Model evaluation for extreme risks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T15:18:17.759688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:18:17.759688Z digest=sha256:a847f69396b411d2694b781393713783fa6f5a32fae67135893d9ac2ef7515ce

Observation 4c740590-8555-45ae-a944-6789247dbe5d · inbound

Towards Responsible Governing AI Proliferation cites this paper.

Towards Responsible Governing AI Proliferation Model evaluation for extreme risks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:48.248429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:48.248429Z digest=sha256:c7403c81f0adce6e7d4a84cfb9d1d281ed2376b70f2c6c2b51245aac1a32e1a4

Observation acc96f4b-8e36-4cc7-9bf2-3fe36371be0f · inbound

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations cites this paper.

Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations Model evaluation for extreme risks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:43.152288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:43.152288Z digest=sha256:efb33612dc451bf2fc943887c5591faf0b120d25558cb838bbcfddafe62e4a13

Observation 6f0d6908-41f4-485e-ad66-2dc58107be26 · inbound

Episodic memory in AI agents poses risks that should be studied and mitigated cites this paper.

Episodic memory in AI agents poses risks that should be studied and mitigated Model evaluation for extreme risks

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-10T17:58:18.139281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:58:18.139281Z digest=sha256:cf3b50520eb6c1e67463ed626c8d58869d70572d368cc4d1ee4970c9e24de95a

Observation 1ed89866-2c47-4d43-b405-9dd9e4c3180b · inbound

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development cites this paper.

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development Model evaluation for extreme risks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T05:34:03.933872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T05:34:03.933872Z digest=sha256:7bc06f80e2e9f2c7f73579acdba0de7a6d88c7d4ae827bfbd3dda1d6b8fdffad

Observation 465a581a-dcfb-4462-9f82-f3ccd21ef63e · inbound

LLM Cyber Evaluations Don't Capture Real-World Risk cites this paper.

LLM Cyber Evaluations Don't Capture Real-World Risk Model evaluation for extreme risks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T22:04:33.559543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:04:33.559543Z digest=sha256:9eb0eeb77221870316b6ab0b1c2f5858496e682317713cb86ffd2aa46de7e637

Observation e882aafa-ceab-4c62-9b51-c96884513ba0 · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Model evaluation for extreme risks

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.319250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.319250Z digest=sha256:444ad04ad48900d63c59845988412aaddf858ca923bc03b9881ad39a3bacfe2d

Observation 609483f9-1026-42f3-8c03-700914fe8f04 · inbound

Enabling External Scrutiny of AI Systems with Privacy-Enhancing Technologies cites this paper.

Enabling External Scrutiny of AI Systems with Privacy-Enhancing Technologies Model evaluation for extreme risks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T05:20:17.538978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:20:17.538978Z digest=sha256:270fd6e74b6e74de0a2dfcdbd67c0c7b79ea61f40f6e6fe58b12f8880d7d9106

Observation 203605c6-35df-4a78-a504-9babfc3fd8af · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Model evaluation for extreme risks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.555404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:1e2d0048d9d25bbc80f2eb59289aa5d86b000c68244107f9ac0d05d3aeb40858

Observation 0d03b0e5-4e34-441b-bdf5-c4ad5266bfe0 · inbound

Understanding and Mitigating Risks of Generative AI in Financial Services cites this paper.

Understanding and Mitigating Risks of Generative AI in Financial Services Model evaluation for extreme risks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T10:19:10.576400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:19:10.576400Z digest=sha256:799eb53dc2d5a8a89db22ba0cb65ccc594106243aefbe2bacd719716c8a991cd

Observation 206158ca-e75c-452e-ac5c-8386ca56610e · inbound

Assessing LLM code generation quality through path planning tasks cites this paper.

Assessing LLM code generation quality through path planning tasks Model evaluation for extreme risks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:12:22.450409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:12:22.450409Z digest=sha256:e164f28583aa9553401e9767fcf020f73d368f480fdb9670abcccba01c80f5fc

Observation 6147bb15-e38f-4062-bff6-212362f481ed · inbound

Evaluating Frontier Models for Stealth and Situational Awareness cites this paper.

Evaluating Frontier Models for Stealth and Situational Awareness Model evaluation for extreme risks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:22:42.688537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:22:42.688537Z digest=sha256:3068f5fba78e4369df53822fe57a583b690053e399ab140f1c92ccb01ddc222f

Observation c07d2801-e39d-4fce-a78b-665c901b99ae · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions Model evaluation for extreme risks

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.516020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.516020Z digest=sha256:872abe38160f8329cc18450c2937c092a5c24798d3845a422110453787962e75

Observation 068c3fb0-56a3-459c-88dc-cfb232f39dda · inbound

Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks cites this paper.

Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks Model evaluation for extreme risks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:51.655116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:51.655116Z digest=sha256:d7854e635b0ab7ee6fe1a71b2871186182f2fcbaf74e40c66b2ce110637d3250

Observation 24a86f83-83d2-4712-9042-e6e893aa06df · inbound

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components cites this paper.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Model evaluation for extreme risks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.649808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.649808Z digest=sha256:5bbb917300f49887b19ec9941af578cf18963471a9fdc09bd316a60b46d136c0

Observation a5823278-1eef-4d29-9d9f-d89cf8709360 · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Model evaluation for extreme risks

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:32:14.669979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:cfb71f3e77cc22b8533bb4d6c5df30b135a8a71ad9d6437adc123c351064926b

Observation ae0ef045-e116-48ed-84e6-14f9f36aac9a · inbound

MalGEN: A Testbed for Modeling and Evaluating Malware Behaviors cites this paper.

MalGEN: A Testbed for Modeling and Evaluating Malware Behaviors Model evaluation for extreme risks

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T11:07:15.347544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:04:23.938028Z digest=sha256:ef63ac5235b54f8738e0a8445370ee8e3f6e4c44d42c158d95794d5387f70d09

Observation 1c6d6086-2acf-4e2d-bf5f-454524113818 · inbound

UCD: Unlearning in LLMs via Contrastive Decoding cites this paper.

UCD: Unlearning in LLMs via Contrastive Decoding Model evaluation for extreme risks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:59.887715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:59.887715Z digest=sha256:02c6e2c3d93ea6f2ba62aa792a8421598639cfef510e910e86df6e114cd0b20d

Observation f7fa0bd3-765b-4d2d-828e-b39219f73bbf · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Model evaluation for extreme risks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.618603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.618603Z digest=sha256:a581bfca2cfd597bd514993946e7f3e0214833a5df1d3100eeb335a9cf0114a6

Observation 7bb39678-e27a-493a-87d7-193ec5fc5c53 · inbound

On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals" cites this paper.

On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals" Model evaluation for extreme risks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:57:43.091137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:57:43.091137Z digest=sha256:838855797b4a06e24eb142d61e87b1b185f1c1a51706039246ef1ae40e83251f

Observation cd4dda47-dcb5-4268-8aac-5b76de9ed102 · inbound

From Turing to Tomorrow: The UK's Approach to AI Regulation cites this paper.

From Turing to Tomorrow: The UK's Approach to AI Regulation Model evaluation for extreme risks

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-06T20:32:41.866495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:32:41.866495Z digest=sha256:8883f8fb60e31556a612d0af32cec807636f20483aae555bcce490c2415e1a5e

Observation 3b823d75-1aec-4bdb-afaa-6f87ed338aaa · inbound

Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI cites this paper.

Domestic frontier AI regulation, an IAEA for AI, an NPT for AI, and a US-led Allied Public-Private Partnership for AI: Four institutions for governing and developing frontier AI Model evaluation for extreme risks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:11:07.257632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:11:07.257632Z digest=sha256:9014e9549943534abe388e6b58e4fe18870828601768d459d4a99c60f048858a

Observation a2704126-6a95-405e-b8d2-e321a4781f2c · inbound

Technical Requirements for Halting Dangerous AI Activities cites this paper.

Technical Requirements for Halting Dangerous AI Activities Model evaluation for extreme risks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:50:30.606349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:50:30.606349Z digest=sha256:8950d1ef76d95094dd14ab902d95b9ff2f22fee3bb7cb2043bad8bd94875cf34

Observation 30fbed42-5157-40ac-b14a-eeaa3bcee6a0 · inbound

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors cites this paper.

ExploreGS: Explorable 3D Scene Reconstruction with Virtual Camera Samplings and Diffusion Priors Model evaluation for extreme risks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:41.829098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:41.829098Z digest=sha256:f582896dc9e4914d2ffd369a0a6ed015f8ff8471d772d50c8784482e6ba77177

Observation 3c1b739c-c238-469b-a268-36ebc6ea9d53 · inbound

Designing Incident Reporting Systems for Harms from General-Purpose AI cites this paper.

Designing Incident Reporting Systems for Harms from General-Purpose AI Model evaluation for extreme risks

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.268321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T00:16:17.173186Z digest=sha256:718b45ed31f25a9de30d81aadffc089ae0be8cf6ed3fb50252d19974a7ff36df

Observation e229ee5b-ed62-464e-b88a-83c79af78af3 · inbound

Internal Deployment in the AI Act cites this paper.

Internal Deployment in the AI Act Model evaluation for extreme risks

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T17:54:18.544842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T17:51:47.841707Z digest=sha256:0df2b7188c988551f0f0ae9ae4376e64acbf3ddce7f2c73b706bf4da1605330b

Observation ae434422-e6ff-4489-91f0-5e33c4893222 · inbound

LLM-Guided Prompt Evolution for Password Guessing cites this paper.

LLM-Guided Prompt Evolution for Password Guessing Model evaluation for extreme risks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.210679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:50:46.308625Z digest=sha256:a51e0dfde01623786978200a1fcf0a9d5e8f2f38b5d1e69cbcb224a74ae78399

Observation 1f6a3b20-f614-4c80-a058-4dff33e7f0d8 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Model evaluation for extreme risks

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.409992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:beaad45901a667257bbccef2bfee588a62a024e114df8bc19de8714ad2614863

Observation ef47b39b-b3c0-486c-99b4-91d58f4a3192 · inbound

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards cites this paper.

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Model evaluation for extreme risks

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.365784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T21:58:05.584559Z digest=sha256:056392063fe98e9ff1d0e341b6d2f2f9c658b2afe414de16668e058cb663e70a

Observation e57dc426-ca56-4c82-b65b-78808d36c75f · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Model evaluation for extreme risks

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.432362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:bc186e95e8acb587a09e16cbc3bc021cab2f2146f35f1f1a5bd8f1fdd82cb2ff

Observation fa8dc43b-6271-499b-9812-a4d0c020d8c0 · inbound

Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM cites this paper.

Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM Model evaluation for extreme risks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.196298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T16:55:19.775120Z digest=sha256:7f0215bf9d8d160e18b77e70afeafc5b9baa3915d9f14d12e5bd0af91cd46dcc

Observation 70fa7b83-c04f-41ad-b983-ff44680cfb62 · inbound

Artificial Jagged Intelligence as Uneven Optimization Energy Allocation Capability Concentration, Redistribution, and Optimization Governance cites this paper.

Artificial Jagged Intelligence as Uneven Optimization Energy Allocation Capability Concentration, Redistribution, and Optimization Governance Model evaluation for extreme risks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:09.016460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T14:13:59.908810Z digest=sha256:66efabf113da3d7a956ecb2550b8b775c0e0a314c53a9974e13f9f31fd5f8094

Observation bcc58897-32c0-4dcf-93e3-8d1564eb52cd · inbound

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts cites this paper.

A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts Model evaluation for extreme risks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:40:43.534363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:11:29.066362Z digest=sha256:e50f3da3bc80ffede109473e772388eccb53cd79868cd3f99266d96322ecf407

Observation 612496ef-2aad-468a-b2f1-e8b1fdc66928 · inbound

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels cites this paper.

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Model evaluation for extreme risks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T21:39:24.651511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T12:07:02.778631Z digest=sha256:f80a49d6b69e271bbccd1ae6f84116181ed9f4249ff3fae3fe1e5dd6dad8e56b

Observation 4abdc819-bec3-4252-9092-d4bf6a8d8648 · inbound

Overtrained, Not Misaligned cites this paper.

Overtrained, Not Misaligned Model evaluation for extreme risks

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:47:26.248329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T06:45:52.544674Z digest=sha256:738d2253261aa93926d2295552bdef8860d29d89c35e349f6b6037af141c7dff

Observation 2a4f006a-c173-44c6-8881-bf25d2053da4 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Model evaluation for extreme risks

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.277926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:c06e250803cffa75b2273bac07c41c4783f6ebf9ddc1c1415cdb55e15d734e6f

Observation 0edff6ea-befb-4647-851b-bc7182eea175 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Model evaluation for extreme risks

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.474151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:2bd9e9d32ffbbd1d9456f794b62eba549970631842664dc9a813ebb2d5a7f1c5

Observation 6aec64f0-8e08-4930-af60-720b9a55ea30 · inbound

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands cites this paper.

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands Model evaluation for extreme risks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:03.991229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T20:53:04.274840Z digest=sha256:77bd1478ca798d7d1bf825224bbbb615a48ca114f0ce1c09f8f6127ed26be922

Observation ce90f471-643a-4f7f-a4ff-9105b12525d2 · inbound

Measuring Safety Alignment Effects in Autonomous Security Agents cites this paper.

Measuring Safety Alignment Effects in Autonomous Security Agents Model evaluation for extreme risks

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:28:05.922999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T04:24:46.357505Z digest=sha256:c451e9f46a99a449d53e2181334101daaf3c19ba5c7c52d75fb43366a4aa8a13

Observation d0953d61-28a0-4d0a-bf04-78c563c4c583 · inbound

Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security cites this paper.

Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security Model evaluation for extreme risks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:03:54.371283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T02:01:42.033718Z digest=sha256:6ad0c53f792cfa1283d32d0ad50f9d277a9e40163206de9e6c8a614368accfc9

Observation 501e47f3-025b-45c7-a4e5-82ee61a3cc84 · inbound

Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security cites this paper.

Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security Model evaluation for extreme risks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:03:54.276808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T02:01:42.033718Z digest=sha256:8d21a93b37e3095807c062eda14664bc2fedf9de94652741bcd156c39ba4e968

Observation 1463a8d7-e34d-4643-adf6-ffc0360b188b · inbound

Consistency Training while Mitigating Obfuscation via Rate Matching cites this paper.

Consistency Training while Mitigating Obfuscation via Rate Matching Model evaluation for extreme risks

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:32:18.172629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T14:25:43.147442Z digest=sha256:cd6a39f723a11906e2edc79752e4e820a77f064c199bd81b2454da4b2b7b4f9d

Observation 89d9bf51-6c35-487a-ac47-5266fa246336 · inbound

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing cites this paper.

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing Model evaluation for extreme risks

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.473340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:17:01.173495Z digest=sha256:eb3908af5dd0ab09d8bbe52d26d0ccbf301c6ad1e45092350d8914dab0f30955

Observation bce19943-8a43-49cf-9f23-fb169f0bc0cb · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs Model evaluation for extreme risks

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T01:41:29.305785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:25a58719cda6b9b6e266bf33dcd0cd9486b2df735ba35d319e6726bc0c28bb51

Observation 2aad5fe4-fc00-42c9-9231-205c1d2855b6 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Model evaluation for extreme risks

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.287703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:77d806b57bc35a2debf02dfb9c9e0d62077930d688c82abd60f7d85a36af3ee0

Observation fae8241c-6b94-4ea4-895c-e5540ed2c83c · inbound

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework cites this paper.

AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework Model evaluation for extreme risks

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:59.936017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T23:42:20.304205Z digest=sha256:864f6203441ebba935af8edf56a7f23fbd2ba446ac596eab2c6379762ba9137e

Observation c2a1ea71-96f9-4b62-9533-dff80c7ec0c2 · inbound

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map cites this paper.

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map Model evaluation for extreme risks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.494876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-03T10:54:37.039277Z digest=sha256:9bec92b8605f564d5dc0f4ce87aef82f232c779d48634146abf81a1d2f3a083e

Observation 8c1c4471-e678-41dd-81b0-679306396f64 · inbound

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies cites this paper.

Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies Model evaluation for extreme risks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T02:36:01.385664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:36:01.385664Z digest=sha256:42aad11670733c1f02661855e3ef0e5e2be19ec5366ba5f68dd44f6491b1cd9a

Observation bcf9b8d3-76a9-4a20-8afa-d14f60399198 · inbound

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI cites this paper.

Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI Model evaluation for extreme risks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T01:45:06.581823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:45:06.581823Z digest=sha256:a575da4541589cb1868a16adb2cc8d9b20eca8cb9dfdfa7dd106bc50f9c84cd4

Observation 215fec64-896f-44cb-9cd7-3387bec92bf5 · inbound

Open Problems in AI Incident Governance cites this paper.

Open Problems in AI Incident Governance Model evaluation for extreme risks

Reference 107

Resolution
unresolved
no resolver link, observed 2026-07-11T07:50:18.332768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:50:18.332768Z digest=sha256:dd9d37764d5dc605acc06b87f7a0c7aef630ea1d5b6d1c9e7a96ea4953064bf7

Observation 49d6e1a2-c8ae-4098-810d-ef6e0cd631dc · inbound

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations cites this paper.

NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations Model evaluation for extreme risks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T11:20:52.129640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:20:52.129640Z digest=sha256:644f91f492874e22b6c371bebf62d1a16756a563d1f0c7e3374f962cbd205921

Observation ad1855e4-32f5-4754-96fd-6e96196fc2d4 · inbound

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI cites this paper.

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI Model evaluation for extreme risks

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-02T16:36:34.814502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:36:34.814502Z digest=sha256:373d19b790b944de9a9b361aa65f32700147dc2374c046faa99ff4eddcb550c6

Observation 91461674-e372-43c8-b8fd-57db70a1d14f · inbound

ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents cites this paper.

ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents Model evaluation for extreme risks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T23:24:19.561447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:24:19.561447Z digest=sha256:d7479216aec33d209744c75dcf47143964c133601c1f917b525739af51760af2

Observation 1c06141e-dd1b-40cd-b3f9-df88df1e4cb1 · inbound

Accountability Asymmetry and Structural Trust in Autonomous AI Systems cites this paper.

Accountability Asymmetry and Structural Trust in Autonomous AI Systems Model evaluation for extreme risks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:51:31.831708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:51:31.831708Z digest=sha256:75c4867d4ae17e59dc8060a1c1b8208ef80b4817405c9bca80934550c30d0c86

Observation fea38624-2cc2-4078-bf04-a49900f7dedd · inbound

"Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents cites this paper.

"Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents Model evaluation for extreme risks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:27:27.394338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:27:27.394338Z digest=sha256:8fe0641c39243f8a31d83a2a766bcce6096f23fcd40db5c705997de09506f62d