Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T01:04:40.936182Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.03070.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T01:04:40.936182Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 486f349b-b9d2-42d9-8d97-0c6bd8f05cbd · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard gpt-oss-120b & gpt-oss-20b Model Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00822f6b-9cd3-486a-a667-1ad9952ecd6b · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 921e6ce0-e3b2-4dbe-9683-f8857cdb3622 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68433ac3-357a-459b-b927-30ab800cbf9e · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against universal jailbreaks, February 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e9af83b-b20c-4971-8b95-8ac4c98bfb4e · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ebe1614-0dac-47a8-b3d5-f23ef00d43fd · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe28a52-b414-451e-999a-f397cc2723b0 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd399f56-0609-4170-915b-3919fcb5cd6c · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c42b4f-2a40-4877-8248-ec8866a67713 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Boundary point jailbreaking of black-box llms, 2026
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e302ba0-1a01-4499-8b35-d7db3113287f · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 974e515b-6c36-4a76-b531-59b2b6b748e1 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard The safety gap toolkit: Evaluating hidden dangers of open-source models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2cc9033-fcb5-404e-b5bc-6a51103e706d · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50f5ab3b-2cb7-4d9f-b6dd-cb0e3db16493 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dac27c1-ce00-407f-b11b-92a491cf6dca · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 774cb533-c83e-4232-9a8b-840c9f9bfdc6 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac7b7765-dcc5-4c5c-a68d-0f42c6539bd1 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard God Has Helped Us, and So Will AI
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c60ab4c-7a46-4812-8b03-d95718c6c93e · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard McKenzie, Oskar J
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5352fe49-0ae7-4ba0-be5b-86492d0ddb08 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Common elements of frontier AI safety policies, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1277843-1da5-4094-9db4-f6d75e1bfed1 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d6afd3-527e-4ea0-bb38-de08a31a3502 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard ChatGPT Agent System Card, July 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67ef4c4d-23e3-4bdf-9a3c-4d8d35ffc6a5 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c2008a-fb9e-4f1a-a5e5-861605ed6319 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Exposing the systematic vulnerability of open-weight models to prefill attacks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ef9201-8fba-4294-bb23-4a0ac23f70e8 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Qwen3.5: Towards native multimodal agents, 2026
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffed2393-7a4e-41ac-aa69-45d8c2381186 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8fe0585a-8534-4ddd-82f2-fbb94be7e2b4 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f94878-4a23-4e07-b4ef-da0b4d1def24 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89692e7a-d5e4-4df7-8875-d5e7e26c8edc · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Jailbroken Frontier Models Retain Their Capabilities
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 677b28d0-f0f8-4248-88f9-255d12d54f4b · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 407c913d-b43a-4987-b53d-247c77fac7eb · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Sorry,” “I can’t help with that, but…
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 273bbf32-5b57-4029-9f26-a0ffb085e131 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard not jailbroken
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01daf259-7b40-443d-a04f-4bc89046a70b · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard jailbroken (1)
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3de8b6a-0ba0-4112-a24b-acb5836644cc · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e870fbac-2b7d-433c-99c5-f8fe535c39f4 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce5d3ada-f4f9-4bef-a33a-18668093c0ff · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0)
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 051e85b6-83a9-447b-a859-cf1d9c650a1f · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4cf8811-900f-4081-9503-7d1530ef92f7 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b63ec2b0-961e-4470-96ae-ee7e7d6f1cf7 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Suppose we find ten working jailbreaks against each of two models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1511973-57fb-4595-a23c-3e1d34c46a7b · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Finding a jailbreak that works once is easy; proving it works reliably is expensive
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b7a4f3f6-6596-4d57-830c-85d2addbc8f7 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard A smart attacker does not run every candidate over the full sample
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33117418-4b40-46b2-92dc-61512cbe38de · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16677385-b223-4e30-bbce-beb5e9ad92c4 · outbound
AI Security Leaderboard: Methodology, Results and Minimal Standard Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%)
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.