Pith. sign in

Paper Citation Record · LEDGER

AI Security Leaderboard: Methodology, Results and Minimal Standard

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.03070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03070 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:04:40.936182Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 486f349b-b9d2-42d9-8d97-0c6bd8f05cbd · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

AI Security Leaderboard: Methodology, Results and Minimal Standard gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.800351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.800351Z digest=sha256:159369ee8ac2bd3cba93a568c82dc1fb9840f81857e4c1ca39564f29a2a58195

Observation 00822f6b-9cd3-486a-a667-1ad9952ecd6b · outbound

This paper cites Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.717747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.804895Z digest=sha256:e543cdce5daab72330b104be853d8a783b4ed2d44221296f6602edb1a9ae882d

Observation 921e6ce0-e3b2-4dbe-9683-f8857cdb3622 · outbound

This paper cites Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.708262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.808590Z digest=sha256:266a897ac74cbac57fa1711dbc2bf5b748d81c5f0ad6655c6f02f5966256a13f

Observation 68433ac3-357a-459b-b927-30ab800cbf9e · outbound

This paper cites Constitutional Classifiers: Defending against universal jailbreaks, February 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against universal jailbreaks, February 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.699247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.812066Z digest=sha256:5747994226827e006dc3fb7204faf68bdfbff0af3b08f0ad721b3624106d2f9a

Observation 5e9af83b-b20c-4971-8b95-8ac4c98bfb4e · outbound

This paper cites Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.689793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.815425Z digest=sha256:0e1af100e3d8e51bb92a5c2954038281134b3a11fdf841f187702680b703de1d

Observation 8ebe1614-0dac-47a8-b3d5-f23ef00d43fd · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

AI Security Leaderboard: Methodology, Results and Minimal Standard Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.818991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.818991Z digest=sha256:c8a35f29e3cbd1687ad6f7206d9afdcefb6405dbcef225d4725f24a7278860fa

Observation dbe28a52-b414-451e-999a-f397cc2723b0 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.822825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.822825Z digest=sha256:326e6500227b6fb2f790fbf126a19c89bebad38b56ee57e3e612696e355e2f2a

Observation bd399f56-0609-4170-915b-3919fcb5cd6c · outbound

This paper cites Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.825990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.825990Z digest=sha256:12b7eb81399c144fa696e99cc6dc28fde3e51a18775fb69d9ccd4ef5d6c78fda

Observation d4c42b4f-2a40-4877-8248-ec8866a67713 · outbound

This paper cites Boundary point jailbreaking of black-box llms, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Boundary point jailbreaking of black-box llms, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.829056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.829056Z digest=sha256:746adf0ca362c08d0260190ffd4c5b0649a705eeb9b4638b7b09a6943fa5067c

Observation 8e302ba0-1a01-4499-8b35-d7db3113287f · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.679664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.832135Z digest=sha256:ad1a805ce3d5e37cd35b7abe6a5ba46a8097731e23fd443e4fb39f2b2941e6c5

Observation 974e515b-6c36-4a76-b531-59b2b6b748e1 · outbound

This paper cites The safety gap toolkit: Evaluating hidden dangers of open-source models.

AI Security Leaderboard: Methodology, Results and Minimal Standard The safety gap toolkit: Evaluating hidden dangers of open-source models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.669906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.835276Z digest=sha256:d045dfc957812eab5dea926e4bde4db6f8bd97c0621b18b52086403540ccbc38

Observation c2cc9033-fcb5-404e-b5bc-6a51103e706d · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.660329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.838399Z digest=sha256:2369e0cc665979832dc6f533ee42a136e3a31cb9e7b39ec2c01f57aaf6813468

Observation 50f5ab3b-2cb7-4d9f-b6dd-cb0e3db16493 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.841373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.841373Z digest=sha256:1f24b969730ba489ff939ca4f978f67ffd6ab841e4313e6a3fa5976a624ba096

Observation 5dac27c1-ce00-407f-b11b-92a491cf6dca · outbound

This paper cites Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark.

AI Security Leaderboard: Methodology, Results and Minimal Standard Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.844948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.844948Z digest=sha256:d1ea776b82005355f7e3ad4c6289c1f22af065587bb60b14db94d5a2659c2f29

Observation 774cb533-c83e-4232-9a8b-840c9f9bfdc6 · outbound

This paper cites Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.650799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.848365Z digest=sha256:5bf0b54963f3d039c09ba0d4c1ace91295d1b560385dbf2fb3dda27354f9a456

Observation ac7b7765-dcc5-4c5c-a68d-0f42c6539bd1 · outbound

This paper cites God Has Helped Us, and So Will AI.

AI Security Leaderboard: Methodology, Results and Minimal Standard God Has Helped Us, and So Will AI

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.641029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.851461Z digest=sha256:128bf3dbaf7a27cc742b61bc20c46fcbe80fcd0fe5e917b50dc72e278a109eeb

Observation 5c60ab4c-7a46-4812-8b03-d95718c6c93e · outbound

This paper cites McKenzie, Oskar J.

AI Security Leaderboard: Methodology, Results and Minimal Standard McKenzie, Oskar J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.854468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.854468Z digest=sha256:75351183e6c58bac71242ab3a5c899c4c5930b34c111bc3605e83d7c6957883f

Observation 5352fe49-0ae7-4ba0-be5b-86492d0ddb08 · outbound

This paper cites Common elements of frontier AI safety policies, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Common elements of frontier AI safety policies, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.631341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.857439Z digest=sha256:2cf1174027c34dce41e5d653515bd0d1d1f533c5f0278caeb2afff9dc1686c7e

Observation b1277843-1da5-4094-9db4-f6d75e1bfed1 · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.860705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.860705Z digest=sha256:ed6ee0b44fe02938874bed2b82c0377b1fde57598c18f51fcc98a284e53526b4

Observation 46d6afd3-527e-4ea0-bb38-de08a31a3502 · outbound

This paper cites ChatGPT Agent System Card, July 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard ChatGPT Agent System Card, July 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.621501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.864122Z digest=sha256:c5393d2958d6fd8969b0a6229260320564a2098724e3feb57e188900a6110a8e

Observation 67ef4c4d-23e3-4bdf-9a3c-4d8d35ffc6a5 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.867237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.867237Z digest=sha256:d936fae6e0e710c6de461caaadbd2f99c492475b79788ef3c6646937ae819ec5

Observation 94c2008a-fb9e-4f1a-a5e5-861605ed6319 · outbound

This paper cites Exposing the systematic vulnerability of open-weight models to prefill attacks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Exposing the systematic vulnerability of open-weight models to prefill attacks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.871031Z digest=sha256:94fb7da4f42cb49e038c52b89e8b20405aeb4d99c61bcdfc1fbd29b295d8a078

Observation b8ef9201-8fba-4294-bb23-4a0ac23f70e8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Qwen3.5: Towards native multimodal agents, 2026

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.612012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.874332Z digest=sha256:3bcc82ee17739bac113182bb460309878b678e62c2ac28fd5303a7d9cf6e5f83

Observation ffed2393-7a4e-41ac-aa69-45d8c2381186 · outbound

This paper cites Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.602321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.877689Z digest=sha256:9360693cca2f95b677e406bcc23b72bc9c4a0dee1d8b50e4eb7ee2d7d74d7d83

Observation 8fe0585a-8534-4ddd-82f2-fbb94be7e2b4 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.880760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.880760Z digest=sha256:edd60c5a02004d8d993204262be8364ae2fc3046516052d8fd729e5a24c92750

Observation d3f94878-4a23-4e07-b4ef-da0b4d1def24 · outbound

This paper cites Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.884319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.884319Z digest=sha256:3c8de03690d04c9f9b95fcb8380f9a937c192155528b0650ac707ecb23fa0cfe

Observation 89692e7a-d5e4-4df7-8875-d5e7e26c8edc · outbound

This paper cites Jailbroken Frontier Models Retain Their Capabilities.

AI Security Leaderboard: Methodology, Results and Minimal Standard Jailbroken Frontier Models Retain Their Capabilities

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:04:40.979412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.887782Z digest=sha256:9328d5553e34fa74dbf0cc2194e765c6f0a438aafa9e6776addd0ae0e1e94836

Observation 677b28d0-f0f8-4248-88f9-255d12d54f4b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.891362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.891362Z digest=sha256:bd163063f5ce3ac37ef626fd17063da9908e48463f565d5f4b8750e9b4593e5f

Observation 407c913d-b43a-4987-b53d-247c77fac7eb · outbound

This paper cites Sorry,” “I can’t help with that, but….

AI Security Leaderboard: Methodology, Results and Minimal Standard Sorry,” “I can’t help with that, but…

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.591935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.895326Z digest=sha256:d74111d90c3fcdf6141d89c2c7d39fae25a455303a341c27efd202929372bd28

Observation 273bbf32-5b57-4029-9f26-a0ffb085e131 · outbound

This paper cites not jailbroken.

AI Security Leaderboard: Methodology, Results and Minimal Standard not jailbroken

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.582787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.898975Z digest=sha256:c9a2635768ef02c235284d62441eceee0ef7a69e706f4fa5f227b89322480712

Observation 01daf259-7b40-443d-a04f-4bc89046a70b · outbound

This paper cites jailbroken (1).

AI Security Leaderboard: Methodology, Results and Minimal Standard jailbroken (1)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.574059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.902515Z digest=sha256:76e648e7840cd9d87f6a8108218a659d702cd66ef02068addeb68127a7521d1d

Observation b3de8b6a-0ba0-4112-a24b-acb5836644cc · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.565345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.906031Z digest=sha256:0c626e92be323c1703b74204ee78ecb302db4a3026153cfea834c4d7ee069de8

Observation e870fbac-2b7d-433c-99c5-f8fe535c39f4 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.555538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.909296Z digest=sha256:e26d6f5914eb3c5a4a2ccb8c1f584ae3a089a051dc4f9dc38a72abf415ad4bd3

Observation ce5d3ada-f4f9-4bef-a33a-18668093c0ff · outbound

This paper cites the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0).

AI Security Leaderboard: Methodology, Results and Minimal Standard the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0)

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T01:04:41.545319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.912861Z digest=sha256:bb0eee97b324ef0ff211943b1b056afd687bc8620a3bebeb4e58dd3aa2d91f2a

Observation 051e85b6-83a9-447b-a859-cf1d9c650a1f · outbound

This paper cites The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price.

AI Security Leaderboard: Methodology, Results and Minimal Standard The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.535433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.917008Z digest=sha256:50e28887f9b00b718d94a57cedf794c5ebcdbff8ca7b9cbae16e85c05e242aa1

Observation f4cf8811-900f-4081-9503-7d1530ef92f7 · outbound

This paper cites If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops.

AI Security Leaderboard: Methodology, Results and Minimal Standard If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.525030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.920005Z digest=sha256:50dc8035fde7608e58f8e6f0b224eee4cc2d82a601fd70e7b1bab85a405f8cc1

Observation b63ec2b0-961e-4470-96ae-ee7e7d6f1cf7 · outbound

This paper cites Suppose we find ten working jailbreaks against each of two models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Suppose we find ten working jailbreaks against each of two models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.514526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.923218Z digest=sha256:66380a1207636d53dd9a9943d5004b756c1ff885fe814865e256922f1ae576f1

Observation e1511973-57fb-4595-a23c-3e1d34c46a7b · outbound

This paper cites Finding a jailbreak that works once is easy; proving it works reliably is expensive.

AI Security Leaderboard: Methodology, Results and Minimal Standard Finding a jailbreak that works once is easy; proving it works reliably is expensive

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.504661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.926614Z digest=sha256:60d230d8be5bf6020ed30be6f29eb4fe64093252c606cc8bea0c9d8c6e67c83d

Observation b7a4f3f6-6596-4d57-830c-85d2addbc8f7 · outbound

This paper cites A smart attacker does not run every candidate over the full sample.

AI Security Leaderboard: Methodology, Results and Minimal Standard A smart attacker does not run every candidate over the full sample

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.494329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.929731Z digest=sha256:4a36ddbd39a62b3ad7abe73d05be55e481b6008cee930164e84eda7ba494f6cf

Observation 33117418-4b40-46b2-92dc-61512cbe38de · outbound

This paper cites Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain).

AI Security Leaderboard: Methodology, Results and Minimal Standard Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.484182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.932922Z digest=sha256:19d19e4c1b00f171abb31c0a026e8e5a9843db54cbe6b5445a0b2363ce529ec0

Observation 16677385-b223-4e30-bbce-beb5e9ad92c4 · outbound

This paper cites Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%).

AI Security Leaderboard: Methodology, Results and Minimal Standard Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.473005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.936182Z digest=sha256:99ee699c2558b36981b09a8bac8517fe50ffe5909824aaf3b9bba0b47a273ad6

Pith citing papers

No inbound Pith citation observations are available.