Pith. sign in

Paper Citation Record · LEDGER

AI Security Leaderboard: Methodology, Results and Minimal Standard

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.03070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03070 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:04:40.936182Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 486f349b-b9d2-42d9-8d97-0c6bd8f05cbd · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

AI Security Leaderboard: Methodology, Results and Minimal Standard gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.800351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.800351Z digest=sha256:c910d6740607c727d1fe7083c2c44bc084b57228ec229ff4fd6178f4ee123cc1

Observation 00822f6b-9cd3-486a-a667-1ad9952ecd6b · outbound

This paper cites Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of OpenAI’s GPT-5.5 cyber capabilities, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.717747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.804895Z digest=sha256:aa6b9485a0f51fe1097969bd95fd984c391f5d51d16d08ffcdb8b67545d33b6c

Observation 921e6ce0-e3b2-4dbe-9683-f8857cdb3622 · outbound

This paper cites Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Our evaluation of Claude Mythos Preview’s cyber capabilities, 2026

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.708262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.808590Z digest=sha256:dbfff0f234a9b95da62f8b6ee2aaeda4849008ff6790474e57625504b763ee0f

Observation 68433ac3-357a-459b-b927-30ab800cbf9e · outbound

This paper cites Constitutional Classifiers: Defending against universal jailbreaks, February 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against universal jailbreaks, February 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.699247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.812066Z digest=sha256:cd572565dde05bc77b801b49186679a5598ec07ea5faa504edb8aa83b410ec7c

Observation 5e9af83b-b20c-4971-8b95-8ac4c98bfb4e · outbound

This paper cites Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Disrupting the first reported ai-orchestrated cyber espionage campaign, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.689793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.815425Z digest=sha256:33be1ae28524e837b65fc8392b852cf2c3b4b91e2ca5a24bc7ae6a55458b9c5d

Observation 8ebe1614-0dac-47a8-b3d5-f23ef00d43fd · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

AI Security Leaderboard: Methodology, Results and Minimal Standard Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.818991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.818991Z digest=sha256:38ddf6acabaffae3289cbb7652f99a5cd34c1ede27075b11ee34a3c6a3c00d8a

Observation dbe28a52-b414-451e-999a-f397cc2723b0 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.822825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.822825Z digest=sha256:8de621b5046d85e128897b09a7b89b23d21b1bf955a9062a98c861a750ef452e

Observation bd399f56-0609-4170-915b-3919fcb5cd6c · outbound

This paper cites Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional classifiers++: Efficient production-grade defenses against universal jailbreaks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.825990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.825990Z digest=sha256:705623db4a4fb3a0e0ae3e189f5fbe995a5b738db62fc2c2268b97f9f7b6808f

Observation d4c42b4f-2a40-4877-8248-ec8866a67713 · outbound

This paper cites Boundary point jailbreaking of black-box llms, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Boundary point jailbreaking of black-box llms, 2026

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.829056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.829056Z digest=sha256:99c38cbd1b286444b152d176f2a8839553e022f98dadd3a667cc2cca62e5b817

Observation 8e302ba0-1a01-4499-8b35-d7db3113287f · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.679664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.832135Z digest=sha256:8cd0ee5e9b0b72c3483e560f33e01c20bb93466e5f971eeab1313e8254a9bc66

Observation 974e515b-6c36-4a76-b531-59b2b6b748e1 · outbound

This paper cites The safety gap toolkit: Evaluating hidden dangers of open-source models.

AI Security Leaderboard: Methodology, Results and Minimal Standard The safety gap toolkit: Evaluating hidden dangers of open-source models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.669906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.835276Z digest=sha256:f85d7d2fc24288cf8cd0df2db09019e0238d4a711c1af10befbd5a790d7c948b

Observation c2cc9033-fcb5-404e-b5bc-6a51103e706d · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.660329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.838399Z digest=sha256:81e9d8b260b6d1723d61ea1fab463263a2484e6ca5b177386ebfdf5b4c3f8abb

Observation 50f5ab3b-2cb7-4d9f-b6dd-cb0e3db16493 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.841373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.841373Z digest=sha256:f111912b1e085d6654590a151aa87c49a4b67dd9f6d943331339c4fe41d19b07

Observation 5dac27c1-ce00-407f-b11b-92a491cf6dca · outbound

This paper cites Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark.

AI Security Leaderboard: Methodology, Results and Minimal Standard Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.844948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.844948Z digest=sha256:7c170337b3e82397e5481762e9d3c2dc95a8ad44719115c30936e320ecf12d49

Observation 774cb533-c83e-4232-9a8b-840c9f9bfdc6 · outbound

This paper cites Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Openai let chatgpt aid and abet mass shooters, florida lawsuit claims, 2026

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.650799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.848365Z digest=sha256:188fcd1e628f2f975038bcb8ada3e192fc5d707f327147c6ed48b3239283552e

Observation ac7b7765-dcc5-4c5c-a68d-0f42c6539bd1 · outbound

This paper cites God Has Helped Us, and So Will AI.

AI Security Leaderboard: Methodology, Results and Minimal Standard God Has Helped Us, and So Will AI

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.641029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.851461Z digest=sha256:813e910fd0fabc3918c1a04bee2598790493715ce58399b88e51da78ace6a582

Observation 5c60ab4c-7a46-4812-8b03-d95718c6c93e · outbound

This paper cites McKenzie, Oskar J.

AI Security Leaderboard: Methodology, Results and Minimal Standard McKenzie, Oskar J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.854468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.854468Z digest=sha256:d63bfdb6ad8e8ceafd8b293c2ba244ab31b40d11272dafece1ae3fe63cad992f

Observation 5352fe49-0ae7-4ba0-be5b-86492d0ddb08 · outbound

This paper cites Common elements of frontier AI safety policies, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Common elements of frontier AI safety policies, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.631341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.857439Z digest=sha256:9744270591fb20f714e6dea2000d40004a8dd6862afd84f16f9bf63a3035fbcc

Observation b1277843-1da5-4094-9db4-f6d75e1bfed1 · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.860705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.860705Z digest=sha256:bbae58d340c1a62caed18a706df9ff3c8ccdd1aa2932c7e0f241d3e02ecea6c2

Observation 46d6afd3-527e-4ea0-bb38-de08a31a3502 · outbound

This paper cites ChatGPT Agent System Card, July 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard ChatGPT Agent System Card, July 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.621501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.864122Z digest=sha256:6873dfbe4dee00a4a92b99262584b1adcea82df6dbe3ef0fcc263b888dde983c

Observation 67ef4c4d-23e3-4bdf-9a3c-4d8d35ffc6a5 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

AI Security Leaderboard: Methodology, Results and Minimal Standard Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.867237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.867237Z digest=sha256:11322f02baa6cf9a11b4446313cb8a30adb81fe345058afff2e5b599596849a5

Observation 94c2008a-fb9e-4f1a-a5e5-861605ed6319 · outbound

This paper cites Exposing the systematic vulnerability of open-weight models to prefill attacks.

AI Security Leaderboard: Methodology, Results and Minimal Standard Exposing the systematic vulnerability of open-weight models to prefill attacks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.871031Z digest=sha256:25a92275e13f55b0b37d06c0fdb09b6aca79ad2b7b86b414d75f47709c79fd5a

Observation b8ef9201-8fba-4294-bb23-4a0ac23f70e8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Qwen3.5: Towards native multimodal agents, 2026

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.612012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.874332Z digest=sha256:95192a0a5a9e360d9fa259e63bbc185c9c3b051db08294e8371da2f0bf75ea17

Observation ffed2393-7a4e-41ac-aa69-45d8c2381186 · outbound

This paper cites Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025.

AI Security Leaderboard: Methodology, Results and Minimal Standard Green beret who exploded cybertruck in las vegas used ai to plan blast, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.602321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.877689Z digest=sha256:aaedc885c495e5bf35f5412afbc5aeb391f0a5f2b0ed2cb9ec07bd415f6a6e2b

Observation 8fe0585a-8534-4ddd-82f2-fbb94be7e2b4 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

AI Security Leaderboard: Methodology, Results and Minimal Standard The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.880760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.880760Z digest=sha256:96b37e1a226c73ef30032d34282507e09aa94d117eba712047b62653638557d0

Observation d3f94878-4a23-4e07-b4ef-da0b4d1def24 · outbound

This paper cites Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026.

AI Security Leaderboard: Methodology, Results and Minimal Standard Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.884319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.884319Z digest=sha256:a08dac56d3cafd3282a0d922de466e9bed43c266409b2ea806937bf75155eaa4

Observation 89692e7a-d5e4-4df7-8875-d5e7e26c8edc · outbound

This paper cites Jailbroken Frontier Models Retain Their Capabilities.

AI Security Leaderboard: Methodology, Results and Minimal Standard Jailbroken Frontier Models Retain Their Capabilities

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:04:40.979412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.887782Z digest=sha256:4c78f450fcdcc5cf85e5db78a28afeea70568134693fb0b7c287c0dc0e3f720a

Observation 677b28d0-f0f8-4248-88f9-255d12d54f4b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T01:04:40.891362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:04:40.891362Z digest=sha256:d41e49ab0f8d09f8856f70c73c393fd59139171764990f13d396a35afce294d0

Observation 407c913d-b43a-4987-b53d-247c77fac7eb · outbound

This paper cites Sorry,” “I can’t help with that, but….

AI Security Leaderboard: Methodology, Results and Minimal Standard Sorry,” “I can’t help with that, but…

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.591935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.895326Z digest=sha256:dfa9d114d04031ab59b793890e855f4b99891664fb5c6051be6d79ed585a63b9

Observation 273bbf32-5b57-4029-9f26-a0ffb085e131 · outbound

This paper cites not jailbroken.

AI Security Leaderboard: Methodology, Results and Minimal Standard not jailbroken

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.582787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.898975Z digest=sha256:efe23c25cbf00a94215c728f3e1186bec6c857c9dda1c1fa97e504a6c7992f95

Observation 01daf259-7b40-443d-a04f-4bc89046a70b · outbound

This paper cites jailbroken (1).

AI Security Leaderboard: Methodology, Results and Minimal Standard jailbroken (1)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.574059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.902515Z digest=sha256:6ad63292ae0a642da3739a380ef87b3ca5364cbe12630ceaa4a5fdbbe47cd887

Observation b3de8b6a-0ba0-4112-a24b-acb5836644cc · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.565345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.906031Z digest=sha256:0dd81f910770c5884e1cc2e36b05d1a7d0fe25dfcc5a21b7d0da8022028fc043

Observation e870fbac-2b7d-433c-99c5-f8fe535c39f4 · outbound

This paper cites an unresolved cited work.

AI Security Leaderboard: Methodology, Results and Minimal Standard Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T01:04:41.555538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.909296Z digest=sha256:3e6bd9104df36b63ddfe315b55b5ee69a9eb16402c347353f3dcd250d56883c5

Observation ce5d3ada-f4f9-4bef-a33a-18668093c0ff · outbound

This paper cites the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0).

AI Security Leaderboard: Methodology, Results and Minimal Standard the jailbreak technique involves simplifying the attacker goal too much, label as ”not jailbroken” (0)

Reference 34

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T01:04:41.545319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.912861Z digest=sha256:3c389cdc1be7e4a39f8823b827ea04a2d9c881b6a650743a431907c23520aad1

Observation 051e85b6-83a9-447b-a859-cf1d9c650a1f · outbound

This paper cites The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price.

AI Security Leaderboard: Methodology, Results and Minimal Standard The cost of one query used to test whether a jailbreak works, from its token usage and the model’s API price

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.535433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.917008Z digest=sha256:47827a8f571e4304176d00c7d45ffc47b2d318a20addeee16c2421f0335abbbf

Observation f4cf8811-900f-4081-9503-7d1530ef92f7 · outbound

This paper cites If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops.

AI Security Leaderboard: Methodology, Results and Minimal Standard If working jailbreaks are common in the pool, the attacker finds a universal one sooner, so the cost drops

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.525030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.920005Z digest=sha256:a9bc8e28d8212691d686e073c8a70d8c6d8e4d23b6f1fb02159689d10b7de57b

Observation b63ec2b0-961e-4470-96ae-ee7e7d6f1cf7 · outbound

This paper cites Suppose we find ten working jailbreaks against each of two models.

AI Security Leaderboard: Methodology, Results and Minimal Standard Suppose we find ten working jailbreaks against each of two models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.514526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.923218Z digest=sha256:7fe9b918e79d17b8c2739666c732068bc03751b2f8edb7a7b11c5f7447e7091b

Observation e1511973-57fb-4595-a23c-3e1d34c46a7b · outbound

This paper cites Finding a jailbreak that works once is easy; proving it works reliably is expensive.

AI Security Leaderboard: Methodology, Results and Minimal Standard Finding a jailbreak that works once is easy; proving it works reliably is expensive

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.504661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.926614Z digest=sha256:221f71042f5ae0443d4e9ecd5b3594445eee0b731ab1877084ffbbe4b53e9118

Observation b7a4f3f6-6596-4d57-830c-85d2addbc8f7 · outbound

This paper cites A smart attacker does not run every candidate over the full sample.

AI Security Leaderboard: Methodology, Results and Minimal Standard A smart attacker does not run every candidate over the full sample

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.494329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.929731Z digest=sha256:ea8d783fcf74456625e67d2b9453556f7392f1830b75968158a8076f026b1e81

Observation 33117418-4b40-46b2-92dc-61512cbe38de · outbound

This paper cites Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain).

AI Security Leaderboard: Methodology, Results and Minimal Standard Test the candidate on a small number of samples (e.g., in our testing, we use 8 per domain)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.484182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.932922Z digest=sha256:7804cb036970ec6ffd70fb51711538b0655f0812568799cf280715bd50f5cfc2

Observation 16677385-b223-4e30-bbce-beb5e9ad92c4 · outbound

This paper cites Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%).

AI Security Leaderboard: Methodology, Results and Minimal Standard Run each surviving candidate on many more samples, e.g., on the order of a hundred to confirm it is universal (in our testing, it jailbreaks at least 75%)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:04:41.473005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T01:04:40.936182Z digest=sha256:bd13e29d78d53dc9ea9b6a188980898ab7360333bcff7f795a323188337d7768

Pith citing papers

No inbound Pith citation observations are available.