Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.09624.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:57:41.323535Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 11f76a94-d370-4d21-be28-6eadc3d8b2c6 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Detecting Language Model Attacks with Perplexity
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f89b53b5-dca4-4b80-b5ba-2301e2351e20 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b83a876-54e8-415d-9f83-f43a69bc846b · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Many-shot jailbreaking.NeurIPS, 37:129696–129742, 2024
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a04cc204-349b-43ec-b09d-ddc3bea168d4 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Refusal in language models is mediated by a single direction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d2f74b0-54f9-4ce0-88f5-f5513b2b2cb5 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a4970e7e-fd06-4e0d-bd95-81a509f36856 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51222e43-59ca-49bb-a4c2-74c796d1d1f5 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8eb405-9d1f-4f3a-9839-316a080594e0 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9802d500-019d-42fb-aa2b-d47c3f598744 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Designing and interpret- ing probes with control tasks
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 043c4b59-721a-44df-a834-1fed402e9e18 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Best-of-N Jailbreaking
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b555bad-6552-4f44-8bb1-b2a7197f9a9f · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Attention Tracker: Detecting Prompt Injection Attacks in LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 287afe03-ecfc-4207-9b98-fa687975c303 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6fd031-542e-4f9a-8a8d-d847d0dc8f38 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aea13e5c-18fb-4152-8bf2-93a40c74b98b · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks WildTeaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 160cb7af-1975-46cf-9fd4-df20e894231f · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 576ff26a-77ce-40e3-85ba-3127b401354d · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks What features in prompts jailbreak LLMs? investigating the mechanisms behind attacks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b88a2ff3-48d9-437b-8f6e-4c6d4157d603 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A simple unified framework for detecting out- of-distribution samples and adversarial attacks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83c8478d-1212-4d6f-8f7b-1d48cc5fbd13 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The power of scale for parameter-efficient prompt tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6dcb5a50-2ca9-45d2-831e-0245c8884244 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Prefix-tuning: Optimiz- ing continuous prompts for generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a46e70c-dcd8-49cf-8743-d965ef4f9f7e · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Towards under- standing jailbreak attacks in LLMs: A representation space analysis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 377f6f05-7647-4b7b-9270-e986d404555d · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN-Turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7dc85007-6352-4923-b4ec-b9c434b1d8e3 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation be4a1f39-a700-40e1-a3a4-3b7bf060fc36 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks The tight constant in the Dvoretzky– Kiefer–Wolfowitz inequality.The Annals of Probability, 18(3):1269–1283, 1990
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 55173b51-9ee4-4f7c-93a6-1d736bb659f6 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks HarmBench: A standardized evaluation framework for automated red teaming and robust re- fusal
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0234a541-4364-4790-a6b3-1e78e02536ea · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Tree of attacks: Jailbreaking black-box LLMs automatically
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30f0bf98-5fd6-4687-b53c-c1224faa306a · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5b8e7e8-33f7-49e8-98d9-ac29d0d6d92d · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Rebuff: A self-hardening prompt injection detector
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 003fb30d-91a6-4429-8abd-2ca3371ff7c0 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks I can’t believe it’s not robust: Catas- trophic collapse of safety classifiers under embedding drift.arXiv preprint arXiv:2603.01297, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f717bca-7e3f-46ac-8393-3b7154d6af70 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Do Anything Now
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 99c6b68e-5387-4f80-8cf3-a42747243495 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 230cbbd7-198a-457f-943e-17a00cd94097 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks A StrongREJECT for empty jailbreaks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63302013-a04c-44d4-b23e-efff0de17d93 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877cd1fd-863c-4801-9533-8dec62aaf5c0 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 968afd6f-ffa7-4c97-bb32-657975f181b4 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77cabf51-e992-43cd-ab3e-84a6343143ba · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Representation Engineering: A Top-Down Approach to AI Transparency
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c6f2e7b-e024-4d6f-b550-23e3816fc8a9 · outbound
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.