Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:12.847460Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.04365.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:12.847460Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:54:06.538826Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T13:26:59.311234Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed19b6ad-0d61-44f6-82e2-19592791cf44 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Bowman, Ethan Perez, Roger B
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b845d7d-de3a-4664-96bf-60899d8be88b · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66d6f79-7d2c-4b13-b5e1-5eb1f2785950 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Zico Kolter
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d07ac786-cf88-49bb-96db-b3fbc6d406e1 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46fb15d-4d9e-4bac-8e2d-91c6fcac36ae · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48d1a98c-12c5-459b-af6c-feb27d1f27f5 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Token highlighter: Inspecting and mitigating jailbreak prompts for large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf8fcf33-b58b-4596-9a6d-f3fd30712868 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b3f59d-53ea-4897-9fb7-2616726aaa25 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e8a31d-8f11-4111-b10a-3e6f44fddd25 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c10190c4-b752-494d-a3d4-1bc75b923637 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b3415f-05b0-4731-a310-adfb9bc3f1ba · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa504a37-0b53-40b5-90e8-832619a453e6 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs GPT-4 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15698224-46b7-4e96-9c96-bffdb51d7ee5 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 803559cc-ce6a-414a-801e-0e4e2297a345 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7a98138-4888-49b4-919a-bd33944a6e4b · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Qwen2.5 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c8d635-0166-48a9-8601-43fe20c4b908 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed5bec15-6373-4d12-a047-fdc5659bfdd5 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Yu, Qingsong Wen, and Yang Liu
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b8c7c9a-59fd-4c5b-854f-6071827ddde2 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbroken: How Does LLM Safety Training Fail?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281b06a9-2921-4fd4-9b5a-43addc867693 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Defending chatgpt against jailbreak attack via self-reminders
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44031e8f-993a-4c66-8791-99f1696146f0 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d4ba01d-58e9-4a56-8fad-03f5838ddbae · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e392ee3c-fb5c-481c-9615-e62c0bd6c656 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Low-Resource Languages Jailbreak GPT-4
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba7b66aa-1fbd-4999-9ebd-7bdd9a34e7b0 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046c14d0-ecf8-4438-a0aa-a7d59829a5b7 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86198f14-e088-479d-9735-6c0ccc9a442a · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f7bb32d-8ba1-40fc-a6e7-876053682a09 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 234e7a8c-c05b-45d5-8c90-ef33670f17bb · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50d55ebd-4441-4f3a-8d0a-1931f66db025 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9d4c793-97d0-4547-a529-97ef0be96397 · outbound
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f1b3b36-483b-4781-ada2-c313b1e68a85 · inbound
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c705765-eeba-40e8-95fc-02ca55f8dee6 · inbound
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.