Pith. sign in

Paper Citation Record · LEDGER

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.04365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04365 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:12.847460Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:54:06.538826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:26:59.311234Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed19b6ad-0d61-44f6-82e2-19592791cf44 · outbound

This paper cites Bowman, Ethan Perez, Roger B.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Bowman, Ethan Perez, Roger B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.871531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:09.120032Z digest=sha256:123cbbfb627de2ddf83ff73a2d9e20a2da1aed69a08c4b3859cbbaff80aba757

Observation 5b845d7d-de3a-4664-96bf-60899d8be88b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.296540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.296540Z digest=sha256:1713592e07cfd0fe9c27bd5f772399c8c7e2fc0173f8986a16a3005d8fb67e89

Observation f66d6f79-7d2c-4b13-b5e1-5eb1f2785950 · outbound

This paper cites Zico Kolter.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Zico Kolter

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.525171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:09.403630Z digest=sha256:9f099f8e25a0faf935e1d8f4aff956e26d56274d0dc91f73da0d2e785255086f

Observation d07ac786-cf88-49bb-96db-b3fbc6d406e1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.473833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.473833Z digest=sha256:cd764124e3ed921bda476832a790d26511b3a97cf46910be24cce4bc143cd4ed

Observation d46fb15d-4d9e-4bac-8e2d-91c6fcac36ae · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:16.254930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:09.551089Z digest=sha256:e134233e319d81d036887981f91b3cc9e926ee3cc78935a0911bd359d4dbd916

Observation 48d1a98c-12c5-459b-af6c-feb27d1f27f5 · outbound

This paper cites Token highlighter: Inspecting and mitigating jailbreak prompts for large language models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Token highlighter: Inspecting and mitigating jailbreak prompts for large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:09.611247Z digest=sha256:2a8a55497b387d2bc97bb8e102b3786d2c26664a44429c3dcae5b9854ef94630

Observation bf8fcf33-b58b-4596-9a6d-f3fd30712868 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.684400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.684400Z digest=sha256:3e8e9afe4e40993b0e1a417088c9b94b29a5d86ef6cc25ff0bd498cd1bff8cf5

Observation 87b3f59d-53ea-4897-9fb7-2616726aaa25 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.839439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.839439Z digest=sha256:0a5a9aefd2c29b834f1df3d373c4a049ff03ad1bbf0b0794910b6c1dd4681524

Observation 65e8a31d-8f11-4111-b10a-3e6f44fddd25 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.982978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.982978Z digest=sha256:f7e6555edb31908b473877af16dffc4aaac3adef027fb6bb82b56a4c211bcd08

Observation c10190c4-b752-494d-a3d4-1bc75b923637 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.094408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.094408Z digest=sha256:9dd88704d9456d1caf328441d183f8303ecade0a7bad6fa5ba2e55bea4fa9809

Observation 46b3415f-05b0-4731-a310-adfb9bc3f1ba · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.231324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.231324Z digest=sha256:4fda77ec137f72d21bb16960f2841c8bde20eee7025dd254f597f0639cdd325c

Observation aa504a37-0b53-40b5-90e8-832619a453e6 · outbound

This paper cites GPT-4 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.352809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.352809Z digest=sha256:9f74c615113bf8e3ab00e47e3879dc99860e9079e885126bad9d5c537e0150aa

Observation 15698224-46b7-4e96-9c96-bffdb51d7ee5 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.497216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.497216Z digest=sha256:240265f318dcb07bbfa7e063cf757261fa61a7ef24a3c415709827396d94b69e

Observation 803559cc-ce6a-414a-801e-0e4e2297a345 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:15.761278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:10.684287Z digest=sha256:0838787cf12ea0bd5af4f2f638b8447e7ebc25d069391d6f41a6adce6dbef5ff

Observation a7a98138-4888-49b4-919a-bd33944a6e4b · outbound

This paper cites Qwen2.5 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.804884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.804884Z digest=sha256:de0b99d3ef78feda11cd81227bed0bac83d78813620eb85e6288f5d00dcc7c3c

Observation 34c8d635-0166-48a9-8601-43fe20c4b908 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.522650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:10.923995Z digest=sha256:1415e055ef0af71526c3402371636e8e7b4cee9fc88237c208f178307b3c98fc

Observation ed5bec15-6373-4d12-a047-fdc5659bfdd5 · outbound

This paper cites Yu, Qingsong Wen, and Yang Liu.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Yu, Qingsong Wen, and Yang Liu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.229719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:11.081110Z digest=sha256:c013fa90b2f91f964c80ce89e855ab5d2ce62b18c6f61cccd2a21f76b3cd22be

Observation 5b8c7c9a-59fd-4c5b-854f-6071827ddde2 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.189415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.189415Z digest=sha256:4c30ae901c78fbdcec879e29892c243293d1991bd280090ab2e61fed96ff769d

Observation 281b06a9-2921-4fd4-9b5a-43addc867693 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Defending chatgpt against jailbreak attack via self-reminders

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.987500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:11.364637Z digest=sha256:bc26446808af85cdb558794b17be18fd698e41f7180a61d0c93b50995fb79436

Observation 44031e8f-993a-4c66-8791-99f1696146f0 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.695109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:11.530520Z digest=sha256:9222d1a2b8aa8ef807446d9731dacb134c4afc94af5f11d60b6b1bcf56e0f3df

Observation 3d4ba01d-58e9-4a56-8fad-03f5838ddbae · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.686321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.686321Z digest=sha256:80367cb469ecce8faee0deb6cd55b9151e5ca5890c70e07426ba98541e26a5d5

Observation e392ee3c-fb5c-481c-9615-e62c0bd6c656 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Low-Resource Languages Jailbreak GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.798190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.798190Z digest=sha256:8e723cccda5cef0608233237b014e399e82515aa03a3ec6b287f36131edcd3ec

Observation ba7b66aa-1fbd-4999-9ebd-7bdd9a34e7b0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.959670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.959670Z digest=sha256:2b5242d1bd07e070264d135e8098ec90c676ac19f9a2eed649d2055c4f355d0e

Observation 046c14d0-ecf8-4438-a0aa-a7d59829a5b7 · outbound

This paper cites Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.418062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.083145Z digest=sha256:1a85c5999cf35c8fe76b905a08561b4fa90df5979dc294859317c0df05143866

Observation 86198f14-e088-479d-9735-6c0ccc9a442a · outbound

This paper cites The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.162079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.219701Z digest=sha256:185307231c004fbbb49de0aca424d1b2e03b233a6ad74dfb5f25d425aef0cba2

Observation 8f7bb32d-8ba1-40fc-a6e7-876053682a09 · outbound

This paper cites With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:13.924016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.366585Z digest=sha256:e194901256a1edd50de4453cc956aa37c3643b8c03c5228b8c34c316763986e4

Observation 234e7a8c-c05b-45d5-8c90-ef33670f17bb · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.681569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.524515Z digest=sha256:e8e122b307c59c5c2eb26d54f3c844367f1d54411703db1ac3ebd309a56dc2ef

Observation 50d55ebd-4441-4f3a-8d0a-1931f66db025 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.410898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.695682Z digest=sha256:5156386dbd530366a4eb29911e5dfead644396a7a2d6987e6ca3a17993124057

Observation e9d4c793-97d0-4547-a529-97ef0be96397 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.209306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:54:12.847460Z digest=sha256:6bd471c8f3c4b0c2e33b37cdea409af8d301ff6ba56b292cbc9b2cc7a13386d1

Pith citing papers

Observation 5f1b3b36-483b-4781-ada2-c313b1e68a85 · inbound

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention cites this paper.

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:54:06.538826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:54:06.538826Z digest=sha256:f6086f5444c959a433c25791cbce8646f090685542ba5929e9896a4b9a2319e2

Observation 3c705765-eeba-40e8-95fc-02ca55f8dee6 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:59.312672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:98e956fab5332734d789539f4785763c72cc8d90c4ab10508d818677aa113142