Pith. sign in

Paper Citation Record · LEDGER

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2507.04365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04365 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:54:12.847460Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:54:06.538826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:26:59.311234Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed19b6ad-0d61-44f6-82e2-19592791cf44 · outbound

This paper cites Bowman, Ethan Perez, Roger B.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Bowman, Ethan Perez, Roger B

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.871531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:09.120032Z digest=sha256:b869441cb87df079bdfb32c322f6a684e1492f3ce43ac3860be6cdbabf1def1c

Observation 5b845d7d-de3a-4664-96bf-60899d8be88b · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.296540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.296540Z digest=sha256:1713592e07cfd0fe9c27bd5f772399c8c7e2fc0173f8986a16a3005d8fb67e89

Observation f66d6f79-7d2c-4b13-b5e1-5eb1f2785950 · outbound

This paper cites Zico Kolter.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Zico Kolter

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.525171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:09.403630Z digest=sha256:361569bca68877e275207de20e2cbbc16945720dad78e540fdc2f7ff34ddd484

Observation d07ac786-cf88-49bb-96db-b3fbc6d406e1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.473833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.473833Z digest=sha256:cd764124e3ed921bda476832a790d26511b3a97cf46910be24cce4bc143cd4ed

Observation d46fb15d-4d9e-4bac-8e2d-91c6fcac36ae · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:16.254930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:09.551089Z digest=sha256:e861da7e2246678063ee297dfac124e9e03c20da06a211968508081f97180560

Observation 48d1a98c-12c5-459b-af6c-feb27d1f27f5 · outbound

This paper cites Token highlighter: Inspecting and mitigating jailbreak prompts for large language models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Token highlighter: Inspecting and mitigating jailbreak prompts for large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:16.027866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:09.611247Z digest=sha256:cd0eefc6cd1426d30ce384f3da545598ecdeb2c72655f690fc7014c56694391d

Observation bf8fcf33-b58b-4596-9a6d-f3fd30712868 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.684400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.684400Z digest=sha256:3e8e9afe4e40993b0e1a417088c9b94b29a5d86ef6cc25ff0bd498cd1bff8cf5

Observation 87b3f59d-53ea-4897-9fb7-2616726aaa25 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.839439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.839439Z digest=sha256:0a5a9aefd2c29b834f1df3d373c4a049ff03ad1bbf0b0794910b6c1dd4681524

Observation 65e8a31d-8f11-4111-b10a-3e6f44fddd25 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:09.982978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:09.982978Z digest=sha256:f7e6555edb31908b473877af16dffc4aaac3adef027fb6bb82b56a4c211bcd08

Observation c10190c4-b752-494d-a3d4-1bc75b923637 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.094408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.094408Z digest=sha256:9dd88704d9456d1caf328441d183f8303ecade0a7bad6fa5ba2e55bea4fa9809

Observation 46b3415f-05b0-4731-a310-adfb9bc3f1ba · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.231324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.231324Z digest=sha256:4fda77ec137f72d21bb16960f2841c8bde20eee7025dd254f597f0639cdd325c

Observation aa504a37-0b53-40b5-90e8-832619a453e6 · outbound

This paper cites GPT-4 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.352809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.352809Z digest=sha256:9f74c615113bf8e3ab00e47e3879dc99860e9079e885126bad9d5c537e0150aa

Observation 15698224-46b7-4e96-9c96-bffdb51d7ee5 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.497216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.497216Z digest=sha256:240265f318dcb07bbfa7e063cf757261fa61a7ef24a3c415709827396d94b69e

Observation 803559cc-ce6a-414a-801e-0e4e2297a345 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:15.761278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:10.684287Z digest=sha256:4e3bf89a379337eedd3553a4bf91a84118461270dae83e1bb12d1710abab229a

Observation a7a98138-4888-49b4-919a-bd33944a6e4b · outbound

This paper cites Qwen2.5 Technical Report.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:10.804884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:10.804884Z digest=sha256:de0b99d3ef78feda11cd81227bed0bac83d78813620eb85e6288f5d00dcc7c3c

Observation 34c8d635-0166-48a9-8601-43fe20c4b908 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.522650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:10.923995Z digest=sha256:663dc40eccb7645a59c9b94a9fab7cf34efc0c3af3d5fcd567941208772d49cb

Observation ed5bec15-6373-4d12-a047-fdc5659bfdd5 · outbound

This paper cites Yu, Qingsong Wen, and Yang Liu.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Yu, Qingsong Wen, and Yang Liu

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:15.229719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:11.081110Z digest=sha256:c1c3fa11df53e8a618ba12deefca9eea505f7bf5977e8adc996ea14fb2243760

Observation 5b8c7c9a-59fd-4c5b-854f-6071827ddde2 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.189415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.189415Z digest=sha256:4c30ae901c78fbdcec879e29892c243293d1991bd280090ab2e61fed96ff769d

Observation 281b06a9-2921-4fd4-9b5a-43addc867693 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Defending chatgpt against jailbreak attack via self-reminders

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.987500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:11.364637Z digest=sha256:00c0f407c0d395f92974fcc77521cbb711222c6755150a3efc613611ace088dd

Observation 44031e8f-993a-4c66-8791-99f1696146f0 · outbound

This paper cites Safedecoding: Defending against jailbreak attacks via safety-aware decoding.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Safedecoding: Defending against jailbreak attacks via safety-aware decoding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.695109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:11.530520Z digest=sha256:91da076dc30bfffe752b54138e0b1996e8d1da158b710e6df5ba14c4c330cf93

Observation 3d4ba01d-58e9-4a56-8fad-03f5838ddbae · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.686321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.686321Z digest=sha256:80367cb469ecce8faee0deb6cd55b9151e5ca5890c70e07426ba98541e26a5d5

Observation e392ee3c-fb5c-481c-9615-e62c0bd6c656 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Low-Resource Languages Jailbreak GPT-4

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.798190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.798190Z digest=sha256:8e723cccda5cef0608233237b014e399e82515aa03a3ec6b287f36131edcd3ec

Observation ba7b66aa-1fbd-4999-9ebd-7bdd9a34e7b0 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:54:11.959670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:54:11.959670Z digest=sha256:2b5242d1bd07e070264d135e8098ec90c676ac19f9a2eed649d2055c4f355d0e

Observation 046c14d0-ecf8-4438-a0aa-a7d59829a5b7 · outbound

This paper cites Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Each matrix has dimensions d × d, and there are four such matrices: Memory for attention matrices = 4d2

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.418062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.083145Z digest=sha256:caf611d86cc879c620118c002d243f5bbfe6fe6b2764b16059b6dbb1a51a60eb

Observation 86198f14-e088-479d-9735-6c0ccc9a442a · outbound

This paper cites The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs The first maps the input dimension d to an intermediate dimension 4d, and the second maps back to d

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:14.162079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.219701Z digest=sha256:472f358a64585d02dfb98727a79142b8b55710084111d917444720ef8afa1903

Observation 8f7bb32d-8ba1-40fc-a6e7-876053682a09 · outbound

This paper cites With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs With n + m tokens in total (e.g., n input tokens and m output tokens), the memory required for Keys and Values per layer is: Key/Value Memory per layer (bytes) = 4(n + m)d

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:54:13.924016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.366585Z digest=sha256:78e6d0b622e1db868bcb6f1e9b87c6e056db89943d3c5fb0cb08ca0b0af00a81

Observation 234e7a8c-c05b-45d5-8c90-ef33670f17bb · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.681569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.524515Z digest=sha256:4974b8906e75eb0538d55ccedc2aebe0e4c55ee982c3f8ab93cf847e8febc61b

Observation 50d55ebd-4441-4f3a-8d0a-1931f66db025 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.410898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.695682Z digest=sha256:99bce530228ccefd3be007565165afccebea784efd444255de8583588f1ee4b6

Observation e9d4c793-97d0-4547-a529-97ef0be96397 · outbound

This paper cites an unresolved cited work.

Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:54:13.209306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:54:12.847460Z digest=sha256:2bd5520641a235508a611f8127064263d35c1a88b0144c0cf58c29cd5e1b36e4

Pith citing papers

Observation 5f1b3b36-483b-4781-ada2-c313b1e68a85 · inbound

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention cites this paper.

CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:54:06.538826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:54:06.538826Z digest=sha256:f6086f5444c959a433c25791cbce8646f090685542ba5929e9896a4b9a2319e2

Observation 3c705765-eeba-40e8-95fc-02ca55f8dee6 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:59.312672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:70618fe136f4042814b403354edac6bf78a56f7a52d9b70e0986408f24799d60