Pith. sign in

Paper Citation Record · LEDGER

Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.09289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09289 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:00:23.917123Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9f78405e-baf5-4d77-80a5-1cc14b926903 · inbound

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit cites this paper.

JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:00:23.917123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:00:23.917123Z digest=sha256:1a0bf552ab73c5beabe823dc12edc61ddfc980e9eb0cacd082103c9d25757515

Observation 20ac46f7-0427-4642-bd30-a8b3e6065a4a · inbound

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks cites this paper.

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:32.348147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:32.348147Z digest=sha256:b3807a595894bed24eb410e3791728628ddb93037545a74f054f18f5d40a7008

Observation 65eae594-778f-44e5-b6be-372c71ee8da4 · inbound

Obfuscated Activations Bypass LLM Latent-Space Defenses cites this paper.

Obfuscated Activations Bypass LLM Latent-Space Defenses Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.439328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.439328Z digest=sha256:69cedee0c107f9759dc82d8eeab9772889ab9057d0f1e4e63da26ad0e9328706

Observation 19b22bfb-0094-447f-bea7-770a03f248e7 · inbound

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models cites this paper.

Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:12.905062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:12.905062Z digest=sha256:6183b1ae9087755df0172b0d75dd41e64fb1d639ebd5ccdd36cedc2b3f0c5459

Observation 4efdb86d-c17f-407e-a93f-312281ef46f5 · inbound

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions cites this paper.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.210046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.210046Z digest=sha256:a67c93c6fe245e6cebfd004c95a9c74b98f1813a6a3ba375400e8b2d00c1daa8

Observation c9e68394-6eef-4cf3-a4b0-e7d8d5b7f83a · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:11.902723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:11.902723Z digest=sha256:ea6a7854474d21ec3e78caa7b5f20085540ea7dc0f54b94fb0c876f8049cc6c4

Observation 062a7351-50e3-44b0-a0e1-ff7da1d591b2 · inbound

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety cites this paper.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.309978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.309978Z digest=sha256:4d882619fa0b0bb53b101550b2f6bc6665e53253426dd28c0ef4eafdff6707c3

Observation 2102cf77-cdb4-4cf3-9487-0404703cf185 · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:15.337553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:15.337553Z digest=sha256:145b5c188bcff5f44fcfa39f9c00eca394adcda6402a9a6193938f2386fb7b9b

Observation d1d58fc7-c254-4680-8429-63a6f04834d2 · inbound

SATORI: Static Test Oracle Generation for REST APIs cites this paper.

SATORI: Static Test Oracle Generation for REST APIs Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:48.609438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:48.609438Z digest=sha256:1324c5d98b68330524e827764f0ff4bd58589067d26493f9c96dccdacef138d4

Observation 6e19f11a-9ecc-4966-819d-6c4b8ca4132f · inbound

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models cites this paper.

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:32:27.687194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:32:27.687194Z digest=sha256:9da8b73d0606b2b58f35d2a94dd14a49f06d3e7f82e036ba092c03811d05a572

Observation 02a251e3-3079-411f-be53-4311f8855fdf · inbound

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems cites this paper.

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:12.115088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:12.115088Z digest=sha256:06b4725c489e4b8482150ee80d8e855053926387e54cf5a05c5b60e83b3e1379

Observation 5d89409f-fc51-40d1-bc23-ae7bdff66039 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.788037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:773ecce02b1f2df7af69405a651b4064732f61ce5292e04692a4422fe1f27a6f

Observation cd8931e8-aacf-4b8f-988c-a67b7df68f92 · inbound

Why Do Large Language Models Generate Harmful Content? cites this paper.

Why Do Large Language Models Generate Harmful Content? Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:04.915542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:31:13.545599Z digest=sha256:435acd5b057bbe88d4819f0571b26bfcf0c9a1f37cf1d9a28af47bc8ea595df8

Observation beecf59b-3f44-4637-ba4c-de0fd2b9d54f · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:22:18.924697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:ec1bfad3008bdc158a05b8fbc93938e8f6068ca3fac4561b8244e6101fb3d939

Observation fcf626c2-ee70-44d3-a918-c569af0fc9f5 · inbound

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks cites this paper.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.195732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.195732Z digest=sha256:7f8e859e203aaa70b1649d7b8cd795ccc274f3de7647962ccc79b03e0f0181ef