Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:43.290503Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2504.13052.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:43.290503Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 97ca7d1c-0d5e-4c38-b937-a3c5c8653d16 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118e1112-7d1c-4038-9cf1-fce9188c2b53 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Constitutional AI: Harmlessness from AI Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 811a994f-9aea-4927-95f7-ad8f026b6b9b · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c7c8c7-4524-4028-a740-2539b93b3b0c · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d880b1e-11fa-4584-9f35-c8120ce55f4a · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d3b2548c-b64d-47fb-ba9d-5a0edf2cdb43 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 652e7996-236d-4768-a90b-9f7bc16bca9d · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3a3249-ff8b-473c-a62f-25e760b3eceb · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms PaLM: Scaling Language Modeling with Pathways
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e87f891-c18e-4205-b13f-2a22924b6ac8 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d6015ac9-89f9-4acb-b4b5-84bc9fe685e2 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d87ae1a4-a71d-4e46-a44c-4e4e780eddf4 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9806ca26-e42d-4d14-9dea-65d88ec9637b · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 266546f7-6626-4e6e-9e03-25fdf72376ca · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d525fef6-adf0-4d78-b47e-90141a4fa4d7 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b194a0e7-825a-42a4-bc14-3bccf7c95fbd · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cd4ef144-6e1c-4893-98e5-78c173712ea7 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 308f7068-d749-40fe-8af6-a0e05830fabc · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Dimensionless Policies based on the Buckingham $\pi$ Theorem: Is This a Good Way to Generalize Numerical Results?
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e39b708-40ce-4aa9-96c4-7f7bbfb56129 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb12c34-8145-4982-9f69-91fa706c6313 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee33961-7857-401f-bb49-9a725cc6d57a · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7bcdf1c7-5f21-4f98-80ae-771f9ac3ef26 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ac4a5a95-5af9-498c-a05e-83a84d76d493 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ec0796-942a-4426-9333-925076d2bd28 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1b977f-11de-42e1-ad99-21424737cd7a · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Red Teaming Language Models with Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92bad072-0b99-4894-a0fc-03834dc80dde · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a6c8530-3350-42f1-9823-34a7cd62de4b · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3375a036-5715-4335-8b84-4de6ba3d4ec9 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 604df649-8fcc-46de-9a5e-abf9a300014c · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2672420b-f5dc-4646-9501-7e3f02d6788e · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ad23a658-881c-4669-861b-e69216d9075b · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 350f2fc8-3cbd-4df7-9d0c-666506826f6b · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Jailbroken: How Does LLM Safety Training Fail?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9c0d9b-d5ca-4b42-a2cb-82e98b14aead · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b45a4906-2582-42b8-955e-5885222aea9a · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Intention Analysis Makes LLMs A Good Jailbreak Defender
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7cae675-c0a7-422b-902b-ceca414e8fef · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 60888ade-9167-4b87-9409-905624aa94f1 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms In Proceedings of the 7th linguistic annotation workshop and interoperability with discourse
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b139c57a-92d3-4950-82f9-dcb1782ef004 · outbound
GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.