Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:11:37.884734Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2504.18872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:11:37.884734Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e8e19bdc-4eaf-4e62-94f0-8525ee4d0b8e · outbound
Latent Adversarial Training Improves the Representation of Refusal Refusal in Language Models Is Mediated by a Single Direction
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f07c3d3-e847-4657-82db-9a88eb1c6366 · outbound
Latent Adversarial Training Improves the Representation of Refusal latent\_adversarial\_training, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 99f24ba7-0d12-4217-ba35-2dca7db00136 · outbound
Latent Adversarial Training Improves the Representation of Refusal Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfd8403-661a-4daa-9039-b58b17411291 · outbound
Latent Adversarial Training Improves the Representation of Refusal Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d50944d6-633b-422e-a924-cfd51cd52358 · outbound
Latent Adversarial Training Improves the Representation of Refusal LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81d0c4c4-fdd1-4eec-be1b-9523cb1459d7 · outbound
Latent Adversarial Training Improves the Representation of Refusal Llama 2 7B Chat , 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 09070545-108f-4b8a-81a9-b6b6dd6e5941 · outbound
Latent Adversarial Training Improves the Representation of Refusal The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6902d673-ae34-418d-ad02-370b873474d0 · outbound
Latent Adversarial Training Improves the Representation of Refusal GPT-4 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660e8bfb-2a4d-465c-87f0-5d2169bb1d0d · outbound
Latent Adversarial Training Improves the Representation of Refusal Steering Llama 2 via Contrastive Activation Addition
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af2426f-642e-4353-96d0-a7c20bfc3c0f · outbound
Latent Adversarial Training Improves the Representation of Refusal Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43855377-9f47-48f7-8d99-cf4afd16afe2 · outbound
Latent Adversarial Training Improves the Representation of Refusal Stanford Alpaca: An Instruction-following LLaMA model , 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f7721851-abe9-4320-a3f0-961e43d5e154 · outbound
Latent Adversarial Training Improves the Representation of Refusal Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ff8dc4-fde5-4294-86d1-12f67372e810 · outbound
Latent Adversarial Training Improves the Representation of Refusal Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41fd8d41-096d-4924-9ef0-80841f02c37b · outbound
Latent Adversarial Training Improves the Representation of Refusal Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d2d191-ac13-4e64-8277-4ae63aa08ec0 · outbound
Latent Adversarial Training Improves the Representation of Refusal write newline
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb46f6b-b291-49cd-99c9-16cbbae715ca · outbound
Latent Adversarial Training Improves the Representation of Refusal @esa (Ref
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d0334d-6335-4d5b-9187-743c72c641c1 · outbound
Latent Adversarial Training Improves the Representation of Refusal Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcde0731-1e3b-476f-8bdb-d49d5ea73e4f · outbound
Latent Adversarial Training Improves the Representation of Refusal Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.