Pith. sign in

Paper Citation Record · LEDGER

Latent Adversarial Training Improves the Representation of Refusal

As of 20 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2504.18872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18872 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:11:37.884734Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8e19bdc-4eaf-4e62-94f0-8525ee4d0b8e · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Latent Adversarial Training Improves the Representation of Refusal Refusal in Language Models Is Mediated by a Single Direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.822570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.822570Z digest=sha256:0b431f1a63fae291874bac47caff1dc9e1d52f0f82b69a185ba01b54e3d63dc2

Observation 1f07c3d3-e847-4657-82db-9a88eb1c6366 · outbound

This paper cites latent\_adversarial\_training, 2024.

Latent Adversarial Training Improves the Representation of Refusal latent\_adversarial\_training, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:11:38.085267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T10:11:37.827332Z digest=sha256:71256ffd33c302a62abc1b3f00d70261b57877b9ece48a877957629c438fd6bd

Observation 99f24ba7-0d12-4217-ba35-2dca7db00136 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Latent Adversarial Training Improves the Representation of Refusal Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.831020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.831020Z digest=sha256:4a29ea53cd185d238c6052f79d702eb1c0efef5ed929b49d54f118f9cdc1e186

Observation 9cfd8403-661a-4daa-9039-b58b17411291 · outbound

This paper cites Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks.

Latent Adversarial Training Improves the Representation of Refusal Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.835041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.835041Z digest=sha256:b5aa222f98b4e4cb7109f6b5d40a6adfcc562dcdc59a83d3dedd18aa57c83c18

Observation d50944d6-633b-422e-a924-cfd51cd52358 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Latent Adversarial Training Improves the Representation of Refusal LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.839207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.839207Z digest=sha256:b79195e3191520561152384f43509ca5f03f3ec2b88491409e5d0ff8bb1af2b6

Observation 81d0c4c4-fdd1-4eec-be1b-9523cb1459d7 · outbound

This paper cites Llama 2 7B Chat , 2023.

Latent Adversarial Training Improves the Representation of Refusal Llama 2 7B Chat , 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:11:38.072677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T10:11:37.842845Z digest=sha256:74f6eb32177b9106a4787dac87d5b4f74f89063a0931ad63e6fc5a7863ec7a28

Observation 09070545-108f-4b8a-81a9-b6b6dd6e5941 · outbound

This paper cites The Llama 3 Herd of Models.

Latent Adversarial Training Improves the Representation of Refusal The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.846329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.846329Z digest=sha256:d433c826617037820c7668fe7a397e89b6e1e69670187927dd7850aea6d96fc1

Observation 6902d673-ae34-418d-ad02-370b873474d0 · outbound

This paper cites GPT-4 Technical Report.

Latent Adversarial Training Improves the Representation of Refusal GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.849545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.849545Z digest=sha256:92d3d00634b9917f479ff798a8c18b9989d5a39e1ec7654535cfe2a5591a316f

Observation 660e8bfb-2a4d-465c-87f0-5d2169bb1d0d · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Latent Adversarial Training Improves the Representation of Refusal Steering Llama 2 via Contrastive Activation Addition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.853271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.853271Z digest=sha256:8760d5a0273e0da9414cb3fcb2de7eea384444b6fbaeb991c03c475f84a2dc62

Observation 1af2426f-642e-4353-96d0-a7c20bfc3c0f · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Latent Adversarial Training Improves the Representation of Refusal Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.856549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.856549Z digest=sha256:e9e3917db2732e45a5911cf6eb64321d439a625f9ac3b04e6a259af7957528c1

Observation 43855377-9f47-48f7-8d99-cf4afd16afe2 · outbound

This paper cites Stanford Alpaca: An Instruction-following LLaMA model , 2023.

Latent Adversarial Training Improves the Representation of Refusal Stanford Alpaca: An Instruction-following LLaMA model , 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:11:38.062124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T10:11:37.859571Z digest=sha256:76aaaa98dbc391f7bba7f6981854885a0b3f1c62689704bccd5304352b9edb3f

Observation f7721851-abe9-4320-a3f0-961e43d5e154 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Latent Adversarial Training Improves the Representation of Refusal Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.862434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.862434Z digest=sha256:4d198fc8f92a044bbbfcf3b93830d59dd89a3c79c147ab7f92b7162b8d40b76f

Observation 61ff8dc4-fde5-4294-86d1-12f67372e810 · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Latent Adversarial Training Improves the Representation of Refusal Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.865408Z digest=sha256:136cf8574dd9f1ef522ff70f5f449f7d698b6cbabae2f283776885b72ac3e917

Observation 41fd8d41-096d-4924-9ef0-80841f02c37b · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Latent Adversarial Training Improves the Representation of Refusal Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.868113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.868113Z digest=sha256:97bb2a2bf18fc010a0d55ab1f255cd4a76a6d78bb681b8728ed4be7e43f18b45

Observation 00d2d191-ac13-4e64-8277-4ae63aa08ec0 · outbound

This paper cites write newline.

Latent Adversarial Training Improves the Representation of Refusal write newline

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.872694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.872694Z digest=sha256:0472d6111fa331d63ae8c27c0020c6ff3f21ff98967efb2cecc7950a5e0f0594

Observation 2cb46f6b-b291-49cd-99c9-16cbbae715ca · outbound

This paper cites @esa (Ref.

Latent Adversarial Training Improves the Representation of Refusal @esa (Ref

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.877140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.877140Z digest=sha256:98ea0c1326c4c7d905b97777909b111758f762e87a5d44daca3215600e7bf464

Observation 13d0334d-6335-4d5b-9187-743c72c641c1 · outbound

This paper cites an unresolved cited work.

Latent Adversarial Training Improves the Representation of Refusal Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.880913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.880913Z digest=sha256:0ad4c82aa2a9948c989c778f47ba84993b11859569b8e1ed31c2d3737ec37cae

Observation dcde0731-1e3b-476f-8bdb-d49d5ea73e4f · outbound

This paper cites an unresolved cited work.

Latent Adversarial Training Improves the Representation of Refusal Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:11:37.884734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:11:37.884734Z digest=sha256:63769a9fdeb150944d68413799969289885636f09d28d3f64e0c07c86dfb7949

Pith citing papers

No inbound Pith citation observations are available.