Pith. sign in

Paper Citation Record · LEDGER

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning

As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2606.16349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.16349 v3

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T13:51:23.302296Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:49:26.207330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T17:48:45.643373Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 517d5d7e-7b84-4ede-8b16-100604b4c251 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:e235f177e28135652b541443fa94e1af7919342478705ac1789461e39af09d3d

Observation 234f4c65-ad39-4d23-8f51-4d524fa98b6e · outbound

This paper cites A StrongREJECT for empty jailbreaks.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning A StrongREJECT for empty jailbreaks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:3f832927d8143cfb6717cab9a29fe46a441b8457355dc7bda79f7b7f0b6a80a4

Observation 6a38767c-91b1-4561-8dd9-908e50a71823 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:0f4e928b7373c554aefe05d95f205fb830f5ec5cdafa4b2b5e7d643e56448907

Observation 527a698d-9823-42b4-b34a-ec034aa2c4a9 · outbound

This paper cites Refusal in language models is mediated by a single direction.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Refusal in language models is mediated by a single direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:db164baf0bd49eb7a053ca02f7bad75a66bf800c317b3108e73a683e29583554

Observation be8d6c9d-66d7-4b68-97db-4c7a5048161f · outbound

This paper cites COSMIC: Generalized refusal direction identification in LLM activations.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning COSMIC: Generalized refusal direction identification in LLM activations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:6e5cd7f20b17120a2610219cd6b7e0e42badb93bed5f7c9a0d1d8b1a3b7f2a99

Observation 93059131-1573-4f46-a876-7574380d2a42 · outbound

This paper cites The geometry of refusal in large language models: Concept cones and representational independence.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning The geometry of refusal in large language models: Concept cones and representational independence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:9c90d8e573f1c8d2cc2cb2935c1884a6e564b94d490c52ea16933eb5029cc8af

Observation 636a9cab-1665-43fb-8dc6-08fbbfa9c357 · outbound

This paper cites LLMs Encode Harmfulness and Refusal Separately.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning LLMs Encode Harmfulness and Refusal Separately

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:b81557c49fd2b86162c44d4d9c285e6c2ffa89e64a264531b679dd5a90f574e5

Observation ae90c5f3-a7c3-4594-99da-f9279fa17980 · outbound

This paper cites Mistral 7B.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:79e6f73545a3377cf774307bb513cf793b7f0c94fb5f7184ed21f0f67b7d486f

Observation 40fa5604-5615-4480-9470-40ae5d26cfdc · outbound

This paper cites The Llama 3 Herd of Models.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:28b646d9b4085ee2334e52c851f69add0827341b97d28a14411243f0b70c4303

Observation 1e1eecbd-f3c3-42cb-afcf-30d735cf7307 · outbound

This paper cites Qwen2.5 Technical Report.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Qwen2.5 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:7637fed629f528a5a8833be95be14e179353d14f42571ad6f680153caab36d4f

Observation 98964eaa-527d-488b-af0a-de80dd4ca404 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:942cb70a17d84983a6ca0f1281fab9a563524dade42208975044f9d9e29224b3

Observation b8461ac4-f350-46b6-98da-e48874fea09f · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:269dcf0b1b22d0e3baa50622c9db348a0e3ad1f34f9c136df5ea3f432d070e80

Observation cdcb5004-6bae-45ef-abc4-c524fdf095c2 · outbound

This paper cites Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:1f53f51d30eb9db81ca9efa2e6f57cdb825244397256b790fd4f976a631425bb

Observation 8d8693e8-7b69-4405-aedc-5acd64c57409 · outbound

This paper cites Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:854a6d8ba2d61850a15e49f59cae37529c399694d897fe7fda55af42919e869d

Observation 8e0ad023-5950-4a2d-9d13-8aaed034aadb · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:03a9bc61c7adb98cf199a9bc01238bb346681951019c23b50754b110c28e58d5

Observation 8cab7028-1ce8-4184-851a-58c7fd8e3341 · outbound

This paper cites FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:38635ee14a41a39078300fbe12e870ab468a0f260c48c819d9b19097d4dba7bc

Observation 8799d1f8-ce76-45b6-b232-40e028ce072e · outbound

This paper cites an unresolved cited work.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:ff8493255fd903f4ae8fc5801361e3fc3e81504863710d57e60b6f06e786bb1a

Observation 6b59134e-298a-4831-a814-00e79c71256f · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In��� ������� ������������� ���������� �� �������� ���������������, 2024.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In��� ������� ������������� ���������� �� �������� ���������������, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:8d32d6e64e7261d54555a075c851a4679a13677b60e9459e93a80dfbab6f3122

Observation e10e4143-9f50-4b2e-ab0d-b8dd5087b1b4 · outbound

This paper cites Robust LLM safeguarding via refusal feature adversarial training.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Robust LLM safeguarding via refusal feature adversarial training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:5c5dbe58e01608835a03a8fea4f4e25893f19f5e1a6b8c29ca455b29c1724052

Observation e78c5e00-dff5-44fe-a5e2-8bc5177ceb23 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:db48dd4e8aedf8620373d4064a803ea03ff4bcba19caa4984fcb8428994d9a47

Observation de712837-19de-409e-9cab-cc3856103e18 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:51ec394944b89b160b5f31e6621c4e0466435566d012de96a8b7c2739be45b8f

Observation 8bddd9a4-b930-4113-9a99-5134a473fb5c · outbound

This paper cites Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:d4871c9ca4ecbc54a7aba4b76c97b7fe6b9269d3c58ac17ae296f270f8c7b441

Observation f23cf5d4-accd-4444-9ee9-b58ff2c50178 · outbound

This paper cites an unresolved cited work.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:cd37c48bd52e4d62e5647238c0b8b64c2bab1f148f70eac83d0205fc8d6ef5f1

Observation b02d906c-690a-4dc7-9b00-247d2fd9d436 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:114e5683ce24f37763665ff8b7573e1a7574f8eb32cea1157412ef904ed60f12

Observation 149c929f-a39f-4043-9ba9-933d92b01a9b · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T13:51:23.302296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:51:23.302296Z digest=sha256:13d7151e158ddbac995523d9d34ebb097dde1c41f98798bb964948de4022f54e

Pith citing papers

Observation 14ccfa8d-37c3-4791-9edb-ab9fd7655f5b · inbound

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning cites this paper.

From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:48:45.644591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T03:49:26.207330Z digest=sha256:50c46ab5d563a1fa535b88dc824a50ab5fd4823ab6b6e2a3597e1f2811dd777c