Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T13:51:23.302296Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2606.16349.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T13:51:23.302296Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:49:26.207330Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T17:48:45.643373Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 517d5d7e-7b84-4ede-8b16-100604b4c251 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning HarmBench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 234f4c65-ad39-4d23-8f51-4d524fa98b6e · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning A StrongREJECT for empty jailbreaks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a38767c-91b1-4561-8dd9-908e50a71823 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 527a698d-9823-42b4-b34a-ec034aa2c4a9 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Refusal in language models is mediated by a single direction
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8d6c9d-66d7-4b68-97db-4c7a5048161f · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning COSMIC: Generalized refusal direction identification in LLM activations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93059131-1573-4f46-a876-7574380d2a42 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning The geometry of refusal in large language models: Concept cones and representational independence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636a9cab-1665-43fb-8dc6-08fbbfa9c357 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning LLMs Encode Harmfulness and Refusal Separately
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae90c5f3-a7c3-4594-99da-f9279fa17980 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Mistral 7B
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fa5604-5615-4480-9470-40ae5d26cfdc · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1eecbd-f3c3-42cb-afcf-30d735cf7307 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Qwen2.5 Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98964eaa-527d-488b-af0a-de80dd4ca404 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8461ac4-f350-46b6-98da-e48874fea09f · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdcb5004-6bae-45ef-abc4-c524fdf095c2 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8693e8-7b69-4405-aedc-5acd64c57409 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e0ad023-5950-4a2d-9d13-8aaed034aadb · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cab7028-1ce8-4184-851a-58c7fd8e3341 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8799d1f8-ce76-45b6-b232-40e028ce072e · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b59134e-298a-4831-a814-00e79c71256f · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In��� ������� ������������� ���������� �� �������� ���������������, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10e4143-9f50-4b2e-ab0d-b8dd5087b1b4 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Robust LLM safeguarding via refusal feature adversarial training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e78c5e00-dff5-44fe-a5e2-8bc5177ceb23 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de712837-19de-409e-9cab-cc3856103e18 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bddd9a4-b930-4113-9a99-5134a473fb5c · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f23cf5d4-accd-4444-9ee9-b58ff2c50178 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02d906c-690a-4dc7-9b00-247d2fd9d436 · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 149c929f-a39f-4043-9ba9-933d92b01a9b · outbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ccfa8d-37c3-4791-9edb-ab9fd7655f5b · inbound
From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning From Refusal Geometry to Safety Geometry: Harmfulness--Refusal Coupling under Dynamic Adversarial Fine-Tuning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.