Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:27:01.450708Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2607.00572.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-12T09:27:01.450708Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 349eceff-ca19-405d-aa13-ffafb802bfca · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83399b44-1fe2-4e2b-a775-713ea8e74fb3 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Autodan: Generating stealthy jailbreak prompts on aligned large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c02e87-485a-499b-827d-2102c6ba037d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a86ca43f-89bb-4fdd-99ab-12c410d5826e · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment do anything now
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52be77e-5c12-4aa2-97b3-005456893548 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreaking black box large language models in twenty queries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c008a9-f87b-4c91-8d0a-f0075c9c134d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c0c71f-5a2f-41df-a541-20d2fb310765 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa9d951a-6564-41a9-8275-953571c67020 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27503855-02e2-46ec-97d3-38ae4b05ea29 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665009a6-5061-498d-a52f-7cd49a0425e8 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Codeat- tack: Revealing safety generalization challenges of large language models via code completion
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea688beb-b2cf-4c88-aa9b-b9d278dcf0f6 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Artprompt: Ascii art-based jailbreak attacks against aligned llms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae2c09a-a602-46be-bd55-397322efeb11 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreaking leading safety-aligned llms with simple adaptive attacks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcbf4ab3-a4f7-4c57-bc73-1e290657cd6d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9cd261-1492-4741-89cd-51dca90a7c5c · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Safety alignment should be made more than just a few tokens deep
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f97bc49c-dcea-4f65-b036-863b21157e77 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Stair: Improving safety alignment with introspective reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1898b460-1e07-49ad-ab28-1f57da9c091d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Programming refusal with conditional activation steering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26dcd498-52c9-4db3-9679-28eed5c8992a · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Scans: Mitigating the exaggerated safety for llms via safety-conscious activation steering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e245bf66-3f2e-4b02-be2f-fe1f6ccfedad · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Improving alignment and robustness with circuit breakers.Adv
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da09daaf-fbf6-4b19-ad51-06eacd8e58b4 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Representation bending for large language model safety
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69947b70-976f-4ec3-a387-7733c7d15841 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6606df10-e0a1-48af-8020-804716d7ad8c · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment On effects of steering latent representation for large language model unlearning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078b703f-4dfe-407c-bbe7-825e0c504931 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Refusal in language models is mediated by a single direction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e793413e-b877-4bfc-9e75-3620df25a06c · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Llms encode harmful- ness and refusal separately
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b37914-aac6-40b3-aa41-472a20e869dc · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment A General Language Assistant as a Laboratory for Alignment
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e3f1af-fcd8-4bca-8f13-4967d973802a · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training language models to follow instructions with human feedback.Adv
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b422ab98-86f6-430e-9a1e-cf1a63729379 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23f19ad7-6ff3-4f1a-800f-6944d36c0655 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The linear representation hypothesis and the geometry of large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a28ed0d-bcd1-4e43-bd1d-4c3acac9bcd5 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Toy Models of Superposition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c35159-4181-46cc-8221-aa26a3fc1e6e · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988a7bfb-2af5-4d0b-bc63-b444476738aa · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Steering Language Models With Activation Engineering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eb021f2-0ce7-49cf-9dcc-8bec0cc05b89 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Steering llama 2 via contrastive activation addition
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2e8620-f2eb-4eb7-9026-0d533a6933fb · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Representation Engineering: A Top-Down Approach to AI Transparency
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ba58ae-c9bb-4bad-b3b7-487290664f63 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Diff-in-means concept editing is worst-case optimal, 2023
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6ecb4a-7d27-4aaf-8d1f-08588c88b3b9 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Inference- time intervention: Eliciting truthful answers from a language model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b698df7a-9a0e-435c-9eef-6052af0a0baf · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Linear Representations of Sentiment in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6125dd17-9811-4507-8f75-bbc056b69bf1 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Improving instruction-following in language models through activation steering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21c4caa-636d-4332-a5a8-6ecce864516d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The Llama 3 Herd of Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64110347-6aae-4c8a-b8f4-7b8c2a8a66f5 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Qwen2.5 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7289d67b-a329-457a-893c-5d740e857a7b · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Openai usage policies, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e191348a-b2f8-4bd7-bdb3-0a45f027740a · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Exponential moving average of weights in deep learning: Dynamics and benefits.Transactions on Machine Learning Research Journal, pages 1–27, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76b4bd4e-5dbd-4b83-9d5f-248055b13a35 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Rethinking safety in llm fine-tuning: An optimization perspective
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 083bb1af-6660-4453-9fa6-0bbf76522683 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Pku-saferlhf: Towards multi-level safety alignment for llms with human preference
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a30b095-c8e2-4179-b04e-af8f77ee6e08 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Fine-tuning aligned language models compromises safety, even when users do not intend to! In Int
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298806c9-ec88-44af-8b06-47b6f0d355ed · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0feb2e48-1960-44fa-9fad-a01c7c9339c7 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8801fbff-ae13-430a-bc7a-d0c566d08dc7 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The art of saying no: Contextual noncompliance in language models.Adv
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04cb9ba6-c668-4086-b1b6-3c0247e3698d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Measuring massive multitask language understanding.Int
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7237b0be-5ba8-4f9a-ad04-8726142e9cdd · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Aligning ai with shared human values.Int
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f100205f-c54d-45f2-abf2-dac03049f1c9 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training Verifiers to Solve Math Word Problems
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1ffa12d-b3e4-454e-b6e2-0275d27b11e5 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Instruction-Following Evaluation for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 105ac1ae-c33a-4b8f-af3a-49a46010d14f · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Evaluating Large Language Models Trained on Code
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd5f164-e71b-4a1f-8e1a-5c9d263a7c05 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8279766b-24ed-4974-8501-084773785cde · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment A survey on llm-as-a-judge.The Innovation, 2024
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4baf2020-c1ff-44de-a9a4-125e91f9dda5 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Enhancing chat language models by scaling high-quality instructional conversations
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35aaa368-611b-4f4d-b352-875ea2acc8c4 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The geometry of refusal in large language models: Concept cones and representational independence
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49bf8d41-7f30-4c56-9767-48ea2811a35d · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The hidden dimensions of llm alignment: A multi-dimensional analysis of orthogonal safety directions
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 880abd99-68a2-464a-8c35-fa7ee0f4eac0 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Differentiated directional intervention: A framework for evading llm safety alignment
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ccba616-4e61-476c-91b7-fff80981294a · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Alphasteer: Learning refusal steering with principled null-space constraint.arXiv preprint arXiv:2506.07022, 2025
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63c3331-8162-426c-b48a-a1485a18b73f · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Analysing the generalisation and reliability of steering vectors.Adv
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3310b1-0b6d-451e-abe6-b2f66f4e3822 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Or-bench: An over-refusal benchmark for large language models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1df60ea9-085c-4ab9-ac56-e71a3fed136b · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Universal jailbreak suffixes are strong attention hijackers.arXiv preprint arXiv:2506.12880, 2025
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59596b0f-d842-424f-b74b-4dcfe998b6f0 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Flipattack: Jailbreak llms via flipping
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f580215-e422-4885-814f-b6f6f14c63da · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Many-shot jailbreaking.Adv
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458ee6af-9556-4d03-adde-58a99e8babef · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a05a620-e575-4fd3-ade9-1f907754593e · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Mistral 7B
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61529521-92bb-4218-9808-e8fc093a1ff6 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb6a24c-dd3c-430d-b446-7173d7548168 · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Gemma 2: Improving Open Language Models at a Practical Size
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f9127c-41b8-409a-860e-f2d783815fdb · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Hashimoto
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 449b96a0-5286-48ee-b5e9-afdf06f6be8c · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Adv
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ec44a6-e41e-41b4-a96a-f41ba9dac89b · outbound
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment What are some good books on Roman history?
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.