Pith. sign in

Paper Citation Record · LEDGER

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

As of 5 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2607.00572.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.00572 v3

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T09:27:01.450708Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved70
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 349eceff-ca19-405d-aa13-ffafb802bfca · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:c401b4fb583a73674b95cd299ea2c1fbb02f4419eda073fdf9ef6726240f05d5

Observation 83399b44-1fe2-4e2b-a775-713ea8e74fb3 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:a9b19fe17a19b2910cc0af353ddaeb2f6abf8611ea36cd20b87f4a1a95625a03

Observation a7c02e87-485a-499b-827d-2102c6ba037d · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:52b44d8172b3aa363309528cdd225cc961f50d205d8c60cf63d6e0cbdf520411

Observation a86ca43f-89bb-4fdd-99ab-12c410d5826e · outbound

This paper cites do anything now.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment do anything now

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:500205ac57d28ee92e017051c833cda9c86ce9cbabd3f495350ceb6a1bf9f7c7

Observation d52be77e-5c12-4aa2-97b3-005456893548 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreaking black box large language models in twenty queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5330f37fd0a9858cc58ddaa00eb6c59239145cb97947ce0eda47e1f53fc9a0c2

Observation a9c008a9-f87b-4c91-8d0a-f0075c9c134d · outbound

This paper cites Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:cd0afeabc88dedf74cb958574c27c1347a2709f137078720af419040b7d6e793

Observation 48c0c71f-5a2f-41df-a541-20d2fb310765 · outbound

This paper cites Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:cbf629bd1334aca160c9dd372d71f74243ebda0093595ec56924886f3c1ce69d

Observation fa9d951a-6564-41a9-8275-953571c67020 · outbound

This paper cites Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:7e42eecab46f573b5969ed8f143e3af81b7e1b49c7e2a90fe7dce510be6bb2fc

Observation 27503855-02e2-46ec-97d3-38ae4b05ea29 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:08b99f3dc1e9e7ec423916e6d1e302188b6eba671d9793cc11268bc804101671

Observation 665009a6-5061-498d-a52f-7cd49a0425e8 · outbound

This paper cites Codeat- tack: Revealing safety generalization challenges of large language models via code completion.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Codeat- tack: Revealing safety generalization challenges of large language models via code completion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:3c1af8b0ab10451e341882046d6e74cd3fc7fd7d5303defd83421db4423aa33c

Observation ea688beb-b2cf-4c88-aa9b-b9d278dcf0f6 · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:7884f41c5ea286fb29231c1ade3a0d252ecb9db73aa1541a7bee80c85223934c

Observation fae2c09a-a602-46be-bd55-397322efeb11 · outbound

This paper cites Jailbreaking leading safety-aligned llms with simple adaptive attacks.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreaking leading safety-aligned llms with simple adaptive attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:8fd8080349ae84a809fe813fdcd501fc883c0fc4e9668f667c2352f4c7b4b327

Observation fcbf4ab3-a4f7-4c57-bc73-1e290657cd6d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:f9ec57e7e3cb1fff8c7e55d9ff251314c0a0be7aeaf36cbf6fb26dcbbb2ec771

Observation 4e9cd261-1492-4741-89cd-51dca90a7c5c · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Safety alignment should be made more than just a few tokens deep

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:7f8af86a13f66461f8bfaffff0c149cab95d5648b06b4d4d82ab65f27562cb31

Observation 9b2a054d-3b7f-4ee8-bccf-a3d47cd92f47 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5a02788bb1b4a9de7f60ee80500b4964dece5483cbdeea3189e617aa18b55fd4

Observation f97bc49c-dcea-4f65-b036-863b21157e77 · outbound

This paper cites Stair: Improving safety alignment with introspective reasoning.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Stair: Improving safety alignment with introspective reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:48cf77c7e3c69c26564e7db908435ffd3a3ef26420c82f8dc6ea7c04324b6f82

Observation 1898b460-1e07-49ad-ab28-1f57da9c091d · outbound

This paper cites Programming refusal with conditional activation steering.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Programming refusal with conditional activation steering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:3e766e34583336b89ba23538754c6b98a0ccd75df544f61e7659ee13521fa20f

Observation 26dcd498-52c9-4db3-9679-28eed5c8992a · outbound

This paper cites Scans: Mitigating the exaggerated safety for llms via safety-conscious activation steering.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Scans: Mitigating the exaggerated safety for llms via safety-conscious activation steering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:8e01a9bf0b4650cc86afe7e9342a8d24741663c7f4eaa7e8a111f1f573c592ac

Observation e245bf66-3f2e-4b02-be2f-fe1f6ccfedad · outbound

This paper cites Improving alignment and robustness with circuit breakers.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Improving alignment and robustness with circuit breakers.Adv

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:f87209c465d0a9c3a84c2e5486c2c4d47be42e458e914518139fb49184e2ae44

Observation da09daaf-fbf6-4b19-ad51-06eacd8e58b4 · outbound

This paper cites Representation bending for large language model safety.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Representation bending for large language model safety

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:cb1e064a9e4d8ef2f8794dd3bbdd66a1de6ad4a3af52e613decdf349ee36aa92

Observation 69947b70-976f-4ec3-a387-7733c7d15841 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:9785d7ebe8ff363ca372582c7d748a15c792b71c47211e9bce62a5a9b6171d8f

Observation 6606df10-e0a1-48af-8020-804716d7ad8c · outbound

This paper cites On effects of steering latent representation for large language model unlearning.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment On effects of steering latent representation for large language model unlearning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:27ad86ad21a3210a33a8a19a815823d54fd1ec049934b4529b9eb6f960bf294e

Observation 078b703f-4dfe-407c-bbe7-825e0c504931 · outbound

This paper cites Refusal in language models is mediated by a single direction.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Refusal in language models is mediated by a single direction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5a2cfc47cfac68fae9cfeeb5fda18c78a49a45e12f233dcdb92d59716f72fcea

Observation e793413e-b877-4bfc-9e75-3620df25a06c · outbound

This paper cites Llms encode harmful- ness and refusal separately.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Llms encode harmful- ness and refusal separately

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5866dac64f61baf4ac80cf21136393dd14d2fa341b02d6ad24847cb4f747dcbb

Observation 36b37914-aac6-40b3-aa41-472a20e869dc · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment A General Language Assistant as a Laboratory for Alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:0cc8244d2bddcc215a968a911a46e03695cf9acc0e8dda611f530364686e540c

Observation 65e3f1af-fcd8-4bca-8f13-4967d973802a · outbound

This paper cites Training language models to follow instructions with human feedback.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training language models to follow instructions with human feedback.Adv

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:2292bd7c233772e9947d5a6f55e7d45eb6234e56fecb410c12b178bd4519952f

Observation b422ab98-86f6-430e-9a1e-cf1a63729379 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:14f6125c9dd3553ac9d4c8c4ee4fb1041a21788555e1c6123fd099541d4d9f17

Observation 23f19ad7-6ff3-4f1a-800f-6944d36c0655 · outbound

This paper cites The linear representation hypothesis and the geometry of large language models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The linear representation hypothesis and the geometry of large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:f401ca015c92d26477d86143bb1940fa80b4351356f82d0d9c001b20838cfb84

Observation 8a28ed0d-bcd1-4e43-bd1d-4c3acac9bcd5 · outbound

This paper cites Toy Models of Superposition.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Toy Models of Superposition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:2531fa428a2a3feff13e01c0df4900ce51a27a1b41928893a5038156229b400b

Observation a5c35159-4181-46cc-8221-aa26a3fc1e6e · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:c256f4feef57acdb31732d98875f70cf8cbc7efc54c1328de6e5edc967d6bf6e

Observation 988a7bfb-2af5-4d0b-bc63-b444476738aa · outbound

This paper cites Steering Language Models With Activation Engineering.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Steering Language Models With Activation Engineering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:0b617c3b73d2d1d68e81c641a39884b54e6ba77fee3517e40bfb7beda75f205a

Observation 2eb021f2-0ce7-49cf-9dcc-8bec0cc05b89 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Steering llama 2 via contrastive activation addition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:d9e5918e531f223f0f870ab83d8a5196148e7bde840cb2d3d9d0b4a13e801159

Observation 1e2e8620-f2eb-4eb7-9026-0d533a6933fb · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:937d8714e72b3f04201b712c4abfeaaa7e0c326e1a65232215c00ca6d379e4c1

Observation 75ba58ae-c9bb-4bad-b3b7-487290664f63 · outbound

This paper cites Diff-in-means concept editing is worst-case optimal, 2023.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Diff-in-means concept editing is worst-case optimal, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:04ba3daf7f25d1a1345f87a6f8c8f64587d88f4e8f7776a12cb8ddeda83acda0

Observation de6ecb4a-7d27-4aaf-8d1f-08588c88b3b9 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Inference- time intervention: Eliciting truthful answers from a language model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:b682d24a0ccba247fca95d003e0e3e9e969e4befc6d2920bc59c496f4aef9808

Observation b698df7a-9a0e-435c-9eef-6052af0a0baf · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Linear Representations of Sentiment in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:49b9d0e04acd21d7a86cea44d35a5dfcd853ab7e59071264a98e8019a21e2ffd

Observation 6125dd17-9811-4507-8f75-bbc056b69bf1 · outbound

This paper cites Improving instruction-following in language models through activation steering.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Improving instruction-following in language models through activation steering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:e7a0805988c592004d2c1e74b278698d425c7a5cf8b5fb207940e13c778382ab

Observation c21c4caa-636d-4332-a5a8-6ecce864516d · outbound

This paper cites The Llama 3 Herd of Models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The Llama 3 Herd of Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:c7e82eb3f9754850f4895e62e522fb7bfe9cc4e523f992cc6bf175853dca5329

Observation 64110347-6aae-4c8a-b8f4-7b8c2a8a66f5 · outbound

This paper cites Qwen2.5 Technical Report.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:bcfa60cbf6d89e3fdcfaf1943889b98c105d456367e31e81ab0c8da19ee27d99

Observation 7289d67b-a329-457a-893c-5d740e857a7b · outbound

This paper cites Openai usage policies, 2025.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Openai usage policies, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:d698e60e8c17926f7ca55e60121f9ee3b02dd13f9c9367f4402090e7ec140baa

Observation e191348a-b2f8-4bd7-bdb3-0a45f027740a · outbound

This paper cites Exponential moving average of weights in deep learning: Dynamics and benefits.Transactions on Machine Learning Research Journal, pages 1–27, 2024.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Exponential moving average of weights in deep learning: Dynamics and benefits.Transactions on Machine Learning Research Journal, pages 1–27, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:acf0499aa0ddc8407eef4e1461376a8acd1bf3612c962686c6b4dd6744d01b36

Observation 76b4bd4e-5dbd-4b83-9d5f-248055b13a35 · outbound

This paper cites Rethinking safety in llm fine-tuning: An optimization perspective.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Rethinking safety in llm fine-tuning: An optimization perspective

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:483369a5b70b9493b378b0ffcf8ab7407f75215f444fb3ac8ee45815661d0de2

Observation 083bb1af-6660-4453-9fa6-0bbf76522683 · outbound

This paper cites Pku-saferlhf: Towards multi-level safety alignment for llms with human preference.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Pku-saferlhf: Towards multi-level safety alignment for llms with human preference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:b43b25eb74b0046670fbe08ecf92c4c304b8073513bf23d46ac567505798f0b0

Observation 7a30b095-c8e2-4179-b04e-af8f77ee6e08 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In Int.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Fine-tuning aligned language models compromises safety, even when users do not intend to! In Int

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:4b2c6c69ffb29741c8c3eced5098ad5ca1233eb53e362fb8f429f4842058e3fe

Observation 298806c9-ec88-44af-8b06-47b6f0d355ed · outbound

This paper cites Jailbreakbench: An open robustness benchmark for jailbreaking large language models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Jailbreakbench: An open robustness benchmark for jailbreaking large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:a874fd6c1a148b7736d4215d39b05fc315a17d7abc2960aa84dcb07e4c7f781d

Observation 0feb2e48-1960-44fa-9fad-a01c7c9339c7 · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:cc7e1e388ceffb9073bd874a73d5567b4cdac72c2952aa9a483363e11b4ac114

Observation 8801fbff-ae13-430a-bc7a-d0c566d08dc7 · outbound

This paper cites The art of saying no: Contextual noncompliance in language models.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The art of saying no: Contextual noncompliance in language models.Adv

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5348b059222f7c4ea0efe166b98125cf3db7ce00712902e986a51d98d9132d54

Observation 04cb9ba6-c668-4086-b1b6-3c0247e3698d · outbound

This paper cites Measuring massive multitask language understanding.Int.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Measuring massive multitask language understanding.Int

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:4d7ff4fca7c8c4d203131028f3366c94d989ed649ec767d62c4741ccb6d87b66

Observation 7237b0be-5ba8-4f9a-ad04-8726142e9cdd · outbound

This paper cites Aligning ai with shared human values.Int.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Aligning ai with shared human values.Int

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:0282f66392d504d144a9fb068b9331fc735b59f8377b37d2a259a7ddee2a6dbd

Observation f100205f-c54d-45f2-abf2-dac03049f1c9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:b586ee0d549582f044fbcf24dedbbd2d8e26c1a1a452033b4268456ce7d2b540

Observation e1ffa12d-b3e4-454e-b6e2-0275d27b11e5 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:ef416aff9104528bbb8e0be01e337ce1ea3b94070862961059a3bd2553b0fa81

Observation 105ac1ae-c33a-4b8f-af3a-49a46010d14f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Evaluating Large Language Models Trained on Code

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:357e5d6fbd805cada04b56b0063d37e5576367a1414062f5d8ec59678c145bd9

Observation 7bd5f164-e71b-4a1f-8e1a-5c9d263a7c05 · outbound

This paper cites Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:12c8532daa3dbdab4f7236460e4767d36699a18a61b5bcc3a9361f87b869bf1e

Observation 8279766b-24ed-4974-8501-084773785cde · outbound

This paper cites A survey on llm-as-a-judge.The Innovation, 2024.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment A survey on llm-as-a-judge.The Innovation, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:658a373f3bde49764704475f59be70fc42ba3ab603428285dd141f334bccf40f

Observation 4baf2020-c1ff-44de-a9a4-125e91f9dda5 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Enhancing chat language models by scaling high-quality instructional conversations

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:9ece8235cb3b261813cbed83d8a1fa47a41554664d995d52b00d24cd0ba9fcd2

Observation 35aaa368-611b-4f4d-b352-875ea2acc8c4 · outbound

This paper cites The geometry of refusal in large language models: Concept cones and representational independence.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The geometry of refusal in large language models: Concept cones and representational independence

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:5a13f09c3812f954aad04554ceb7ccb39f27242c8f687715159e42fe2956b274

Observation 49bf8d41-7f30-4c56-9767-48ea2811a35d · outbound

This paper cites The hidden dimensions of llm alignment: A multi-dimensional analysis of orthogonal safety directions.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment The hidden dimensions of llm alignment: A multi-dimensional analysis of orthogonal safety directions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:cd95cadf6100c652249616cb17f8fe4813434b24a304f04b77b0cef050b960ec

Observation 880abd99-68a2-464a-8c35-fa7ee0f4eac0 · outbound

This paper cites Differentiated directional intervention: A framework for evading llm safety alignment.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Differentiated directional intervention: A framework for evading llm safety alignment

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:a31a077cbf919aa897f5f1461cf8dc02a7ee2ffcc040512984c8ed608f93101e

Observation 7ccba616-4e61-476c-91b7-fff80981294a · outbound

This paper cites Alphasteer: Learning refusal steering with principled null-space constraint.arXiv preprint arXiv:2506.07022, 2025.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Alphasteer: Learning refusal steering with principled null-space constraint.arXiv preprint arXiv:2506.07022, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:d27c9459bd584462ca60b97a26d08e4be76b1aef2df61bddd59c6f3db90dee1c

Observation f63c3331-8162-426c-b48a-a1485a18b73f · outbound

This paper cites Analysing the generalisation and reliability of steering vectors.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Analysing the generalisation and reliability of steering vectors.Adv

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:55d703f28053cbcc2d170a5472c90a02e0458789dee9fc52788c7e3251b4539e

Observation ca3310b1-0b6d-451e-abe6-b2f66f4e3822 · outbound

This paper cites Or-bench: An over-refusal benchmark for large language models.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Or-bench: An over-refusal benchmark for large language models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:906704cb14af25a2d4cb16b8f326d14a413a5f4d8f3abe8a7f17a4fdb018ac24

Observation 1df60ea9-085c-4ab9-ac56-e71a3fed136b · outbound

This paper cites Universal jailbreak suffixes are strong attention hijackers.arXiv preprint arXiv:2506.12880, 2025.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Universal jailbreak suffixes are strong attention hijackers.arXiv preprint arXiv:2506.12880, 2025

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:65a59687e252931fc67706e49f11611ca8bd6931fa640c8b50911d844c7f3fc7

Observation 59596b0f-d842-424f-b74b-4dcfe998b6f0 · outbound

This paper cites Flipattack: Jailbreak llms via flipping.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Flipattack: Jailbreak llms via flipping

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:d43b411093999a1171ea9ef4997968ea8d3c0a4a625d53df837975daf0bceda3

Observation 0f580215-e422-4885-814f-b6f6f14c63da · outbound

This paper cites Many-shot jailbreaking.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Many-shot jailbreaking.Adv

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:f65fcb8733595e75fefd40e86788c402b0b8d366369b2611995d163d9cfc90d3

Observation 458ee6af-9556-4d03-adde-58a99e8babef · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:829450100d6753ec504dce58faa19b9f6859f3921ccd3ae7680cde401624125a

Observation 8a05a620-e575-4fd3-ade9-1f907754593e · outbound

This paper cites Mistral 7B.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Mistral 7B

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:9c4a00b102d8839e6feaa380866550e44c472b986d4d9ab23f01519f5ceeefa5

Observation 61529521-92bb-4218-9808-e8fc093a1ff6 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:ad90bc0a0e139b547de27e3bae568c3a34875b144d9dd92206c8717f9c0bdf12

Observation 5fb6a24c-dd3c-430d-b446-7173d7548168 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Gemma 2: Improving Open Language Models at a Practical Size

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:2c227a57e98ce146ce54087821d8aefe4373658ed56c2c8da0c9bdcf4cea3423

Observation 81f9127c-41b8-409a-860e-f2d783815fdb · outbound

This paper cites Hashimoto.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Hashimoto

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:c76f35693d74e4c905096ce19a70074304f5e5a9e95b99ee8a9dd22f0a70f35e

Observation 449b96a0-5286-48ee-b5e9-afdf06f6be8c · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Adv.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Adv

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:abd86726e5fbaa185e8653ea6ba9d92420150d5daa97044402590d879ac63b37

Observation 14ec44a6-e41e-41b4-a96a-f41ba9dac89b · outbound

This paper cites What are some good books on Roman history?.

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment What are some good books on Roman history?

Reference 71

Resolution
malformed identifier
no resolver link, observed 2026-07-12T09:27:01.450708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:27:01.450708Z digest=sha256:0e8c6bcdc3a1b42d16deaf8f7cdd234700aabe58626dea77f2c30bd92b4894bf

Pith citing papers

No inbound Pith citation observations are available.