Pith. sign in

Paper Citation Record · LEDGER

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2608.03201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03201 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.556402Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27e49704-fa53-4cbf-9e3d-f27f51fd9122 · outbound

This paper cites Re- fusal in language models is mediated by a single direction.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Re- fusal in language models is mediated by a single direction

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.078896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.473502Z digest=sha256:6e091b82b7fe53e3c861525741af3cacf1ed6c18bc03433223f270371c597007

Observation 412f3972-a423-4bca-8c02-7a3d42bcc147 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.476922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.476922Z digest=sha256:d9251f70d054fa455aa38a05485d89f4bd206f34ccc68dae7532a0dcf46f822d

Observation 14b17d5d-295f-462a-b722-b17c870d7169 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Evaluating Large Language Models Trained on Code

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.480069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.480069Z digest=sha256:e8a59784c324c9b2cafab6ce5e66568f2ecc700029ee8a1992e7ec38567bbcf9

Observation 8777735c-fe77-4430-ba0a-b3851f49c04e · outbound

This paper cites Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.071508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.483397Z digest=sha256:a8df0062a910573696d5f840fb55ea5cea81e2dcb61d083bdbaa815aa53fe44d

Observation 2c492514-75ff-45d7-a126-263604525c1f · outbound

This paper cites Shortcut learning in deep neural networks.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Shortcut learning in deep neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.064247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.486387Z digest=sha256:7e126af022609c8808d7dfc0503e43974b15d6a6f6c72fb759aa8bb804f4d1a1

Observation ab9baa83-2016-4e87-9384-e13ce2f75c63 · outbound

This paper cites Gemma 2: Improving open language models at a practical size, 2024.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Gemma 2: Improving open language models at a practical size, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.056614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.489325Z digest=sha256:75fb25f54bc0e28a2eaadee950363856c99e6b099ccadf4df92e8c94f6d6a598

Observation 61869a1f-7c2e-4d51-ba11-463f949eb22e · outbound

This paper cites an unresolved cited work.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-05T23:46:21.049400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.492245Z digest=sha256:3b681d638ed8d9c7af5d35c432458624fb269283f55204778d7d2f81beb7d059

Observation d55693ed-b70e-40a7-a494-f4adad971c6c · outbound

This paper cites The Llama 3 Herd of Models.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.495238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.495238Z digest=sha256:b16224521dd69dae3267fea1981812aea8cc9477dddb574761f36bdf7d9c503a

Observation bc014932-bf8c-480b-89bc-a0f37ece256b · outbound

This paper cites Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in Neural In- formation Processing Systems, 37:8093–8131, 2024.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in Neural In- formation Processing Systems, 37:8093–8131, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.042184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.498191Z digest=sha256:1e69e9098ecb2aaaad292822f0b319e4adc98102a8819d09054b9aeb2638beb3

Observation aaea2452-5997-4ffc-be65-ad824ecae504 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.500842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.500842Z digest=sha256:14566f3882257a968eed7e68a033f871ab1243911e3a1f5aa212e92e0892dae9

Observation 30482cf7-dbe6-488d-9c4c-05bc9328d8ee · outbound

This paper cites Beavertails: Towards improved safety align- ment of llm via a human-preference dataset.Advances in Neural Information Processing Systems, 36:24678–24704,.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Beavertails: Towards improved safety align- ment of llm via a human-preference dataset.Advances in Neural Information Processing Systems, 36:24678–24704,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.034218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.503732Z digest=sha256:258b06db717697f9aa1d399d934577bae367fe4d5ae36660d1b3018549dad861

Observation 0fa892b2-1f9e-4fc2-8e9f-9e5eb123a348 · outbound

This paper cites Safety of large language models beyond english: A systematic lit- erature review of risks, biases, and safeguards.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety of large language models beyond english: A systematic lit- erature review of risks, biases, and safeguards

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.026713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.506623Z digest=sha256:2842c4c87d88e84f044be816f6c8bbb4de39af659954feddce2974f8e8e13527

Observation 57e47eed-913e-447a-ab1e-d610b18c1643 · outbound

This paper cites Inference-time intervention: Elicit- ing truthful answers from a language model.Advances in Neural Information Processing Systems, 36:41451–41530,.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Inference-time intervention: Elicit- ing truthful answers from a language model.Advances in Neural Information Processing Systems, 36:41451–41530,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.019182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.509274Z digest=sha256:5962f73a4bcfc060c53029e3a2147b7a46a6737d36610d71fddaba165ff1850b

Observation 5f655b4d-ed6d-4f9f-b895-4ed3971bae1c · outbound

This paper cites Guardreasoner: Towards reasoning-based llm safeguards.arXiv preprint arXiv:2501.18492, 2025.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Guardreasoner: Towards reasoning-based llm safeguards.arXiv preprint arXiv:2501.18492, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.511914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.511914Z digest=sha256:ae80875f1b5613b70e4bf7ba02a3e3f9d7564962328b876ae94be65e304d125d

Observation 1f420115-b5f4-46b7-acac-d7e6cbdda070 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.514492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.514492Z digest=sha256:ce09ef185685e94123d8a6b53026eb27379e2f73a9f377bb3b78b0d30719f8e7

Observation 74ede1b0-9fc5-4d6c-ade2-a56f036c4ef1 · outbound

This paper cites Training language models to follow instructions with human feedback.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Training language models to follow instructions with human feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.517697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.517697Z digest=sha256:397f0fb711bb03cec4b85b3bc0fee402e82be959b4d3507e24189f0b86a99f16

Observation 5a2f17bd-9433-4bb6-979b-c2c5f86332d2 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety alignment should be made more than just a few tokens deep

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.010787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.520653Z digest=sha256:67ea14d2052dc79cb52e320ad4b3cce56b995a13aa3da3c4fca2ae8cee3ed9be

Observation f23becf8-9b10-4974-8a8a-27d6907a09ec · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:21.002494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.523391Z digest=sha256:e7fcf62584987f2b28380eaefeaf4e736abc9d4208e7905d43589caac8cc1c7c

Observation ed392039-ccda-4b92-8987-fb612f6c028a · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.526280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.526280Z digest=sha256:e19dbb1b5aa43ce432de8dfa10b04c6cd23b896e1add47c00021b1d6815bbd48

Observation cdf6e92b-ca1e-4607-9b5a-fdb65d7cd5d1 · outbound

This paper cites Safety through reasoning: An empirical study of reasoning guardrail models.Findings of the Associa- tion for Computational Linguistics: EMNLP, 2025:21862– 21880, 2025.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety through reasoning: An empirical study of reasoning guardrail models.Findings of the Associa- tion for Computational Linguistics: EMNLP, 2025:21862– 21880, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:20.994463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.529099Z digest=sha256:acba4f6f93f7e5667b432fe5ec06f32b8374794bf88ecabddcb67241c552a2f3

Observation 1ba4260b-cefc-4b30-a406-b7a6b34aa2e9 · outbound

This paper cites Shortcut learning in safety: The impact of keyword bias in safeguards.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Shortcut learning in safety: The impact of keyword bias in safeguards

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:20.986490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.531657Z digest=sha256:3c7b3b3ab6387247c69b9ad6dcc01899ca8f53376111708cbd6377bb42433c43

Observation 7dd512e3-c88b-434c-9b0d-04c7b66fd737 · outbound

This paper cites Jail- broken: How does llm safety training fail?Advances in Neu- ral Information Processing Systems, 36:80079–80110, 2023.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Jail- broken: How does llm safety training fail?Advances in Neu- ral Information Processing Systems, 36:80079–80110, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:20.978408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.534343Z digest=sha256:aa5a790e67ec420200906292fde67a28914f8becaec4a366237c1cc28d4707eb

Observation ea325f40-8252-45be-a68b-2bf9cc12070c · outbound

This paper cites Qwen3 Technical Report.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Qwen3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.536933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.536933Z digest=sha256:7add196553790bb2a9e1981ebaa69d27e664314b31b03722b943de5e8208510e

Observation 251d382c-e487-4edc-a2af-3c1549ea1d5f · outbound

This paper cites Safeseek: Universal attribution of safety circuits in language models.arXiv preprint arXiv:2603.23268, 2026.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safeseek: Universal attribution of safety circuits in language models.arXiv preprint arXiv:2603.23268, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.539768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.539768Z digest=sha256:7b310615c3642d5960218983ab4212895f6353ac8df9175f78e948ef87cbf4cb

Observation 077f18ff-e8be-4268-a724-7044051ea787 · outbound

This paper cites S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.542247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.542247Z digest=sha256:6e3332737a2f65541c8c8c0f2875baef0b3ed46151f2ce29dc2707834cb71751

Observation bfc6d994-3eb0-43ba-83ba-345a5a84029c · outbound

This paper cites Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T23:46:20.969869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.545023Z digest=sha256:4af492099e3fb86043b1a97c5928de201c0ce2ca66ab00da3b0fe0121e63113a

Observation 83c5b1b7-a17d-416c-9347-9ab31cd1660e · outbound

This paper cites AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.547519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.547519Z digest=sha256:96ca11b65a3ecbc46dbc3c62e99ff6c831263d55592dda3a699075195b7d34b1

Observation 820d62e9-055a-458d-b2e8-4c39047555f2 · outbound

This paper cites STAIR: Improving Safety Alignment with Introspective Reasoning.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models STAIR: Improving Safety Alignment with Introspective Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.550587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.550587Z digest=sha256:630cea008f96f3eb3becbb4874a4b7e5f50571205efe9e8393c4513a91886e6e

Observation bc178fed-7253-42ef-866b-51cd6a4120b8 · outbound

This paper cites Qwen3Guard Technical Report.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Qwen3Guard Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:20.553544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:20.553544Z digest=sha256:3c0fb13bd47938cd681b8f6122c3c0386909cffcfbcbdb07d4fdcd1e23f260f2

Observation 7ddfc659-70e3-4213-b109-be4fcefa8cb1 · outbound

This paper cites Llms encode harmfulness and refusal sepa- rately.Advances in Neural Information Processing Systems, 38:140283–140318, 2026.

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Llms encode harmfulness and refusal sepa- rately.Advances in Neural Information Processing Systems, 38:140283–140318, 2026

Reference 30

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T23:46:20.703274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T23:46:20.556402Z digest=sha256:61795b796d7ac1d77b7cdd44f717604fdad0f8934d7761412e1ea00284fc69d5

Pith citing papers

No inbound Pith citation observations are available.