Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.556402Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2608.03201.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:46:20.556402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 27e49704-fa53-4cbf-9e3d-f27f51fd9122 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Re- fusal in language models is mediated by a single direction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 412f3972-a423-4bca-8c02-7a3d42bcc147 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14b17d5d-295f-462a-b722-b17c870d7169 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8777735c-fe77-4430-ba0a-b3851f49c04e · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2c492514-75ff-45d7-a126-263604525c1f · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Shortcut learning in deep neural networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab9baa83-2016-4e87-9384-e13ce2f75c63 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Gemma 2: Improving open language models at a practical size, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61869a1f-7c2e-4d51-ba11-463f949eb22e · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d55693ed-b70e-40a7-a494-f4adad971c6c · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc014932-bf8c-480b-89bc-a0f37ece256b · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Advances in Neural In- formation Processing Systems, 37:8093–8131, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aaea2452-5997-4ffc-be65-ad824ecae504 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30482cf7-dbe6-488d-9c4c-05bc9328d8ee · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Beavertails: Towards improved safety align- ment of llm via a human-preference dataset.Advances in Neural Information Processing Systems, 36:24678–24704,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0fa892b2-1f9e-4fc2-8e9f-9e5eb123a348 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety of large language models beyond english: A systematic lit- erature review of risks, biases, and safeguards
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57e47eed-913e-447a-ab1e-d610b18c1643 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Inference-time intervention: Elicit- ing truthful answers from a language model.Advances in Neural Information Processing Systems, 36:41451–41530,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f655b4d-ed6d-4f9f-b895-4ed3971bae1c · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Guardreasoner: Towards reasoning-based llm safeguards.arXiv preprint arXiv:2501.18492, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f420115-b5f4-46b7-acac-d7e6cbdda070 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ede1b0-9fc5-4d6c-ade2-a56f036c4ef1 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Training language models to follow instructions with human feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2f17bd-9433-4bb6-979b-c2c5f86332d2 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety alignment should be made more than just a few tokens deep
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f23becf8-9b10-4974-8a8a-27d6907a09ec · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed392039-ccda-4b92-8987-fb612f6c028a · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf6e92b-ca1e-4607-9b5a-fdb65d7cd5d1 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safety through reasoning: An empirical study of reasoning guardrail models.Findings of the Associa- tion for Computational Linguistics: EMNLP, 2025:21862– 21880, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ba4260b-cefc-4b30-a406-b7a6b34aa2e9 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Shortcut learning in safety: The impact of keyword bias in safeguards
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7dd512e3-c88b-434c-9b0d-04c7b66fd737 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Jail- broken: How does llm safety training fail?Advances in Neu- ral Information Processing Systems, 36:80079–80110, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea325f40-8252-45be-a68b-2bf9cc12070c · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Qwen3 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251d382c-e487-4edc-a2af-3c1549ea1d5f · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Safeseek: Universal attribution of safety circuits in language models.arXiv preprint arXiv:2603.23268, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 077f18ff-e8be-4268-a724-7044051ea787 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc6d994-3eb0-43ba-83ba-345a5a84029c · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Refuse whenever you feel unsafe: Improving safety in llms via decoupled refusal training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83c5b1b7-a17d-416c-9347-9ab31cd1660e · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820d62e9-055a-458d-b2e8-4c39047555f2 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models STAIR: Improving Safety Alignment with Introspective Reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc178fed-7253-42ef-866b-51cd6a4120b8 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Qwen3Guard Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ddfc659-70e3-4213-b109-be4fcefa8cb1 · outbound
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models Llms encode harmfulness and refusal sepa- rately.Advances in Neural Information Processing Systems, 38:140283–140318, 2026
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.