Pith. sign in

REVIEW 3 cited by

Refusal Behavior in Large Language Models: A Nonlinear Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08145 v1 pith:PD2T6QDG submitted 2025-01-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords refusalbehaviornonlinearalignmentlanguagelargellmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Refusal behavior in large language models (LLMs) enables them to decline responding to harmful, unethical, or inappropriate prompts, ensuring alignment with ethical standards. This paper investigates refusal behavior across six LLMs from three architectural families. We challenge the assumption of refusal as a linear phenomenon by employing dimensionality reduction techniques, including PCA, t-SNE, and UMAP. Our results reveal that refusal mechanisms exhibit nonlinear, multidimensional characteristics that vary by model architecture and layer. These findings highlight the need for nonlinear interpretability to improve alignment research and inform safer AI deployment strategies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Truth judgments in LLMs are causally steerable along multiple independent directions forming a cone, not just one axis, across Qwen and Gemma families.

  2. The Geometry of Harmfulness in LLMs through Subconcept Probing

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Fifty-five harmfulness subconcept directions in Llama-3.1-8B-Instruct form a nearly rank-1 subspace, and steering along the dominant direction cuts jailbreak success but costs accuracy and fails on Qwen.

  3. From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law

    cs.CY 2025-06 conditional novelty 4.0 of 10

    Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.

Pith tools