Pith. sign in

Zhihao Xu, Ruixuan Huang, Changyu Chen, and Xiting Wang

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 12 2025 1

roles

background 1

polarities

background 1

representative citing papers

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

cs.CL · 2026-07-06 · conditional · novelty 6.0

A causal audit via neuron-row zeroing shows attribution methods (LRP, IG) faithfully identify dispensable neurons and can install refusal behavior, while rank-stability proxies systematically miss selector failures.

Safety Targeted Embedding Exploit via Refinement

cs.AI · 2026-07-02 · unverdicted · novelty 6.0

STEER is a gradient-guided attack that iteratively translates refusal-triggering words into low-resource languages to jailbreak LLMs, reaching 93-96.7% success on open models and 35.5% transfer to GPT-4o-mini.

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

cs.LG · 2025-12-07 · unverdicted · novelty 6.0

GSAE improves selective refusal on safety benchmarks by smoothing SAE directions over a co-activation graph and applying them via a two-gate controller, outperforming standard SAEs and baselines on Llama-3 and other models.

Semantic Structure of Feature Space in Large Language Models

cs.CL · 2026-04-29 · unverdicted · novelty 5.0

LLM hidden states encode semantic features whose geometric relations, including axis projections, cosine similarities, low-dimensional subspaces, and steering spillovers, closely mirror human psychological associations.

citing papers explorer

Showing 13 of 13 citing papers.