Pith. sign in

REVIEW 2 cited by

Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.15576 v1 pith:4QUWBVDT submitted 2025-02-21 cs.CL

classification cs.CL
keywords explanationsbehaviorsfeaturessemanticsteeringunderstandingautoencodersbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) excel at handling human queries, but they can occasionally generate flawed or unexpected responses. Understanding their internal states is crucial for understanding their successes, diagnosing their failures, and refining their capabilities. Although sparse autoencoders (SAEs) have shown promise for interpreting LLM internal representations, limited research has explored how to better explain SAE features, i.e., understanding the semantic meaning of features learned by SAE. Our theoretical analysis reveals that existing explanation methods suffer from the frequency bias issue, where they emphasize linguistic patterns over semantic concepts, while the latter is more critical to steer LLM behaviors. To address this, we propose using a fixed vocabulary set for feature interpretations and designing a mutual information-based objective, aiming to better capture the semantic meaning behind these features. We further propose two runtime steering strategies that adjust the learned feature activations based on their corresponding explanations. Empirical results show that, compared to baselines, our method provides more discourse-level explanations and effectively steers LLM behaviors to defend against jailbreak attacks. These findings highlight the value of explanations for steering LLM behaviors in downstream applications. We will release our code and data once accepted.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale

    cs.LG 2025-09 conditional novelty 6.0 of 10

    An instruction-tuned LLaMA 3.1 8B model, trained on LLM-generated pseudo-labels, is claimed to identify IoT device vendors from passive network metadata with 98.25% top-1 accuracy across 2,015 vendors.

  2. EmoPerso: Enhancing Personality Detection with Self-Supervised Emotion-Aware Modelling

    cs.CL 2025-09 conditional novelty 5.0 of 10

    EmoPerso improves MBTI personality detection by training an emotion head on heuristic pseudo-labels and using cross-attention with reasoning chains, achieving 81.07% Macro-F1 on Kaggle and 68.60% on Pandora.

Pith tools