REVIEW 10 cited by
LEACE: Perfect linear concept erasure in closed form
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called "concept scrubbing," which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Code is available at https://github.com/EleutherAI/concept-erasure.
Forward citations
Cited by 10 Pith papers
-
Emergent Misalignment Recruits a Pre-existing Persona Subspace
Fine-tuning on narrow bad data recruits a low-rank persona subspace already present in a frozen instruction-tuned model; holding that subspace out of activations prevents broad misalignment, and injecting it into the ...
-
Convergent Linear Representations of Emergent Misalignment
A single activation direction extracted from one misaligned fine-tune can ablate emergent misalignment across models trained on different datasets with different LoRA setups.
-
Slowing Learning by Erasing Simple Features
QLEACE removes all quadratically available class information from a representation, reliably slows feedforward networks, but can inject higher-order information that lets stronger architectures learn faster.
-
ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders
With matched-scale sparse autoencoders, HuBERT-ECG best preserves its ECG representation while ECG-JEPA best exposes clinical measurements through single features — a leader split that repeats on MIMIC-IV-ECG.
-
Emergent Latent-State Computation under Stochastic Volatility
Volatility forecasters develop linearly decodable representations of the next hidden log-volatility state; in long cycles this appears immediately after the input projection and ℓ2 normalization.
-
Improved Representation Steering for Language Models
RePS, a reference-free bidirectional preference optimization objective, improves representation steering and suppression for Gemma models, outperforming language-modeling objectives and approaching prompting performance.
-
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
UnLOK-VQA adds rephrase and neighborhood samples to OK-VQA to measure generalization and specificity of multimodal unlearning, and its evaluation shows that hiding answer tokens in hidden states beats other deletion o...
-
Converting MLPs into Polynomials in Closed Form
Closed-form polynomial approximations of MLPs and GLUs under Gaussian mixture inputs reveal a training phase where networks shift from linear to quadratic behavior.
-
The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text
Masking treatment-defining words before learning adjustment representations preserves overlap and reduces bias in text-as-treatment causal inference.
-
Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning
A perspective paper that advocates direct, post hoc interpretability for multi-agent deep reinforcement learning and offers a taxonomy of where those methods might apply.
Discussion (0). Continue with ORCID to comment.