Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
Advances in neural information processing systems , volume=
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
Gender bias and factual gender knowledge are severely entangled in language model circuits and neurons, making neuron ablation an unreliable method for debiasing.
High-probability generalization bounds for D-SGD are derived at the optimal rate O(1/sqrt(mn) log(1/δ)) via pointwise uniform stability across convex and non-convex settings.
The paper characterizes deductive stereotyping in LLMs and introduces Fair-GCG to discover injection phrases that improve fairness across benchmarks, reasoning, and real-world tasks.
Pre-training provides a geometric warm start in a single-index model that enables weak-to-strong generalization up to a supervisor-limited bound, with empirical phase-transition evidence in LLMs.
Relative Probability Association Metric (RPAM) measures LM associations via softmax-normalized continuation probabilities and correlates strongly with human associations and downstream LM behavior across three models.
citing papers explorer
-
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
-
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
Gender bias and factual gender knowledge are severely entangled in language model circuits and neurons, making neuron ablation an unreliable method for debiasing.
-
Unveiling High-Probability Generalization in Decentralized SGD
High-probability generalization bounds for D-SGD are derived at the optimal rate O(1/sqrt(mn) log(1/δ)) via pointwise uniform stability across convex and non-convex settings.
-
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
The paper characterizes deductive stereotyping in LLMs and introduces Fair-GCG to discover injection phrases that improve fairness across benchmarks, reasoning, and real-world tasks.
-
On the Blessing of Pre-training in Weak-to-Strong Generalization
Pre-training provides a geometric warm start in a single-index model that enables weak-to-strong generalization up to a supervisor-limited bound, with empirical phase-transition evidence in LLMs.
-
RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
Relative Probability Association Metric (RPAM) measures LM associations via softmax-normalized continuation probabilities and correlates strongly with human associations and downstream LM behavior across three models.