Pith. sign in

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We introduce Steer-Bench, a benchmark for assessing population-specific steering using contrasting Reddit communities. Covering 30 contrasting subreddit pairs across 19 domains, Steer-Bench includes over 10,000 instruction-response pairs and validated 5,500 multiple-choice question with corresponding silver labels to test alignment with diverse community norms. Our evaluation of 13 popular LLMs using Steer-Bench reveals that while human experts achieve an accuracy of 81% with silver labels, the best-performing models reach only around 65% accuracy depending on the domain and configuration. Some models lag behind human-level alignment by over 15 percentage points, highlighting significant gaps in community-sensitive steerability. Steer-Bench is a benchmark to systematically assess how effectively LLMs understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent diverse cultural and ideological perspectives.

fields

cs.AI 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Divergent Response Modes in Frontier Language Models Under Steering Pressure

cs.AI · 2026-08-06 · conditional · novelty 7.0

Six frontier language models show categorical, model-specific response modes under steering pressure, with GPT-5 uniquely withholding reasoning while answering, and a linear probe plus activation steering tracing Llama's derail behavior to its residual stream.

citing papers explorer

Showing 1 of 1 citing paper.

  • Divergent Response Modes in Frontier Language Models Under Steering Pressure cs.AI · 2026-08-06 · conditional · none · ref 1 · internal anchor

    Six frontier language models show categorical, model-specific response modes under steering pressure, with GPT-5 uniquely withholding reasoning while answering, and a linear probe plus activation steering tracing Llama's derail behavior to its residual stream.