{"id":"a85ff413-e57a-41e9-9bf3-aef5d45309cb","arxiv_id":"2505.20645","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"STEER-BENCH is a Reddit-derived benchmark of 5,552 multiple-choice questions on which the best of 13 large language models scores near 65 percent, versus human experts near 81 percent.","lead":"This paper introduces STEER-BENCH, a benchmark that tests whether large language models can adapt their answers to match contrasting online communities built from Reddit forums. Across 13 models, even the best-scoring chatbots reach about 65 percent agreement with the benchmark's community labels, where human experts reach around 81 percent.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 81% human baseline appears to be Cohen's Kappa 0.815 measured on roughly 30 items; without raw accuracy and error bars, the central human-vs-model gap is not established.","rationale":"The reader's weakest assumption already targets GPT-4o silver-label validity and the small human validation sample; my stress-test sharpens this into a concrete, load-bearing defect in the headline comparison. The benchmark's construction and model ranking are still informative, so I would not reject the paper. However, the abstract's 'human experts achieve 81% accuracy' claim appears to be supported only by a Cohen's Kappa coefficient on roughly 30 sampled questions, which is not an accuracy figure and lacks error bars. Since all model accuracies are computed against the same silver labels, a properly documented human baseline is essential for the central 'models lag humans by 15 points' conclusion. A raw-agreement recomputation from the released annotation logs would settle whether the 81% figure is real, and a larger community-member validation would test generalizability. Thus the reader's CONDITIONAL verdict stands, with no change needed.","tokens_in":30112,"tokens_out":4953,"duration_ms":54910,"concrete_test":"Download the released human-annotation logs and recompute the raw human-golden vs GPT-4o-silver percentage agreement per validated item, with a 95% bootstrap confidence interval restricted to the exact items used in Section 4.6. If the raw percentage is not approximately 81%, or if the confidence interval overlaps the best model's 65% accuracy on the same items, the headline human-vs-model gap is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims 'human experts achieve an accuracy of 81% with silver labels,' but Section 4.6 reports only 'the inter-rater agreement between the golden labels and silver labels (GPT-4o generated answers) is 0.815 measured by Cohen's Kappa.' Cohen's Kappa is chance-corrected and is not the same as raw accuracy; the paper never reports the raw human-vs-silver percentage. The validation design samples one multiple-choice question per subreddit pair, so the human baseline rests on at most 30 questions after filtering, while all 5,552 evaluation questions are scored against GPT-4o silver labels. Annotators are described as 'familiar with Reddit culture,' not as members of the target communities, so even the reported agreement measures generic Reddit-literate readers' agreement with GPT-4o, not community-specific expertise. The central 'humans 81% vs best model 65%' claim therefore depends on a baseline that is either mislabeled as accuracy or measured on a tiny sample with no confidence interval, making the 15-point gap unquantified. The Limitations section honestly acknowledges GPT-4o supervision bias, but that admission does not repair the missing human-baseline documentation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces STEER-BENCH, a benchmark for measuring how well LLMs can be steered toward the perspectives of specific online communities. It uses 30 contrasting Reddit subreddit pairs across 19 domains, with GPT-4o used to generate open-ended instruction-response demonstrations and multiple-choice questions with 'silver' labels. The authors evaluate 13 LLMs under in-context learning and supervised fine-tuning, reporting that the best models reach roughly 65% accuracy while human experts allegedly achieve 81% accuracy with silver labels, leaving a 15-point gap. The paper also documents scaling trends, domain-level variation, and differences among model families.","tokens_in":30375,"tokens_out":4098,"duration_ms":45354,"significance":"If its validity holds, STEER-BENCH would be a useful and reusable resource for a genuinely under-explored capability: aligning LLM outputs with community-specific norms and worldviews. The paper has tangible strengths: a relatively large constructed dataset (8,328 paired demonstrations, 5,552 multiple-choice questions), broad model coverage across open and proprietary families, two steering paradigms, public code/data, and a candid Limitations section that explicitly acknowledges GPT-4o supervision bias and binary community framing. These strengths make the benchmark potentially valuable to the community. However, the validity of the headline quantitative claims currently rests on a small, partly misreported human-validation study and on silver labels that are generated and scored by the same model family being evaluated.","major_comments":[{"comment":"The abstract's claim that 'human experts achieve an accuracy of 81% with silver labels' is not supported by the reported analysis. Section 4.6 reports Cohen's Kappa of 0.815 between golden and silver labels; Kappa is a chance-corrected agreement coefficient, not raw accuracy, and the paper never reports the raw percentage of human answers matching GPT-4o's labels. Moreover, the human validation design samples only one multiple-choice question per subreddit pair, so after filtering for annotator familiarity the human baseline rests on at most 30 questions, with no confidence interval. This directly undermines the headline 'humans 81% vs. best model 65%' gap. Please report raw agreement with sample size and confidence intervals, or substantially temper the abstract and conclusion claims.","section":"Abstract and Section 4.6"},{"comment":"All 5,552 evaluation questions are scored against GPT-4o-generated silver labels, while human validation is limited to a small sample of annotators described only as 'familiar with Reddit culture,' not as members of the target communities. As the Limitations section honestly acknowledges, scores therefore partly measure agreement with GPT-4o's interpretation of each community, and models that resemble GPT-4o may be systematically favored. This is not by itself disqualifying, but the paper needs to provide more substantive independent grounding: for example, per-domain or per-pair human-silver agreement rates, raw human accuracy against silver labels, and an analysis of label reliability across ideologically sensitive domains. Without this, the benchmark's validity as a measure of community alignment is not established.","section":"Sections 4.4 and 5.3"},{"comment":"The 'In-topic Few-shot' configuration, which is the main evaluation setting, may suffer from topic-level leakage. The open-ended demonstrations and the multiple-choice questions are generated from the same sampled comments and the same topic keywords. A model can therefore answer a test question by matching surface content or answer patterns in the demonstrations rather than by generalizing a community's perspective to new instances. This could inflate Config 4 and Config 5 scores relative to what 'steerability' should mean. Please either hold out topics or comments during demonstration construction, or provide an analysis showing that test questions are not answerable by simple demonstration matching.","section":"Sections 4.3, 4.4, and 5.2"},{"comment":"The paper makes strong comparative claims—for example, the '53-point gap' between Qwen2.5-32B and Llama-3.2-3B on Abortion, and monotonic within-family scaling trends—without any uncertainty quantification. Per-domain sample sizes are small (e.g., the Abortion domain has only 32 questions in Table 10), so point estimates in Tables 14 and 15 will have wide confidence intervals. Please report bootstrap confidence intervals or standard errors for the main accuracy figures, and ideally for the domain-level tables, before drawing conclusions about which domains are 'easy' or which models dominate in specific areas.","section":"Tables 11, 14, 15 and Section 5.4.1"}],"minor_comments":[{"comment":"The phrase '5,500 multiple-choice question' should be '5,500 multiple-choice questions.'","section":"Abstract"},{"comment":"The related-work citation 'Hendrycks et al.' is incomplete; a year and venue are needed.","section":"Section 2"},{"comment":"The explanation that DeepSeek-v3's uniformly low performance is 'possibly due to its predominantly non-English pretraining' is speculative and not supported by evidence presented in the paper; please either provide supporting analysis or remove the conjecture.","section":"Section 5.4"},{"comment":"The BERTopic step would benefit from explicit hyperparameter settings (e.g., embedding model, minimum topic size), since topic quality directly affects all downstream generation.","section":"Section 4.2"},{"comment":"The human validation section would be clearer if it stated explicitly how many sections survived the familiarity filter, how many multiple-choice questions remained after filtering, and how the 'soft voting' with confidence scores was operationalized.","section":"Section 4.6"}],"recommendation":"major_revision","confidential_remarks":"The abstract's '81% accuracy' appears to conflate Cohen's Kappa with raw accuracy; this is the kind of error that should be corrected prominently. The topic-leakage concern in the in-topic few-shot setting is, in my view, the most substantive validity issue, and it may require new experiments or dataset splits. I would encourage the editor to request a revised version that either provides the missing human-baseline documentation and leakage controls or substantially narrows the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark itself is a solid, reusable resource. The headline claim is not. The abstract says human experts hit 81% accuracy with the silver labels, but Section 4.6 reports a Cohen's Kappa of 0.815 between human golden labels and GPT-4o silver labels. Those are different numbers. Kappa corrects for chance agreement, and raw agreement here is likely higher, but the paper never reports it. The human validation also samples one multiple-choice question per subreddit pair, so after filtering for familiarity you are looking at maybe 30 questions, with no confidence interval. The 'models lag humans by 15 points' conclusion is therefore not established as stated.\n\nWhat is genuinely new: a carefully constructed benchmark for community-specific steerability—30 contrasting subreddit pairs across 19 domains, 8,328 paired instruction-response demonstrations, and 5,552 multiple-choice questions with silver labels. The pipeline builds on Community-Cross-Instruct and COMPO, but the packaging, the scale, and the systematic evaluation of 13 models under both in-context learning and SFT make it a useful addition. The authors do several things right: they anonymize subreddit names during generation to reduce reliance on prior knowledge, they use topic modeling to find shared discussion themes, they discard ambiguous answer options, and they validate a sample with four annotators. The Limitations section is honest about the GPT-4o supervision bias and the binary framing.\n\nThe soft spots beyond the baseline issue: the silver labels come from the same model (GPT-4o) that provides the steering demonstrations, so scores partly measure agreement with GPT-4o's reading of each community. Human validation provides some independent grounding, but the sample is too small to fully offset that. There are also no error bars or significance tests on the model comparisons; some domain-level differences rest on very few questions (Abortion has 32, Music 48). The annotators are 'familiar with Reddit culture,' not necessarily community members, which is a lesser but real limitation.\n\nWho gets value: model developers choosing between steering methods, alignment researchers looking for a community-norm evaluation, and social NLP folks. It deserves a serious referee. The right outcome is probably major revision: fix the human baseline reporting, add raw accuracy and confidence intervals, and re-state the human gap cautiously. The dataset itself is worth having regardless.","headline":"Useful, reusable benchmark, but the headline human-vs-model gap rests on a mislabeled Kappa statistic and a roughly 30-item human sample.","tokens_in":30874,"tokens_out":2889,"would_cite":true,"duration_ms":28309,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces a benchmark for measuring whether LLMs can be steered to adopt a community's viewpoint, built from 30 contrasting subreddit pairs and 5,552 multiple-choice questions; it reports that the best of 13 tested models…","keywords":["steerability","LLM evaluation","community alignment","silver labels","in-context learning","supervised finetuning","Reddit communities"],"falsifier":"A concrete check: take a random sample of the 5,552 multiple-choice questions, have a larger and more diverse panel of human annotators who are deeply familiar with each community produce their own answer labels, then rescore the 13 models against those human labels. If agreement with the GPT-4o silver labels falls well below the reported 0.815, or if rescoring erases the human-over-model gap or changes which model is best, the benchmark's central conclusion would not survive.","tokens_in":29925,"feed_emoji":"🎯","tokens_out":12396,"duration_ms":118881,"temperature":0.7,"pith_summary":"The paper is trying to establish that community-level steerability—whether a language model can adjust its outputs to match a specific group's norms and viewpoints—can be measured in a structured way, and that current LLMs are far from human-level at it. It builds STEER-BENCH from 30 contrasting subreddit pairs across 19 domains, generating over 10,000 instruction-response demonstrations and 5,552 multiple-choice questions whose GPT-4o-generated answers serve as silver labels for what each community would say. Testing 13 models under both in-context learning and supervised fine-tuning, the paper reports that human experts match the silver labels 81% of the time while the best model reaches only about 65%, with some models more than 15 percentage points behind. It also finds that on-topic demonstrations help more than naming the community, that steerability improves with model size within a family, and that fine-tuning on this data does not beat prompting. If the benchmark is valid, it provides a reusable way to track whether models are getting better at respecting diverse cultural and ideological perspectives before they are deployed in community-facing settings.","feed_headline":"LLMs lag humans by 15 points on community viewpoints","feed_subtitle":"New benchmark scores 13 models on community-aligned questions: best hits 65%, human experts 81%.","key_machinery":"The central object is the contrasting community pair: two subreddits that discuss shared topics from different perspectives, such as r/Parenting versus r/Childfree or r/linux versus r/windows. The construction pipeline first uses topic modeling on comments from both communities to find topics both sides actually discuss, then prompts an advanced LLM (the paper uses GPT-4o) to write open-ended instruction-response pairs and multiple-choice questions with two community-specific answers. The evaluation mechanism is simple accuracy against the silver label—the GPT-4o-generated answer treated as ground truth for that community—after a model has been steered by demonstrations, either in the prompt or through fine-tuning. Because each question has two contrasting correct answers, the benchmark can tell whether a model is actually adopting the target community's viewpoint rather than producing a generic answer, and the paper's five prompt configurations separate the effect of community context from the model's pretrained priors.","core_discovery":"The central discovery is a quantitative gap: when a model is steered toward a community by examples of that community's own answers, no tested model agrees with community-aligned labels as often as humans do. Human experts reach 81% agreement with the GPT-4o-generated silver labels, the best LLM reaches about 65%, and weaker or less culturally aligned models fall to roughly 30–40%. The paper argues the gap is not just a prompt problem: adding in-topic few-shot demonstrations and subreddit identifiers raises most models' scores but does not close the gap, and models differ sharply by family and by domain, with ideologically sensitive domains such as abortion and politics showing the largest spreads between strong and weak models. On the paper's terms, these results show that steerability is a real, separable capability that scales within model families but is not yet close to human-level community alignment.","pith_inferences":["The paper leaves implicit that its accuracy numbers are upper bounds on true community alignment, because a model that expresses a community's view in wording different from GPT-4o's would be counted wrong; rescoring with multi-generator or full human labels could produce different rankings.","The binary community design probably makes the task easier than real life, since real communities contain internal disagreements; a model that captures one faction's perspective is penalized unless it matches the single silver label, and a continuous or multi-community version could reveal a different failure profile.","The paired demonstrations are directly usable as preference data—each instruction comes with a target-community answer and a contrastive-community answer—so the benchmark's data could support preference-optimization training, not just evaluation.","If the reported gap persists after label improvements, it would imply that community alignment is set mostly by pretraining data and alignment choices rather than by prompting, which would redirect improvement efforts toward data composition."],"forward_implications":["Without explicit on-topic demonstrations, current LLMs will often fail to reflect the perspective of a specific community even when asked to do so.","Within a model family, larger models are the safer choice for community-sensitive deployment, because steerability improves steadily with scale.","Adding in-topic example answers and a subreddit identifier is a stronger and cheaper steering lever than relying on the model's pretrained knowledge or a bare community name.","Fine-tuning on a few hundred community demonstrations does not yet beat in-context learning, so today's practical steering gains will mostly come from prompting rather than weight updates.","The benchmark supplies a reusable protocol: any future model can be scored on the same 5,552 questions to see whether the human-model gap narrows."],"supporting_citations":[{"why":"provides the topic-modeling method used to identify shared topics between each contrasting community pair.","marker":"Grootendorst, 2022"},{"why":"supplies the community-cross-instruction approach the paper adapts to generate instruction-response pairs and questions from social media comments.","marker":"He et al., 2024b"},{"why":"contributes the COMPO framework that uses contrasting subreddit norms for personalization, which motivates the paired-community design.","marker":"Kumar et al., 2025"},{"why":"documents LLM susceptibility to ideological steering, the behavior STEER-BENCH is designed to measure.","marker":"Chen et al., 2024a"},{"why":"earlier result that prompting models toward demographic viewpoints gives modest gains, which the paper's larger human-model gap extends.","marker":"Santurkar et al., 2023"},{"why":"supports using Reddit communities as a source of contrasting opinions and rhetorical styles.","marker":"Chen et al., 2023"}],"fun_headline_variants":["LLMs trail humans by 15 points on community-aligned answers","Steer-Bench: best LLM hits 65% vs human 81% on community norms","Community alignment gap: humans beat all 13 LLMs by 15+ points","New benchmark shows LLMs fall short on community-specific steering","Steerability test: LLMs hit 65%, humans 81% on community norms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the GPT-4o-generated silver labels faithfully capture each community's perspective, because every model's accuracy is scored against those labels; the human check used only four annotators on sampled topics, and the paper itself flags this risk in its Limitations section.","fun_headline_variants_meta":{"raw":{"variants":["LLMs trail humans by 15 points on community-aligned answers","Steer-Bench: best LLM hits 65% vs human 81% on community norms","Community alignment gap: humans beat all 13 LLMs by 15+ points","New benchmark shows LLMs fall short on community-specific steering","Steerability test: LLMs hit 65%, humans 81% on community norms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000278,"raw_usage":{"total_tokens":1641,"prompt_tokens":918,"completion_tokens":723,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":619}},"tokens_in":534,"tokens_out":723,"duration_ms":6555,"temperature":1.0,"reasoning_tokens":619,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:49:47.946538+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: take a random sample of the 5,552 multiple-choice questions, have a larger and more diverse panel of human annotators who are deeply familiar with each community produce their own answer labels, then rescore the 13 models against those human labels. If agreement with the GPT-4o silver labels falls well below the reported 0.815, or if rescoring erases the human-over-model gap or changes which model is best, the benchmark's central conclusion would not survive.","supporting_citations":[{"cited_title":"Smith, and Hannaneh Hajishirzi","cited_arxiv_id":null,"evidence_quote":"contributes the COMPO framework that uses contrasting subreddit norms for personalization, which motivates the paired-community design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supports using Reddit communities as a source of contrasting opinions and rhetorical styles."}],"review_version":1}