REVIEW 4 major objections 5 minor 46 references
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper introduces a benchmark for measuring whether LLMs can be steered to adopt a community's viewpoint, built from 30 contrasting subreddit pairs and 5,552 multiple-choice questions; it reports that the best of 13 tested models…
desk verdict Useful, reusable benchmark, but the headline human-vs-model gap rests on a mislabeled Kappa statistic and a roughly 30-item human sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the contrasting community pair: two subreddits that discuss shared topics from different perspectives, such as r/Parenting versus r/Childfree or r/linux versus r/windows. The construction pipeline first uses topic modeling on comments from both communities to find topics both sides actually discuss, then prompts an advanced LLM (the paper uses GPT-4o) to write open-ended instruction-response pairs and multiple-choice questions with two community-specific answers. The evaluation mechanism is simple accuracy against the silver label—the GPT-4o-generated answer treated as ground truth for that community—after a model has been steered by demonstrations, either in the prompt or through fine-tuning. Because each question has two contrasting correct answers, the benchmark can tell whether a model is actually adopting the target community's viewpoint rather than producing a generic answer, and the paper's five prompt configurations separate the effect of community context from the model's pretrained priors.
What would settle it
A concrete check: take a random sample of the 5,552 multiple-choice questions, have a larger and more diverse panel of human annotators who are deeply familiar with each community produce their own answer labels, then rescore the 13 models against those human labels. If agreement with the GPT-4o silver labels falls well below the reported 0.815, or if rescoring erases the human-over-model gap or changes which model is best, the benchmark's central conclusion would not survive.
Extended reading notes
Core claim
The central discovery is a quantitative gap: when a model is steered toward a community by examples of that community's own answers, no tested model agrees with community-aligned labels as often as humans do. Human experts reach 81% agreement with the GPT-4o-generated silver labels, the best LLM reaches about 65%, and weaker or less culturally aligned models fall to roughly 30–40%. The paper argues the gap is not just a prompt problem: adding in-topic few-shot demonstrations and subreddit identifiers raises most models' scores but does not close the gap, and models differ sharply by family and by domain, with ideologically sensitive domains such as abortion and politics showing the largest spreads between strong and weak models. On the paper's terms, these results show that steerability is a real, separable capability that scales within model families but is not yet close to human-level community alignment.
Load-bearing premise
The load-bearing premise is that the GPT-4o-generated silver labels faithfully capture each community's perspective, because every model's accuracy is scored against those labels; the human check used only four annotators on sampled topics, and the paper itself flags this risk in its Limitations section.
Editorial extensions
If this is right
- Without explicit on-topic demonstrations, current LLMs will often fail to reflect the perspective of a specific community even when asked to do so.
- Within a model family, larger models are the safer choice for community-sensitive deployment, because steerability improves steadily with scale.
- Adding in-topic example answers and a subreddit identifier is a stronger and cheaper steering lever than relying on the model's pretrained knowledge or a bare community name.
- Fine-tuning on a few hundred community demonstrations does not yet beat in-context learning, so today's practical steering gains will mostly come from prompting rather than weight updates.
- The benchmark supplies a reusable protocol: any future model can be scored on the same 5,552 questions to see whether the human-model gap narrows.
Reading between the lines
- The paper leaves implicit that its accuracy numbers are upper bounds on true community alignment, because a model that expresses a community's view in wording different from GPT-4o's would be counted wrong; rescoring with multi-generator or full human labels could produce different rankings.
- The binary community design probably makes the task easier than real life, since real communities contain internal disagreements; a model that captures one faction's perspective is penalized unless it matches the single silver label, and a continuous or multi-community version could reveal a different failure profile.
- The paired demonstrations are directly usable as preference data—each instruction comes with a target-community answer and a contrastive-community answer—so the benchmark's data could support preference-optimization training, not just evaluation.
- If the reported gap persists after label improvements, it would imply that community alignment is set mostly by pretraining data and alignment choices rather than by prompting, which would redirect improvement efforts toward data composition.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces STEER-BENCH, a benchmark for measuring how well LLMs can be steered toward the perspectives of specific online communities. It uses 30 contrasting Reddit subreddit pairs across 19 domains, with GPT-4o used to generate open-ended instruction-response demonstrations and multiple-choice questions with 'silver' labels. The authors evaluate 13 LLMs under in-context learning and supervised fine-tuning, reporting that the best models reach roughly 65% accuracy while human experts allegedly achieve 81% accuracy with silver labels, leaving a 15-point gap. The paper also documents scaling trends, domain-level variation, and differences among model families.
Significance. If its validity holds, STEER-BENCH would be a useful and reusable resource for a genuinely under-explored capability: aligning LLM outputs with community-specific norms and worldviews. The paper has tangible strengths: a relatively large constructed dataset (8,328 paired demonstrations, 5,552 multiple-choice questions), broad model coverage across open and proprietary families, two steering paradigms, public code/data, and a candid Limitations section that explicitly acknowledges GPT-4o supervision bias and binary community framing. These strengths make the benchmark potentially valuable to the community. However, the validity of the headline quantitative claims currently rests on a small, partly misreported human-validation study and on silver labels that are generated and scored by the same model family being evaluated.
major comments (4)
- [Abstract and Section 4.6] The abstract's claim that 'human experts achieve an accuracy of 81% with silver labels' is not supported by the reported analysis. Section 4.6 reports Cohen's Kappa of 0.815 between golden and silver labels; Kappa is a chance-corrected agreement coefficient, not raw accuracy, and the paper never reports the raw percentage of human answers matching GPT-4o's labels. Moreover, the human validation design samples only one multiple-choice question per subreddit pair, so after filtering for annotator familiarity the human baseline rests on at most 30 questions, with no confidence interval. This directly undermines the headline 'humans 81% vs. best model 65%' gap. Please report raw agreement with sample size and confidence intervals, or substantially temper the abstract and conclusion claims.
- [Sections 4.4 and 5.3] All 5,552 evaluation questions are scored against GPT-4o-generated silver labels, while human validation is limited to a small sample of annotators described only as 'familiar with Reddit culture,' not as members of the target communities. As the Limitations section honestly acknowledges, scores therefore partly measure agreement with GPT-4o's interpretation of each community, and models that resemble GPT-4o may be systematically favored. This is not by itself disqualifying, but the paper needs to provide more substantive independent grounding: for example, per-domain or per-pair human-silver agreement rates, raw human accuracy against silver labels, and an analysis of label reliability across ideologically sensitive domains. Without this, the benchmark's validity as a measure of community alignment is not established.
- [Sections 4.3, 4.4, and 5.2] The 'In-topic Few-shot' configuration, which is the main evaluation setting, may suffer from topic-level leakage. The open-ended demonstrations and the multiple-choice questions are generated from the same sampled comments and the same topic keywords. A model can therefore answer a test question by matching surface content or answer patterns in the demonstrations rather than by generalizing a community's perspective to new instances. This could inflate Config 4 and Config 5 scores relative to what 'steerability' should mean. Please either hold out topics or comments during demonstration construction, or provide an analysis showing that test questions are not answerable by simple demonstration matching.
- [Tables 11, 14, 15 and Section 5.4.1] The paper makes strong comparative claims—for example, the '53-point gap' between Qwen2.5-32B and Llama-3.2-3B on Abortion, and monotonic within-family scaling trends—without any uncertainty quantification. Per-domain sample sizes are small (e.g., the Abortion domain has only 32 questions in Table 10), so point estimates in Tables 14 and 15 will have wide confidence intervals. Please report bootstrap confidence intervals or standard errors for the main accuracy figures, and ideally for the domain-level tables, before drawing conclusions about which domains are 'easy' or which models dominate in specific areas.
minor comments (5)
- [Abstract] The phrase '5,500 multiple-choice question' should be '5,500 multiple-choice questions.'
- [Section 2] The related-work citation 'Hendrycks et al.' is incomplete; a year and venue are needed.
- [Section 5.4] The explanation that DeepSeek-v3's uniformly low performance is 'possibly due to its predominantly non-English pretraining' is speculative and not supported by evidence presented in the paper; please either provide supporting analysis or remove the conjecture.
- [Section 4.2] The BERTopic step would benefit from explicit hyperparameter settings (e.g., embedding model, minimum topic size), since topic quality directly affects all downstream generation.
- [Section 4.6] The human validation section would be clearer if it stated explicitly how many sections survived the familiarity filter, how many multiple-choice questions remained after filtering, and how the 'soft voting' with confidence scores was operationalized.
Circularity Check
Steer-Bench's core accuracy is defined as agreement with GPT-4o's own silver labels, and the abstract's '81% human accuracy' relabels Cohen's Kappa.
-
self definitional
[Section 4.4 (Question-Answer Generation); Section 5.3 (Evaluation Protocol); Section 4.3 (Instruction-Response Generation)]
"GPT-4o is prompted to create such a question q_k and select the correct answer a^A_k and a^B_k, aligned with C_A and C_B respectively. ... We refer to the GPT-4o answer as the silver label. ... We measure the accuracy as the proportion of model responses that match the silver labels."
The target 'community-aligned answer' is not measured against an independent community ground truth: GPT-4o both writes the multiple-choice question and selects the 'correct' answer from the same sampled comments (Section 4.4), and GPT-4o also generates the open-ended demonstrations used to steer the model (Section 4.3). Accuracy is then defined as the fraction of model responses matching GPT-4o's silver label (Section 5.3). Hence the benchmark's central scores measure, by construction, how well a model reproduces GPT-4o's interpretation of the community, rather than an independently verified community norm.
-
other
[Abstract vs. Section 4.6 (Human Validation)]
"human experts achieve an accuracy of 81% with silver labels (Abstract). The inter-rater agreement between the golden labels and silver labels (GPT-4o generated answers) is 0.815 measured by Cohen's Kappa (Section 4.6)."
Cohen's Kappa is a chance-corrected inter-rater agreement statistic, not a raw accuracy percentage. The abstract's 'human experts achieve an accuracy of 81% with silver labels' turns the Section 4.6 Kappa value of 0.815 into a human accuracy baseline, and the paper's headline conclusion that models lag humans by over 15 points depends on comparing model accuracy to this relabeled number. Moreover, the Kappa was computed on at most 30 multiple-choice questions (one per subreddit pair), not on the 5,552 items used for model scoring. The reported data therefore do not establish the claimed human-vs-model accuracy gap; the only independent anchor against the GPT-4o label loop is weakened by a statistic relabeled as accuracy.
full rationale
Steer-Bench's raw material (Reddit comments) is external, and the human annotation study provides a genuinely independent check on a small sample, so the paper is not fully self-referential. However, the operational ground truth for all 5,552 evaluation questions is 'the GPT-4o answer' (Section 4.4), and the steering demonstrations are generated by the same GPT-4o from the same sampled comments (Section 4.3). Accuracy is then defined as agreement with those GPT-4o silver labels (Section 5.3). Thus the central model-accuracy numbers measure self-consistency with GPT-4o's interpretation of the communities, not an independently established community norm; the paper's own Limitations section admits this risk. The abstract's 'human experts achieve an accuracy of 81%' is not a raw accuracy but Cohen's Kappa 0.815 from a roughly 30-item validation, so the headline human-vs-model gap is not established by the reported statistic. I do not see a load-bearing self-citation chain: the cited Community-Cross-Instruct and COMPO are scaffolding for the pipeline, not the justification for correctness. Hence score 6: the core evaluation metric partially reduces by construction, but human validation and the paper's candid limitation statement prevent a score of 8-10.
Assumptions & free parameters
free parameters (5)
- Shared topic minimum comment threshold =
200 per community
- Document subsampling caps =
500,000 per community, 3x cap for larger community
- Number of in-context demonstrations =
12
- Annotator confidence weights for golden labels =
1.0 for Yes, 0.5 for Maybe
- SFT training settings =
2 epochs, batch size 8, LR 8e-6/6e-6
assumptions (5)
- domain assumption Reddit subreddit comments are a faithful proxy for community norms and perspectives.
- domain assumption GPT-4o can generate instructions and answers that faithfully reflect the provided comments.
- domain assumption Human validation on a small sample is sufficient to certify all 5,552 silver labels.
- domain assumption Contrasting subreddit pairs provide a valid binary grounding for community-specific alignment.
- domain assumption Multiple-choice accuracy on a single correct answer measures steerability.
Cite this review
Pith. "Pith review of STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models." pith.science (2026). https://pith.science/paper/5VUSQF7O
@misc{pith2026250520645,
author = {Pith},
title = {Pith review of: STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5VUSQF7O}},
note = {Machine review of arXiv:2505.20645}
}
read the original abstract
Steerability, or the ability of large language models (LLMs) to adapt outputs to align with diverse community-specific norms, perspectives, and communication styles, is critical for real-world applications but remains under-evaluated. We introduce Steer-Bench, a benchmark for assessing population-specific steering using contrasting Reddit communities. Covering 30 contrasting subreddit pairs across 19 domains, Steer-Bench includes over 10,000 instruction-response pairs and validated 5,500 multiple-choice question with corresponding silver labels to test alignment with diverse community norms. Our evaluation of 13 popular LLMs using Steer-Bench reveals that while human experts achieve an accuracy of 81% with silver labels, the best-performing models reach only around 65% accuracy depending on the domain and configuration. Some models lag behind human-level alignment by over 15 percentage points, highlighting significant gaps in community-sensitive steerability. Steer-Bench is a benchmark to systematically assess how effectively LLMs understand community-specific instructions, their resilience to adversarial steering attempts, and their ability to accurately represent diverse cultural and ideological perspectives.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Duarte Alves, Nuno Guerreiro, Jo \ a o Alves, Jos \'e Pombal, Ricardo Rei, Jos \'e de Souza, Pierre Colombo, and Andre Martins. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.744 Steering large language models for machine translation with finetuning and in-context learning . In Findings of the Association for Computational Linguistics: EMNLP 2023, ...
-
[4]
Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki, Sarath Chandar, and Pascal Vincent. 2025. Steering large language model activations in sparse spaces. arXiv preprint arXiv:2503.00177
arXiv 2025
-
[5]
Yuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin, Lu Lin, Fenglong Ma, and Jinghui Chen. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/58cbe393b4254da8966780a40d023c0b-Paper-Conference.pdf Personalized steering of large language models: Versatile steering vectors through bi-directional preference optimization . In Advances in Neural In...
work page 2024
-
[6]
Kai Chen, Zihao He, Rong-Ching Chang, Jonathan May, and Kristina Lerman. 2023. Anger breeds controversy: analyzing controversy and emotions on reddit. In International conference on social computing, behavioral-cultural modeling and prediction and behavior representation in modeling and simulation, pages 44--53. Springer
work page 2023
-
[7]
Kai Chen, Zihao He, Jun Yan, Taiwei Shi, and Kristina Lerman. 2024 a . https://doi.org/10.18653/v1/2024.emnlp-main.952 How susceptible are large language models to ideological manipulation? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17140--17161, Miami, Florida, USA. Association for Computational Linguistics
-
[8]
Nuo Chen, Yan Wang, Yang Deng, and Jia Li. 2024 b . The oscars of ai theater: A survey on role-playing with language models. arXiv preprint arXiv:2407.11484
arXiv 2024
Show all 46 references
-
[9]
Yi Dong, Zhilin Wang, Makesh Sreedhar, Xianchao Wu, and Oleksii Kuchaiev. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.754 S teer LM : Attribute conditioned SFT as an (user-steerable) alternative to RLHF . In Findings of the Association for Computational Linguistics: ...
2023 doi
-
[10]
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for C...
2019
-
[11]
Aleksandra Edwards and Jose Camacho-Collados. 2024. https://aclanthology.org/2024.lrec-main.879/ Language models for text classification: Is in-context learning enough? In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources a...
2024
-
[12]
Hamideh Ghanadian, Isar Nejadgholi, and Hussein Al Osman. 2024. Socially aware synthetic data generation for suicidal ideation detection using large language models. IEEe Access, 12:14350--14363
2024
-
[13]
Maarten Grootendorst. 2022. Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794
2022 arXiv
-
[14]
Ping Guo, Yubing Ren, Yue Hu, Yanan Cao, Yunpeng Li, and Heyan Huang. 2024. https://doi.org/10.1145/3626772.3657819 Steering large language models for cross-lingual information retrieval . In Proceedings of the 47th International ACM SIGIR Conference on Research and Developmen...
2024
-
[15]
a m \"a l \
Perttu H \"a m \"a l \"a inen, Mikke Tavast, and Anton Kunnari. 2023. Evaluating large language models in generating synthetic hci research data: a case study. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1--19
2023
-
[16]
Qianyu He, Jie Zeng, Qianxi He, Jiaqing Liang, and Yanghua Xiao. 2024 a . From complex to simple: Enhancing multi-constraint complex instruction following ability of large language models. arXiv preprint arXiv:2404.15846
2024 arXiv
-
[17]
Zihao He, Minh Duc Chu, Rebecca Dorn, Siyi Guo, and Kristina Lerman. 2024 b . https://doi.org/10.18653/v1/2024.emnlp-main.945 Community-cross-instruct: Unsupervised instruction generation for aligning large language models to online communities . In Proceedings of the 2024 Con...
2024 doi
-
[18]
Zihao He, Siyi Guo, Ashwin Rao, and Kristina Lerman. 2024 c . https://doi.org/10.18653/v1/2024.findings-acl.395 Whose emotions and moral sentiments do language models reflect? In Findings of the Association for Computational Linguistics: ACL 2024, pages 6611--6631, Bangkok, Th...
2024 doi
-
[19]
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning ai with shared human values. In International Conference on Learning Representations
-
[20]
Smith, and Hannaneh Hajishirzi
Sachin Kumar, Chan Young Park, Yulia Tsvetkov, Noah A. Smith, and Hannaneh Hajishirzi. 2025. https://aclanthology.org/2025.naacl-long.419/ C om PO : Community preferences for language model personalization . In Proceedings of the 2025 Conference of the Nations of the Americas ...
2025
-
[21]
Cheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram, and Xing Xie. 2024 a . Culturellm: Incorporating cultural differences into large language models. Advances in Neural Information Processing Systems, 37:84799--84838
2024
-
[22]
Junyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2024 b . https://doi.org/10.18653/v1/2024.naacl-long.405 The steerability of large language models toward data-driven personas . In Proceedings of the 2024 Con...
2024 doi
-
[23]
Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason E Weston, and Mike Lewis. 2024 c . https://openreview.net/forum?id=1oijHJBRsT Self-alignment with instruction backtranslation . In The Twelfth International Conference on Learning Representations
2024
-
[24]
Ziyi Liu, Priyanka Dey, Zhenyu Zhao, Jen-tse Huang, Rahul Gupta, Yang Liu, and Jieyu Zhao. 2025. Can llms grasp implicit cultural values? benchmarking llms' metacognitive cultural intelligence with cq-bench. arXiv preprint arXiv:2504.01127
2025
-
[25]
Renze Lou, Kai Zhang, and Wenpeng Yin. 2024. Large language model instruction following: A survey of progresses and challenges. Computational Linguistics, 50(3):1053--1095
2024
-
[26]
Keming Lu, Bowen Yu, Chang Zhou, and Jingren Zhou. 2024. https://doi.org/10.18653/v1/2024.acl-long.423 Large language models are superpositions of all characters: Attaining arbitrary role-play via self-alignment . In Proceedings of the 62nd Annual Meeting of the Association fo...
2024 doi
-
[27]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, a...
2022
-
[28]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383--2392
2016
-
[29]
Yiting Ran, Xintao Wang, Rui Xu, Xinfeng Yuan, Jiaqing Liang, Yanghua Xiao, and Deqing Yang. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.853 Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data . In Findings of the ...
2024 doi
-
[30]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. https://proceedings.mlr.press/v202/santurkar23a.html Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, volume...
2023
-
[31]
Murray Shanahan, Kyle McDonell, and Laria Reynolds. 2023. Role play with large language models. Nature, 623(7987):493--498
2023
-
[32]
Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. https://aclanthology.org/2023.emnlp-main.814/ Character- LLM : A trainable agent for role-playing . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153--13187, Singapor...
2023
-
[33]
Taiwei Shi, Kai Chen, and Jieyu Zhao. 2024 a . https://doi.org/10.18653/v1/2024.naacl-long.422 Safer-instruct: Aligning language models with automated preference data . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lin...
2024 doi
-
[34]
Weiyan Shi, Ryan Li, Yutong Zhang, Caleb Ziems, Sunny Yu, Raya Horesh, Rog \'e rio Abreu De Paula, and Diyi Yang. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.288 C ulture B ank: An online community-driven knowledge base towards culturally aware language technologi...
2024 doi
-
[35]
Haoran Sun, Lixin Liu, Junjie Li, Fengyu Wang, Baohua Dong, Ran Lin, and Ruohui Huang. 2024. Conifer: Improving complex constrained instruction-following ability of large language models. arXiv preprint arXiv:2404.02823
2024 arXiv
-
[36]
Zhi Rui Tam, Cheng-Kuang Wu, Yi-Lin Tsai, Chieh-Yen Lin, Hung-yi Lee, and Yun-Nung Chen. 2024. https://doi.org/10.18653/v1/2024.emnlp-industry.91 Let me speak freely? a study on the impact of format restrictions on large language model performance. In Proceedings of the 2024 C...
2024 doi
-
[37]
Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024 a . https://doi.org/10.18653/v1/2024.findings-acl.878 R ole ...
2024 doi
-
[38]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...
2023 doi
-
[39]
Zifeng Wang, Chun-Liang Li, Vincent Perot, Long Le, Jin Miao, Zizhao Zhang, Chen-Yu Lee, and Tomas Pfister. 2024 b . https://doi.org/10.18653/v1/2024.findings-naacl.235 C odec LM : Aligning language models with tailored synthetic data . In Findings of the Association for Compu...
2024 doi
-
[40]
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024. https://openreview.net/forum?id=CfXh93NDgH Wizard LM : Empowering large pre-trained language models to follow complex instructions . In The Twelfth Internatio...
2024
-
[41]
Zhiyuan Zeng, Jiatong Yu, Tianyu Gao, Yu Meng, Tanya Goyal, and Danqi Chen. 2024. https://openreview.net/forum?id=tr0KidwPLc Evaluating large language models at evaluating instruction following . In The Twelfth International Conference on Learning Representations
2024
-
[42]
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and 1 others. 2023. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792
2023
-
[43]
Tao Zhang, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, and 1 others. 2024. Cfbench: A comprehensive constraints-following benchmark for llms. arXiv preprint arXiv:2408.01122
2024
-
[44]
Hao Zhao, Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. 2025. https://openreview.net/forum?id=STEEDDv3zI Is in-context learning sufficient for instruction following in LLM s? In The Thirteenth International Conference on Learning Representations
2025
-
[45]
Zijie Zhong, Linqing Zhong, Zhaoze Sun, Qingyun Jin, Zengchang Qin, and Xiaofan Zhang. 2025. https://aclanthology.org/2025.coling-main.46/ S ynthe T 2 C : Generating synthetic data for fine-tuning large language models on the T ext2 C ypher task . In Proceedings of the 31st In...
2025
-
[46]
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.