REVIEW 3 major objections 4 minor 60 references
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Within motivational interviewing, a preference-optimized LLM counselor trained to suppress confrontation reliably loses goal persistence across three base models, while the relational-attunement gain is inconsistent, and training against ca
desk verdict Careful, well-designed study of a real alignment side effect; the GP measurement worry is real but doesn't overturn the result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-axis rubric anchored in the Motivational Interviewing Treatment Integrity code—Goal Persistence (0–3) and Relational Attunement (0–3), thresholded at 2 to define four quadrants (rolling-with, capitulation, confrontation, collapse)—used by an automatic judge to score responses, plus DPO preference sets whose rejected pool is selected by a single lever λ (fraction confrontation), with on-policy negatives and a three-way firewall (generator, training-label judge, and evaluation judge from disjoint model families). The rubric's GP axis required a quote-the-words refinement to make the 1-versus-2 boundary checkable. Evaluation is reported as blind pairwise win-rate against
What would settle it
Re-run the 142 test comparisons with a panel of trained MI coders providing the pairwise GP preference instead of the LLM judge; if the human-majority GP win-rate for the confrontation-penalized variant does not fall below 0.5, the reported trade-off is an artifact of the rubric.
Extended reading notes
Core claim
The central discovery is a gated trade-off. Scoring counselor responses on two MITI-anchored axes—goal persistence (keeps the session on the change the client is weighing) and relational attunement (honors the client's autonomy)—the paper builds DPO preference sets that differ only in which failure supplies the rejected response, using on-policy negatives, and evaluates blind pairwise win-rates against each base. Penalizing confrontation lowers goal persistence below parity on every base and in every seed run, while raising attunement on two of three bases; penalizing capitulation moves neither axis, because aligned models almost never capitulate on-policy. The paper interprets this as evide
Load-bearing premise
The load-bearing premise is that the automatic judge's pairwise Goal Persistence preference measures a stable, human-validable construct; the paper's own human recheck shows item-level agreement on GP is weak (weighted kappa 0.13 among coders, 0.38 judge-vs-consensus), so if GP win-rates do not track a real construct the central drop could be a rubric artifact.
Editorial extensions
If this is right
- A one-sided preference signal against a single MI failure mode does not teach the target behavior; it provokes the opposite failure, so practical resistance-aware training needs a signal against both capitulation and confrontation.
- On-policy negatives are required to move generation; scripted off-policy negatives are learned in the loss but leave behavior unchanged.
- Inference-time prompting can raise attunement substantially (RA win-rate 0.81) with no significant goal-persistence drop, so some attunement gains are available without optimization cost.
- The effect of penalizing confrontation is seed-stable and reproducible across the three bases; relaxing the KL anchor amplifies the goal-persistence drop, consistent with preference over-optimization.
- The trade-off is gated by the base's on-policy failure profile: models that rarely capitulate show no response to anti-capitulation training.
Reading between the lines
- The GP/RA split maps naturally onto the technical versus relational pathways of MI; if that mapping holds, the trade-off suggests a structural tension inside MI practice, not just an artifact of DPO, and multi-objective alignment methods that optimize both signals jointly are the obvious next test.
- The weak human agreement on single-utterance GP coding hints that goal persistence may be a session-level construct; a session-level measurement could shrink or enlarge the observed trade-off.
- The prompt-only result implies that at inference time a model can be steered toward attunement without sacrificing persistence, so the optimization cost may be avoidable in deployment even though it is real in training.
- The judge-swap robustness suggests the effect is not a preference-label artifact, but the missing human pairwise replication on the GP axis is the decisive check: until trained coders reproduce the below-parity GP win-rate, the trade-off's magnitude rests on an automatic judge.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether preference optimization against a single MI failure mode teaches 'rolling with resistance' or trades one failure for another. It constructs topic-disjoint DPO data from AnnoMI in which the chosen responses are shared and the rejected responses are on-policy samples labeled as capitulation, confrontation, or a mix, selected by a single lever lambda. Across Qwen3-8B, Qwen2.5-7B, and Llama-3.1-8B, it measures blind pairwise win-rates against each base on two MITI-anchored axes, Goal Persistence (GP) and Relational Attunement (RA), using a firewalled LLM judge from a family disjoint from all generators. The main empirical claim is that penalizing confrontation lowers GP below parity in all nine seed runs and on all three bases, while raising RA on two of three bases; penalizing capitulation is inert; and a prompt-only control raises RA without the GP cost, locating the cost in the optimization rather than in attunement itself. Strengths include topic-disjoint held-out evaluation, on-policy negatives, judge-swap robustness, monotone lambda and beta ablations, full per-run counts, and an explicit code/data package.
Significance. If the central claim holds, the paper makes a substantive, clinically grounded contribution to preference-optimization research: it demonstrates a concrete trade-off between the two MITI-derived pathways and gives design guidance (a signal against both failure modes is needed). The experimental discipline is exemplary in several respects: the evaluation firewall, the topic-disjoint split, the judge-swap and beta/lambda ablations, the prompt-only control, and the honest appendix that reports every per-seed run. The significance is conditional on a load-bearing measurement issue: the GP axis, on which the headline negative result is measured, has weak item-level human construct validity. Until that is addressed, the 'robust cost' is an LLM-judge-relative finding rather than a demonstrated property of the MI construct.
major comments (3)
- [§7.3, Table 12 / Appendix E] The central, load-bearing claim is the GP drop below parity in all nine seed runs (Table 5). But the manuscript's own validation shows that the GP axis has low human item-level reliability: mean pairwise quadratic-weighted kappa among coders is 0.13, judge-vs-human-consensus kappa is 0.38, and Fleiss quadrant kappa is 0.11. Since Appendix A states that every reported effect is a difference in an LLM judge's blind pairwise preferences, and since absolute GP/RA scores are near ceiling (Table 11), the headline result rests entirely on the judge's GP policy. The reported 76% pairwise direction agreement is suggestive, but it is not the needed evidence: it does not give a human-majority pairwise GP win-rate for D_conf versus base on the same test contexts. I request a human-majority pairwise replication on a sizeable sample, or at minimum the human-majority pairwise GP win-rate with its confi
- [§7.3 / Appendix D (external criterion)] The external validity evidence validates RA but not GP. The judge's RA separates expert-rated high- from low-quality MI sessions with a large effect (Cliff's delta = 0.71, p < 1e-15), but GP does not separate quality in the same direction (mean GP 1.66 vs 1.95, delta = -0.20). Thus the only externally anchored axis is the one on which no negative result is claimed, while the axis carrying the robust negative result, GP, has no external criterion supporting it. Cross-judge agreement alone (kappa = 0.73 for GP) shows that the judge's policy is stable, not that it measures what a human MI coder would call goal persistence. I ask for an external GP criterion (e.g., session-level behavioral labels that should track persistence, or a human-majority rating on the pairwise task) or, failing that, for the paper to explicitly weaken the 'robust cost' language to 'robust under the automatic judge'.
- [§7.5, Table 8 / Table 13] The conclusion that penalizing capitulation is 'inert' and that the trade-off is 'gated by the base's failure profile' is supported by a well-powered D_cap null only on Qwen3 (499 pairs). The D_cap sets on Qwen2.5 and Llama contain only 20 and 47 pairs, respectively, and with n=142 test contexts and high tie rates these cells have very wide intervals (e.g., Llama GP win 0.433, 95% CI [0.354, 0.515]). The paper's Table 8 footnote appropriately limits their interpretation, but Section 7.5 and the abstract still use the cross-base 'inert' wording as part of the gating explanation. The central D_conf finding is unaffected, but the asymmetry claim should either be backed by larger D_cap runs on the replication bases or explicitly confined to Qwen3.
minor comments (4)
- [§7.5 vs. Table 14] The main text reports Qwen3 lengths as 51.0 vs 48.6 tokens, while Table 14 reports 43.0 vs 40.8 for the same base/D_conf comparison. Since these numbers are used to rule out a length artifact, please reconcile the definitions or correct the inconsistency.
- [Table 13 caption] The phrase 'win counts ties as one half; w/t/l are the raw counts' is ambiguous. State explicitly that the reported win-rate is (wins + 0.5 ties)/n and that w/t/l are raw counts, with ties counted as neither wins nor losses in the McNemar tests.
- [Appendix I] The fixed phrase lists for directive, concession, and reflection markers are described as 'fixed in advance' but not reproduced. Include the full lists in the supplementary materials so the lexical mechanism analysis is reproducible.
- [§6, Table 8] The table footnote says D_mix is not used for any reported run, while the 'reject both' cell is D_lambda_0.50 from the frontier sweep. This is confusing; label the table so it is clear which set corresponds to the reported 'reject both' cell.
Circularity Check
No significant circularity: the GP cost is a held-out, firewalled empirical result with non-circular controls; the LLM-judge validity concern is a measurement limitation, not a derivation that reduces to its inputs.
full rationale
The paper's derivation chain is self-contained and empirically controlled. Preference sets are built from AnnoMI contexts split by topic before generation (Appendix B), with shared positive pools and rejected pools varied only by the failure mode selected through λ (Section 6). The headline result — confrontation-penalizing DPO lowers pairwise GP win-rate below parity on all three bases — is measured on the held-out 142 test contexts by an evaluation judge from a disjoint model family, blind to variant identity, with position randomized (Sections 7.1, 7.3). No reported number is a fitted constant; the rubric was defined before training and validated against external AnnoMI expert labels, session-quality splits, and a human recheck. The non-circularity is directly shown by the paper's own controls: if the GP drop were merely the training signal re-read as evaluation, then Dcap (reject capitulation) would also move GP/RA, but it is inert (GP 0.52/RA 0.53 on Qwen3), and prompt-only would show the cost, but it raises RA (0.81) at a smaller GP cost (0.45), locating the effect in optimization rather than in attunement itself. The one residual concern — the automatic judge's GP construct has weak item-level human agreement (mean pairwise weighted kappa 0.13, judge-vs-consensus 0.38; Appendix E, Table 12) — is a measurement-validity threat that the paper acknowledges explicitly ("item-level human agreement on the GP axis is modest"), not a circularity: the judge was not fitted to produce the GP drop, and the finding is a difference in blind pairwise preferences, not a constant recovered from the training data. There are no load-bearing self-citations: the cited prior work (AnnoMI, MITI, DPO, over-optimization literature) is external and independently checkable, and no uniqueness theorem or prior result by the same authors is invoked to force the choice of GP/RA or the interpretation of the trade-off. The paper therefore does not reduce to its own inputs under any of the enumerated circularity patterns.
Assumptions & free parameters
free parameters (3)
- DPO KL strength beta =
0.5 (chosen; ablated 0.05/0.3/0.5)
- LoRA/training hyperparameters =
rank 16, alpha 32, dropout 0.05, LR 1e-5, 3 epochs, batch 16, seq 1024
- On-policy negative sampling temperature =
0.9
assumptions (6)
- domain assumption AnnoMI expert annotations (majority-vote therapist behavior and client talk-type labels) are valid ground truth for MI constructs.
- domain assumption The GP/RA rubric operationalizes MITI 'roll with resistance' as GP>=2 and RA>=2, with threshold at 2.
- domain assumption Disjoint-family LLM judges provide valid GP/RA scores when using the v2 rubric.
- domain assumption Splitting AnnoMI by topic before generation and never recomputing the map prevents leakage.
- standard math DPO with LoRA adapters on the frozen base is a valid route to test the effect of preference direction.
- domain assumption The three model families are mutually disjoint so the evaluation firewall is effective.
invented entities (2)
-
Goal Persistence (GP) axis (0-3)
-
Relational Attunement (RA) axis (0-3)
independent evidence
Cite this review
Pith. "Pith review of Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing." pith.science (2026). https://pith.science/paper/BH7TPMED
@misc{pith2026260728814,
author = {Pith},
title = {Pith review of: Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing},
year = {2026},
howpublished = {\url{https://pith.science/paper/BH7TPMED}},
note = {Machine review of arXiv:2607.28814}
}
read the original abstract
In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.
Figures
Reference graph
Works this paper leans on
-
[2]
A.; and Krahmer, E
Basar, E.; Sun, X.; Hendrickx, I.; de Wit, J.; Bosse, T.; de Bruijn, G.-J.; Bosch, J. A.; and Krahmer, E. 2025. How Well Can Large Language Models Reflect? A Human Evaluation of LLM -generated Reflections for Motivational Interviewing Dialogues. In Proceedings of the 31st International Conference on Computational Linguistics (COLING), 1964--1982
2025
-
[3]
Casper, S.; Davies, X.; Shi, C.; Gilbert, T. K.; Scheurer, J.; Rando, J.; Freedman, R.; Korbak, T.; Lindner, D.; Freire, P.; Wang, T.; Marks, S.; Segerie, C.-R.; Carroll, M.; Peng, A.; Christoffersen, P.; Damani, M.; Slocum, S.; Anwar, U.; Siththaranjan, A.; Nadeau, M.; Michaud, E. J.; Pfau, J.; Krasheninnikov, D.; Chen, X.; Langosco, L.; Hase, P.; Biyik,...
2023
-
[4]
Dai, J.; Chen, T.; Yang, Y.; Zheng, Q.; and Pan, G. 2025. Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization. In International Conference on Learning Representations (ICLR)
2025
-
[5]
Dai, J.; Pan, X.; Sun, R.; Ji, J.; Xu, X.; Liu, M.; Wang, Y.; and Yang, Y. 2024. Safe RLHF : Safe Reinforcement Learning from Human Feedback. In International Conference on Learning Representations (ICLR)
2024
-
[6]
Dubois, Y.; Galambosi, B.; Liang, P.; and Hashimoto, T. B. 2024. Length-Controlled AlpacaEval : A Simple Way to Debias Automatic Evaluators. In Conference on Language Modeling (COLM)
2024
-
[7]
Gao, L.; Schulman, J.; and Hilton, J. 2023. Scaling Laws for Reward Model Overoptimization. In International Conference on Machine Learning (ICML), 10835--10866
2023
-
[8]
He, Q.; and Maghsudi, S. 2025. Pareto Multi-Objective Alignment for Language Models. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD)
2025
-
[10]
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2511--2522
2023
Show all 60 references
-
[11]
R.; Borsari, B.; Gaume, J.; Hoadley, A.; Gordon, R
Magill, M.; Apodaca, T. R.; Borsari, B.; Gaume, J.; Hoadley, A.; Gordon, R. E. F.; Tonigan, J. S.; and Moyers, T. B. 2018. A Meta-Analysis of Motivational Interviewing Process: Technical, Relational, and Conditional Process Models of Change. Journal of Consulting and Clinical ...
2018
-
[12]
R.; and Rollnick, S
Miller, W. R.; and Rollnick, S. 2013. Motivational Interviewing: Helping People Change. Guilford Press, 3rd edition
2013
-
[13]
J.; P \'e rez-Rosas, V.; Resnicow, K.; and Mihalcea, R
Min, D. J.; P \'e rez-Rosas, V.; Resnicow, K.; and Mihalcea, R. 2024. Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resource...
2024
-
[14]
B.; Rowell, L
Moyers, T. B.; Rowell, L. N.; Manuel, J. K.; Ernst, D.; and Houck, J. M. 2016. The Motivational Interviewing Treatment Integrity Code (MITI 4): Rationale, Preliminary Reliability and Validity. Journal of Substance Abuse Treatment, 65: 36--42
2016
-
[15]
Otani, A. 1989. Client Resistance in Counseling: Its Theoretical Rationale and Taxonomic Classification. Journal of Counseling and Development, 67(8): 458--461
1989
-
[16]
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022. Training Language Models to Follow Instructions with Human Feedback. In Advances in Neural Information Processing Systems (NeurIPS), volume 35, 277...
2022
-
[17]
R.; and Feng, S
Panickssery, A.; Bowman, S. R.; and Feng, S. 2024. LLM Evaluators Recognize and Favor Their Own Generations. In Advances in Neural Information Processing Systems (NeurIPS)
2024
-
[18]
Perez, E.; Ringer, S.; Luko s i \=u t \.e , K.; Nguyen, K.; Chen, E.; et al. 2023. Discovering Language Model Behaviors with Model-Written Evaluations. Findings of the Association for Computational Linguistics (ACL)
2023
-
[19]
D.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In Advances in Neural Information Processing Systems (NeurIPS), volume 36
2023
-
[20]
Ram \'e , A.; Couairon, G.; Dancette, C.; Gaya, J.-B.; Shukor, M.; Soulier, L.; and Cord, M. 2023. Rewarded Soups: Towards Pareto-Optimal Alignment by Interpolating Weights Fine-Tuned on Diverse Rewards. In Advances in Neural Information Processing Systems (NeurIPS), volume 36
2023
-
[21]
Sharma, M.; Tong, M.; Korbak, T.; Duvenaud, D.; Askell, A.; et al. 2024. Towards Understanding Sycophancy in Language Models. In International Conference on Learning Representations (ICLR)
2024
-
[22]
Skalse, J.; Howe, N. H. R.; Krasheninnikov, D.; and Krueger, D. 2022. Defining and Characterizing Reward Gaming. In Advances in Neural Information Processing Systems (NeurIPS)
2022
-
[23]
Sun, X.; Tang, X.; El Ali, A.; Li, Z.; Ren, P.; de Wit, J.; Pei, J.; and Bosch, J. A. 2025. Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies. In Proceedings of the 31st International Conference on Computational Linguistics...
2025
-
[24]
Tajwar, F.; Singh, A.; Sharma, A.; Rafailov, R.; Schneider, J.; Xie, T.; Ermon, S.; Finn, C.; and Kumar, A. 2024. Preference Fine-Tuning of LLM s Should Leverage Suboptimal, On-Policy Data. In International Conference on Machine Learning (ICML), 47441--47474
2024
-
[25]
Wang, H.; Lin, Y.; Xiong, W.; Yang, R.; Diao, S.; Qiu, S.; Zhao, H.; and Zhang, T. 2024 a . Arithmetic Control of LLM s for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards. In Proceedings of the 62nd Annual Meeting of the Association for...
2024
-
[26]
Wang, P.; Li, L.; Chen, L.; Cai, Z.; Zhu, D.; Lin, B.; Cao, Y.; Kong, L.; Liu, Q.; Liu, T.; and Sui, Z. 2024 b . Large Language Models are not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), 9440--9450
2024
-
[27]
R.; He, H.; and Feng, S
Wen, J.; Zhong, R.; Khan, A.; Perez, E.; Steinhardt, J.; Huang, M.; Bowman, S. R.; He, H.; and Feng, S. 2025. Language Models Learn to Mislead Humans via RLHF . In International Conference on Learning Representations (ICLR)
2025
-
[28]
Wu, Z.; Balloccu, S.; Kumar, V.; Helaoui, R.; Reiter, E.; Reforgiato Recupero, D.; and Riboni, D. 2022. Anno- MI : A Dataset of Expert-Annotated Counselling Dialogues. In ICASSP 2022 -- IEEE International Conference on Acoustics, Speech and Signal Processing, 6177--6181
2022
-
[29]
A.; Ostendorf, M.; and Hajishirzi, H
Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023. Fine-Grained Human Feedback Gives Better Rewards for Language Model Training. In Advances in Neural Information Processing Systems (NeurIPS), volume 36
2023
-
[30]
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023. Judging LLM -as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems (NeurIPS), volume 36
2023
-
[31]
Zhou, Z.; Liu, J.; Shao, J.; Yue, X.; Yang, C.; Ouyang, W.; and Qiao, Y. 2024. Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization. In Findings of the Association for Computational Linguistics (ACL), 10586--10613
2024
-
[32]
2013 , edition=
Motivational Interviewing: Helping People Change , author=. 2013 , edition=
2013
-
[33]
Journal of Substance Abuse Treatment , volume=
The Motivational Interviewing Treatment Integrity Code (MITI 4): Rationale, Preliminary Reliability and Validity , author=. Journal of Substance Abuse Treatment , volume=
-
[34]
Wu, Zixiu and Balloccu, Simone and Kumar, Vivek and Helaoui, Rim and Reiter, Ehud and Reforgiato Recupero, Diego and Riboni, Daniele , booktitle=. Anno-
-
[35]
Journal of Counseling and Development , volume=
Client Resistance in Counseling: Its Theoretical Rationale and Taxonomic Classification , author=. Journal of Counseling and Development , volume=
-
[36]
Advances in Neural Information Processing Systems (NeurIPS) , volume=
Training Language Models to Follow Instructions with Human Feedback , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[37]
Advances in Neural Information Processing Systems (NeurIPS) , volume=
Direct Preference Optimization: Your Language Model is Secretly a Reward Model , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[38]
International Conference on Machine Learning (ICML) , pages=
Scaling Laws for Reward Model Overoptimization , author=. International Conference on Machine Learning (ICML) , pages=
-
[39]
Findings of the Association for Computational Linguistics (ACL) , year=
Discovering Language Model Behaviors with Model-Written Evaluations , author=. Findings of the Association for Computational Linguistics (ACL) , year=
-
[40]
International Conference on Learning Representations (ICLR) , year=
Towards Understanding Sycophancy in Language Models , author=. International Conference on Learning Representations (ICLR) , year=
-
[41]
Advances in Neural Information Processing Systems (NeurIPS) , volume=
Rewarded Soups: Towards Pareto-Optimal Alignment by Interpolating Weights Fine-Tuned on Diverse Rewards , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[42]
Findings of the Association for Computational Linguistics (ACL) , pages=
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization , author=. Findings of the Association for Computational Linguistics (ACL) , pages=
-
[43]
arXiv preprint arXiv:2503.11701 , year=
A Survey of Direct Preference Optimization , author=. arXiv preprint arXiv:2503.11701 , year=
-
[44]
Machine Learning and Knowledge Discovery in Databases (ECML PKDD) , year=
Pareto Multi-Objective Alignment for Language Models , author=. Machine Learning and Knowledge Discovery in Databases (ECML PKDD) , year=
-
[45]
International Conference on Learning Representations (ICLR) , year=
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization , author=. International Conference on Learning Representations (ICLR) , year=
-
[46]
Advances in Neural Information Processing Systems (NeurIPS) , volume=
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , author=. Advances in Neural Information Processing Systems (NeurIPS) , volume=
-
[47]
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and Zhuang, Siyuan and Wu, Zhanghao and Zhuang, Yonghao and Lin, Zi and Li, Zhuohan and Li, Dacheng and Xing, Eric and others , booktitle=. Judging
-
[48]
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING) , pages=
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation , author=. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING) , pages=
2024
-
[49]
Proceedings of the 31st International Conference on Computational Linguistics (COLING) , pages=
Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing Strategies , author=. Proceedings of the 31st International Conference on Computational Linguistics (COLING) , pages=
-
[50]
and Krahmer, Emiel , booktitle=
Basar, Erkan and Sun, Xin and Hendrickx, Iris and de Wit, Jan and Bosse, Tibor and de Bruijn, Gert-Jan and Bosch, Jos A. and Krahmer, Emiel , booktitle=. How Well Can Large Language Models Reflect? A Human Evaluation of
-
[51]
Transactions on Machine Learning Research , year=
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback , author=. Transactions on Machine Learning Research , year=
-
[52]
and He, He and Feng, Shi , booktitle=
Wen, Jiaxin and Zhong, Ruiqi and Khan, Akbir and Perez, Ethan and Steinhardt, Jacob and Huang, Minlie and Bowman, Samuel R. and He, He and Feng, Shi , booktitle=. Language Models Learn to Mislead Humans via
-
[53]
Dai, Josef and Pan, Xuehai and Sun, Ruiyang and Ji, Jiaming and Xu, Xinbo and Liu, Mickel and Wang, Yizhou and Yang, Yaodong , booktitle=. Safe
-
[54]
Arithmetic Control of
Wang, Haoxiang and Lin, Yong and Xiong, Wei and Yang, Rui and Diao, Shizhe and Qiu, Shuang and Zhao, Han and Zhang, Tong , booktitle=. Arithmetic Control of
-
[55]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , pages=
Large Language Models are not Fair Evaluators , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , pages=
-
[56]
and Feng, Shi , booktitle=
Panickssery, Arjun and Bowman, Samuel R. and Feng, Shi , booktitle=
-
[57]
Liu, Yang and Iter, Dan and Xu, Yichong and Wang, Shuohang and Xu, Ruochen and Zhu, Chenguang , booktitle=. G-Eval:
-
[58]
Preference Fine-Tuning of
Tajwar, Fahim and Singh, Anikait and Sharma, Archit and Rafailov, Rafael and Schneider, Jeff and Xie, Tengyang and Ermon, Stefano and Finn, Chelsea and Kumar, Aviral , booktitle=. Preference Fine-Tuning of
-
[59]
Advances in Neural Information Processing Systems (NeurIPS) , year=
Defining and Characterizing Reward Gaming , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[60]
arXiv preprint arXiv:2204.05862 , year=
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback , author=. arXiv preprint arXiv:2204.05862 , year=
-
[61]
Length-Controlled
Dubois, Yann and Galambosi, Bal. Length-Controlled. Conference on Language Modeling (COLM) , year=
-
[62]
Journal of Consulting and Clinical Psychology , volume=
A Meta-Analysis of Motivational Interviewing Process: Technical, Relational, and Conditional Process Models of Change , author=. Journal of Consulting and Clinical Psychology , volume=
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.