{"id":"5bb8f84c-2b68-465d-9fae-9ce72a33a9d3","arxiv_id":"2607.28814","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Penalizing confrontation in preference-optimized LLM counselors reliably lowers goal persistence; attunement gains are model-dependent, and penalizing capitulation is inert because on-policy capitulation is rare.","lead":"Researchers trained three LLM counselors to avoid either giving in to or arguing with clients in motivational interviewing, then measured whether they still pursued the client's change goal and respected their autonomy. Penalizing argumentative responses reliably reduced goal persistence, while gains in attunement varied by model and penalizing capitulation did nothing because the models rarely did it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central GP cost rests on an LLM judge with weak item-level construct validity; a human-majority pairwise GP replication is needed.","rationale":"The reader's weakest_assumption identifies the GP judge's construct validity as the load-bearing issue, and I agree. The central empirical outcome is a GP drop; all other findings (RA gain, capitulation inertia, gating) are secondary to that robust cost. The paper's own validity section shows that human item-level agreement on GP is poor (inter-human kappa=0.13, judge-vs-consensus kappa=0.38). While the pairwise human-judge direction agreement of 76% on 50 items is reassuring, it is not a direct estimate of the human-majority GP win-rate and is too small to establish the central effect on its own. The paper also reports that one of three coders systematically conflated persistence with pushiness, which is exactly the construct-confusion that would produce a judge artifact. The other potential concerns (small capitulation arms, unreproducible package URL) are real but secondary: the capitulation arms are explicitly described as underpowered, and the core trade-off does not depend on the inertness claim being exactly zero. The verdict should remain CONDITIONAL pending the proposed human-majority replication. No change to the reader's verdict is needed; if the test passes, the paper could be ACCEPTed, and if it fails, REJECTed.","tokens_in":23800,"tokens_out":4290,"duration_ms":48008,"concrete_test":"Use the 50 pairwise human judgments already collected (or a new stratified sample of ~80 contexts from the 142 test contexts, with three trained coders and majority vote) and compute the human-majority GP win-rate of the confrontation-penalizing variant against the base, not merely agreement with the judge. Pre-register the analysis. If the majority-vote GP win-rate is below 0.5 with a 95% CI excluding 0.5, and per-context direction agreement with the judge is >=80% on decided items, the concern is resolved. If the human-majority win-rate is at or above parity, the central cost is not established. Report the same for RA as a control.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that penalizing confrontation reliably lowers Goal Persistence below parity in all nine seed runs. But GP is measured by an automatic judge, and the manuscript's own validation (Appendix E, Table 12) shows human item-level agreement is low: pairwise quadratic-weighted kappa among coders is 0.13 for GP, and judge-vs-human-consensus weighted kappa is only 0.38. The 'quote-the-words' rubric refinement improved cross-judge reliability (kappa=0.73) but did not correspondingly anchor human agreement. The risk is not just noise; it is that the judge's GP construct may conflate 'keeps the change topic alive' with a surface style (explicit pushes, questions, agenda-keeping phrases) that the DPO update suppresses. The 50-item pairwise human-judge direction agreement of 76% is suggestive but does not directly report the human-majority GP win-rate of Dconf vs base. If a human majority, on the same pairwise task, did not place Dconf below base on GP, the 'robust cost' would be an artifact of the judge rather than a property of the optimization. The external validity check (RA separates high/low-quality sessions; GP does not in the same direction) is reassuring for RA but not for GP—the axis on which the headline negative result is measured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether preference optimization against a single MI failure mode teaches 'rolling with resistance' or trades one failure for another. It constructs topic-disjoint DPO data from AnnoMI in which the chosen responses are shared and the rejected responses are on-policy samples labeled as capitulation, confrontation, or a mix, selected by a single lever lambda. Across Qwen3-8B, Qwen2.5-7B, and Llama-3.1-8B, it measures blind pairwise win-rates against each base on two MITI-anchored axes, Goal Persistence (GP) and Relational Attunement (RA), using a firewalled LLM judge from a family disjoint from all generators. The main empirical claim is that penalizing confrontation lowers GP below parity in all nine seed runs and on all three bases, while raising RA on two of three bases; penalizing capitulation is inert; and a prompt-only control raises RA without the GP cost, locating the cost in the optimization rather than in attunement itself. Strengths include topic-disjoint held-out evaluation, on-policy negatives, judge-swap robustness, monotone lambda and beta ablations, full per-run counts, and an explicit code/data package.","tokens_in":24043,"tokens_out":8389,"duration_ms":91617,"significance":"If the central claim holds, the paper makes a substantive, clinically grounded contribution to preference-optimization research: it demonstrates a concrete trade-off between the two MITI-derived pathways and gives design guidance (a signal against both failure modes is needed). The experimental discipline is exemplary in several respects: the evaluation firewall, the topic-disjoint split, the judge-swap and beta/lambda ablations, the prompt-only control, and the honest appendix that reports every per-seed run. The significance is conditional on a load-bearing measurement issue: the GP axis, on which the headline negative result is measured, has weak item-level human construct validity. Until that is addressed, the 'robust cost' is an LLM-judge-relative finding rather than a demonstrated property of the MI construct.","major_comments":[{"comment":"The central, load-bearing claim is the GP drop below parity in all nine seed runs (Table 5). But the manuscript's own validation shows that the GP axis has low human item-level reliability: mean pairwise quadratic-weighted kappa among coders is 0.13, judge-vs-human-consensus kappa is 0.38, and Fleiss quadrant kappa is 0.11. Since Appendix A states that every reported effect is a difference in an LLM judge's blind pairwise preferences, and since absolute GP/RA scores are near ceiling (Table 11), the headline result rests entirely on the judge's GP policy. The reported 76% pairwise direction agreement is suggestive, but it is not the needed evidence: it does not give a human-majority pairwise GP win-rate for D_conf versus base on the same test contexts. I request a human-majority pairwise replication on a sizeable sample, or at minimum the human-majority pairwise GP win-rate with its confi","section":"§7.3, Table 12 / Appendix E"},{"comment":"The external validity evidence validates RA but not GP. The judge's RA separates expert-rated high- from low-quality MI sessions with a large effect (Cliff's delta = 0.71, p < 1e-15), but GP does not separate quality in the same direction (mean GP 1.66 vs 1.95, delta = -0.20). Thus the only externally anchored axis is the one on which no negative result is claimed, while the axis carrying the robust negative result, GP, has no external criterion supporting it. Cross-judge agreement alone (kappa = 0.73 for GP) shows that the judge's policy is stable, not that it measures what a human MI coder would call goal persistence. I ask for an external GP criterion (e.g., session-level behavioral labels that should track persistence, or a human-majority rating on the pairwise task) or, failing that, for the paper to explicitly weaken the 'robust cost' language to 'robust under the automatic judge'.","section":"§7.3 / Appendix D (external criterion)"},{"comment":"The conclusion that penalizing capitulation is 'inert' and that the trade-off is 'gated by the base's failure profile' is supported by a well-powered D_cap null only on Qwen3 (499 pairs). The D_cap sets on Qwen2.5 and Llama contain only 20 and 47 pairs, respectively, and with n=142 test contexts and high tie rates these cells have very wide intervals (e.g., Llama GP win 0.433, 95% CI [0.354, 0.515]). The paper's Table 8 footnote appropriately limits their interpretation, but Section 7.5 and the abstract still use the cross-base 'inert' wording as part of the gating explanation. The central D_conf finding is unaffected, but the asymmetry claim should either be backed by larger D_cap runs on the replication bases or explicitly confined to Qwen3.","section":"§7.5, Table 8 / Table 13"}],"minor_comments":[{"comment":"The main text reports Qwen3 lengths as 51.0 vs 48.6 tokens, while Table 14 reports 43.0 vs 40.8 for the same base/D_conf comparison. Since these numbers are used to rule out a length artifact, please reconcile the definitions or correct the inconsistency.","section":"§7.5 vs. Table 14"},{"comment":"The phrase 'win counts ties as one half; w/t/l are the raw counts' is ambiguous. State explicitly that the reported win-rate is (wins + 0.5 ties)/n and that w/t/l are raw counts, with ties counted as neither wins nor losses in the McNemar tests.","section":"Table 13 caption"},{"comment":"The fixed phrase lists for directive, concession, and reflection markers are described as 'fixed in advance' but not reproduced. Include the full lists in the supplementary materials so the lexical mechanism analysis is reproducible.","section":"Appendix I"},{"comment":"The table footnote says D_mix is not used for any reported run, while the 'reject both' cell is D_lambda_0.50 from the frontier sweep. This is confusing; label the table so it is clear which set corresponds to the reported 'reject both' cell.","section":"§6, Table 8"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is unusually transparent and well-controlled. My central concern is precisely the one the authors flag in Appendix A: the headline result is a judge-relative quantity on a construct with weak human item-level agreement. If the authors add a human-majority pairwise GP replication that confirms the D_conf-below-base direction, I would be comfortable with acceptance. The current evidence is sufficient for a careful negative-result-in-MI story, but not yet for the strong cross-base 'robust cost' claim as a property of the GP construct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid empirical paper with a genuinely new finding, and the main caveat is the one the authors themselves flag — the Goal Persistence axis is the weakest measured construct. I don't think that kills the result, but it should be front-and-center in any revision.\n\nWhat's new: they isolate which failure mode is penalized in DPO for MI counselors, using preference sets that share positives and differ only in the rejected response. Across three aligned models, penalizing confrontation reliably lowers goal persistence below parity in all nine seed runs, while the attunement gain is inconsistent. Penalizing capitulation is inert because these models rarely capitulate on-policy. The prompt-only control is the key contrast: it raises attunement without the GP cost, so the cost is in the optimization, not in attunement itself. That's a clean, non-obvious result and it's well-supported by the reported numbers: full per-run tables, pooled McNemar, judge-swap reproduction, lambda sweep, beta ablation. The design is unusually careful — topic-disjoint splits, on-policy negatives, firewalled judge families. Credit where due.\n\nThe soft spot is the GP measure. Human pairwise kappa is 0.13 at the item level, judge-vs-consensus only 0.38, and the headline negative result (GP drop) is measured by that judge. The authors are transparent about this — Appendix A says every reported effect is a difference in an LLM judge's blind pairwise preferences. They also give some external anchor: RA separates expert-rated high/low sessions, but GP does not in the same direction, which is reassuring for RA but not for GP. I agree with the stress-test note that a human-majority pairwise GP replication would be the decisive check. Still, the direction is robust across judge swaps, and the 76% pairwise direction agreement with human majority is at least suggestive. I'd treat the GP cost as probable but not fully nailed down.\n\nMinor: the code/data claim has no locatable URL or commit hash, which matters for a paper that says 'every number is recomputable.' And the capitulation arms on the replication bases are underpowered (20-47 pairs) — though the authors call this themselves.\n\nBottom line: worth a serious referee, especially for people working on alignment side effects in domain-specific generation and clinical NLP. It's a careful study of a real safety-relevant side effect in a clinically grounded setting. The main fix is either a human replication of the GP pairwise result or a more conservative framing of the central claim.","headline":"Careful, well-designed study of a real alignment side effect; the GP measurement worry is real but doesn't overturn the result.","tokens_in":24601,"tokens_out":2041,"would_cite":true,"duration_ms":21349,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Within motivational interviewing, a preference-optimized LLM counselor trained to suppress confrontation reliably loses goal persistence across three base models, while the relational-attunement gain is inconsistent, and training against ca","keywords":["motivational interviewing","preference optimization","direct preference optimization","goal persistence","relational attunement","rolling with resistance","LLM-as-judge","on-policy negatives"],"falsifier":"Re-run the 142 test comparisons with a panel of trained MI coders providing the pairwise GP preference instead of the LLM judge; if the human-majority GP win-rate for the confrontation-penalized variant does not fall below 0.5, the reported trade-off is an artifact of the rubric.","tokens_in":23583,"feed_emoji":"🤝","tokens_out":5176,"duration_ms":50768,"temperature":0.7,"pith_summary":"This paper asks whether punishing one failure mode in motivational interviewing—either capitulating to the client or confronting them—teaches an LLM counselor to roll with resistance, or merely provokes the opposite failure. It claims that a single-mode preference signal does not teach rolling with resistance for free: on three aligned instruction models, penalizing confrontation consistently pushes goal persistence below parity (pairwise win-rates 0.38–0.40 against each base, pooled p<1e-5), while the gain in relational attunement appears on only two of the three bases. Penalizing capitulation is inert, because the models rarely capitulate on their own. A prompt-only control raises attunement without the goal-persistence drop, locating the cost in the optimization procedure rather than in attunement itself. The paper matters because preference optimization is the standard alignment tool for shaping interpersonal behavior in LLMs, and a plausible-looking training signal can make a counselor less willing to keep a change agenda alive.","feed_headline":"LLM counselors trade persistence for warmth under one-sided training","feed_subtitle":"Penalizing confrontation drops goal persistence below parity on every model; prompt-only instruction avoids the cost.","key_machinery":"The machinery is a two-axis rubric anchored in the Motivational Interviewing Treatment Integrity code—Goal Persistence (0–3) and Relational Attunement (0–3), thresholded at 2 to define four quadrants (rolling-with, capitulation, confrontation, collapse)—used by an automatic judge to score responses, plus DPO preference sets whose rejected pool is selected by a single lever λ (fraction confrontation), with on-policy negatives and a three-way firewall (generator, training-label judge, and evaluation judge from disjoint model families). The rubric's GP axis required a quote-the-words refinement to make the 1-versus-2 boundary checkable. Evaluation is reported as blind pairwise win-rate against","core_discovery":"The central discovery is a gated trade-off. Scoring counselor responses on two MITI-anchored axes—goal persistence (keeps the session on the change the client is weighing) and relational attunement (honors the client's autonomy)—the paper builds DPO preference sets that differ only in which failure supplies the rejected response, using on-policy negatives, and evaluates blind pairwise win-rates against each base. Penalizing confrontation lowers goal persistence below parity on every base and in every seed run, while raising attunement on two of three bases; penalizing capitulation moves neither axis, because aligned models almost never capitulate on-policy. The paper interprets this as evide","pith_inferences":["The GP/RA split maps naturally onto the technical versus relational pathways of MI; if that mapping holds, the trade-off suggests a structural tension inside MI practice, not just an artifact of DPO, and multi-objective alignment methods that optimize both signals jointly are the obvious next test.","The weak human agreement on single-utterance GP coding hints that goal persistence may be a session-level construct; a session-level measurement could shrink or enlarge the observed trade-off.","The prompt-only result implies that at inference time a model can be steered toward attunement without sacrificing persistence, so the optimization cost may be avoidable in deployment even though it is real in training.","The judge-swap robustness suggests the effect is not a preference-label artifact, but the missing human pairwise replication on the GP axis is the decisive check: until trained coders reproduce the below-parity GP win-rate, the trade-off's magnitude rests on an automatic judge."],"forward_implications":["A one-sided preference signal against a single MI failure mode does not teach the target behavior; it provokes the opposite failure, so practical resistance-aware training needs a signal against both capitulation and confrontation.","On-policy negatives are required to move generation; scripted off-policy negatives are learned in the loss but leave behavior unchanged.","Inference-time prompting can raise attunement substantially (RA win-rate 0.81) with no significant goal-persistence drop, so some attunement gains are available without optimization cost.","The effect of penalizing confrontation is seed-stable and reproducible across the three bases; relaxing the KL anchor amplifies the goal-persistence drop, consistent with preference over-optimization.","The trade-off is gated by the base's on-policy failure profile: models that rarely capitulate show no response to anti-capitulation training."],"fun_headline_variants":["Penalizing confrontation costs LLM counselors goal persistence","LLM counselors lose persistence when trained to avoid confrontation","Prompt-only instruction boosts attunement without persistence cost","Avoiding confrontation in LLM training hurts goal persistence"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the automatic judge's pairwise Goal Persistence preference measures a stable, human-validable construct; the paper's own human recheck shows item-level agreement on GP is weak (weighted kappa 0.13 among coders, 0.38 judge-vs-consensus), so if GP win-rates do not track a real construct the central drop could be a rubric artifact.","fun_headline_variants_meta":{"raw":{"variants":["Penalizing confrontation costs LLM counselors goal persistence","LLM counselors lose persistence when trained to avoid confrontation","Prompt-only instruction boosts attunement without persistence cost","Avoiding confrontation in LLM training hurts goal persistence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000312,"raw_usage":{"total_tokens":1666,"prompt_tokens":855,"completion_tokens":811,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":748}},"tokens_in":599,"tokens_out":811,"duration_ms":7773,"temperature":1.0,"reasoning_tokens":748,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T00:18:17.200354+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 142 test comparisons with a panel of trained MI coders providing the pairwise GP preference instead of the LLM judge; if the human-majority GP win-rate for the confrontation-penalized variant does not fall below 0.5, the reported trade-off is an artifact of the rubric.","supporting_citations":[],"review_version":1}