{"id":"dfaafebc-4f13-4bd7-8b50-758b8c353419","arxiv_id":"2504.17083","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"People prefer different chatbot writing styles depending on their own personality and trust in LLMs, according to two preliminary regression-based studies.","lead":"This paper asks whether the writing style of a chatbot's answers changes which answers people prefer. It finds that style does matter and that a person's personality and trust in LLMs shift which styles they like best, based on a small preliminary experiment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated style-transfer and style-measurement pipelines leave the trait-moderation claim confounded by possible content changes.","rationale":"The central claim has two components: language style influences preference, and individual traits moderate which styles matter. The first is supported by large secondary datasets and is plausible. The distinctive second component rests entirely on Study 2, where the stimulus generation and measurement loop is unvalidated and the sample is very small. The reader's weakest assumption identifies exactly this loop, and I agree that it is the most load-bearing point. My reading does not uncover an internal inconsistency or a fatal flaw; rather, it reinforces the need for a validity check before the trait-moderation results can be taken as evidence about language style. If the manipulation check passes, the conditional verdict stands. If it fails, the specific moderation claims would need to be retracted or reframed as effects of content or other confounded dimensions. For that reason, the reader's CONDITIONAL verdict does not need to change.","tokens_in":10230,"tokens_out":7502,"duration_ms":68960,"concrete_test":"Run a manipulation-check study on the generated stimuli (3 queries × 9 styles × 3 levels). Have two independent annotators, or a held-out judge model, rate every response on all nine style dimensions plus content-equivalence and factuality relative to the baseline. Compute the correlation matrix between intended intensity levels and rated dimensions. If any off-target dimension correlates with an intended level as strongly as the target dimension, or if factuality or content-equivalence varies systematically with an intended level, the style-transfer manipulation is confounded and the Study 2 moderation estimates cannot be interpreted as style effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Study 2's distinctive claim, that individual traits moderate the effect of specific language styles on preference, depends on the zero-shot style-transfer prompts in Appendix B.3 changing only the intended style dimension, and on the style measures in Appendix A.1 isolating that dimension. No manipulation check, annotator agreement, or correlation between intended and measured levels is reported. The prompts themselves suggest off-target changes: Persuasiveness L3 adds strong emotional appeal or reasoning, Richness L3 adds excessive details, tangents, or background information, and Friendliness L3 shifts to an informal register. These plausibly alter content, accuracy, or multiple style dimensions at once. Since the same GPT-4o-Mini pipeline is used for both generation and measurement, coefficients attributed to a single style can absorb content differences. The inferential base is also fragile: 162 valid preference samples from roughly 10 users, no cluster-robust standard errors, no multiple-comparison correction, and an unexplained drop from 600 planned samples. If the style variables are confounded, the trait-moderation pattern in Fig. 1 (right) is not evidence about language style at all.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two studies on whether LLM language style affects user preference in open-ended interaction and whether user traits moderate this effect. Study 1 fits logistic regressions of binary preference on nine measured style features across three existing preference datasets (ArenaPref, ChatbotArena, MultiPref). Study 2 recruits 10 UK-based Prolific users, measures Big-Five traits and trust toward LLMs, and uses a 'Gibbs sampling with people' procedure in which users iteratively manipulate style intensities of GPT-4o-Mini responses; the resulting 162 valid preference samples are analyzed with moderated logistic regression. The authors conclude that LLM language style does influence user preference, that the influential styles vary across user populations, and that individual traits moderate these effects. The paper is explicitly framed as a preliminary study with acknowledged limitations in sample size and demographic diversity.","tokens_in":10433,"tokens_out":7648,"duration_ms":66452,"significance":"If the results held, the paper would provide a useful empirical mapping between specific stylistic dimensions and user preferences, with implications for personalization and for risks of misinformation. The strengths include the use of multiple real interaction datasets, an experimental design that elicits preferences rather than relying only on retrospective ratings, and unusually explicit caveats about sample limitations. However, the current analysis does not yet establish the specific style-preference relations claimed, because the style measures and stimuli are not validated, the reported Study 1 findings are internally inconsistent, and the Study 2 inference is vulnerable to clustering, filtering, and multiple-testing issues. The central claim is plausible and worth investigating, but the evidence as presented is not yet sufficient to support the specific conclusions.","major_comments":[{"comment":"The prose findings in Section 2.1 do not match Table 3. For ArenaPref, the text names Richness, Complexity, and Friendliness as significant, while Table 3 reports Richness (0.680**), Figurativeness (0.581**), and Presentation (0.160*) as the significant coefficients, with Complexity at 0.117 and Friendliness at -0.080. For ChatbotArena, the text names Richness, Presentation, and Figurativeness, but Table 3 shows Richness, Complexity (0.269*), and Friendliness (0.289*) as the significant coefficients. Since RQ.1 is answered from these numbers, the text or the table must be corrected and the discrepancy explained.","section":"Sec. 2.1 and Table 3"},{"comment":"The style measures and style-transfer stimuli are not validated, and the transfer prompts appear to change more than the target dimension. For example, Persuasiveness L3 adds 'strong emotional appeal or reasoning', Richness L3 adds 'excessive details, tangents, or background information', and Friendliness L3 changes to an informal register, each plausibly altering content, accuracy, or multiple style dimensions simultaneously. Because the same GPT-4o-Mini model is used for both generation and classification, a regression coefficient on a single style can absorb these off-target differences. The paper should report a manipulation check (e.g., human or model ratings of each intended dimension and a measure of content or accuracy preservation) before attributing preference shifts to individual styles.","section":"Appendices A.1 and B.3"},{"comment":"Study 2's inferential base is too fragile for the moderation claims. Ten participants were recruited and one is listed as rejected, but the final 162 samples are not explained relative to the planned 60 samples per participant (600 total); no filtering criteria are given. The observations are repeated within users, yet no cluster-robust standard errors or mixed-effects model is used. The moderated regression with 9 styles and 9 interactions already has many parameters for 162 samples, and testing 6 traits by 9 styles without multiple-comparison correction invites false positives. The authors should clarify the number of participants analyzed, describe the filtering, and rerun the key results with cluster-robust inference and corrected significance thresholds.","section":"Sec. 3 and Table 4"},{"comment":"The moderated regression equation is under-specified: y = logit(β0 + Σβ_i x_i + Σβ'_j x_i z_k) does not show how the style interactions are paired with the trait, whether the trait main effect z_k is included, or whether traits are centered. The definition of the reported 'odds shift' 1 - exp(β_i + β'_j z_k) is also not stated consistently with standard odds ratios. Without this information the moderation coefficients and Figure 3 cannot be reproduced, so the central moderation claim is not yet verifiable.","section":"Sec. 3, moderated regression equation"},{"comment":"Table 3 contains a row labeled 'Our Experiment (Study 2)' reporting main-effect coefficients, but Section 3.1 reports only moderation results and never interprets this row. If these are main-effect estimates from Study 2, they need to be described and placed in relation to Study 1; if not, the row should be removed.","section":"Table 3, 'Our Experiment (Study 2)' row"}],"minor_comments":[{"comment":"The formula '1 - exp(β_i)' likely should be 'exp(β_i)' or '1 - exp(-β_i)' to represent a positive odds increase; as written, positive coefficients would yield negative percentages.","section":"Sec. 2.1"},{"comment":"The table lists 'Num. of Participants 10' and 'Num. of Rejected Participants 1' but does not state how many participants contributed to n = 162, nor why 60 planned samples per participant reduced to roughly 18 valid samples each on average.","section":"Appendix B.2, Table 4"},{"comment":"The significance legend is garbled ('** : p < 0.01 ** : p < 0.05 +*: p < 0.10' with duplicate star symbols), and the relationship between the left and right panels and Studies 1 and 2 is not clear from the caption.","section":"Fig. 1"},{"comment":"There is a typo ('two paragprahs'), and the zero-shot style transfer prompts should also specify sampling parameters such as temperature, max tokens, and number of candidate responses for reproducibility.","section":"Appendix B.3"},{"comment":"The operationalization of 'open-ended interaction' using interrogative prefixes and exclusion of math/code/computation keywords is coarse; the paper should acknowledge that this filter may not cleanly isolate open-ended scenarios.","section":"Sec. 2, data selection"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its exploratory nature and the topic is timely. My main concern beyond the technical issues is the internal mismatch between the prose in Section 2.1 and Table 3; this must be resolved before the paper can be considered. The unexplained filtering in Study 2 and the absence of multiple-comparison control also need to be addressed, either through reanalysis or through much more cautious framing. If these points are fixed, the paper could be publishable as a preliminary study, but as it stands the specific claims are not yet supported by the reported evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a short, honest preliminary paper. The new piece—that style-preference effects hinge on user traits—is worth a serious look, but the evidence is thinner than the abstract implies. And there is a direct mismatch between the prose in Section 2.1 and Table 3 that you will want to resolve before citing anything.\n\nWhat's good. Study 1 is a clean re-run of the known style-over-substance regressions on three public preference datasets, and it reproduces population differences, which is a useful sanity check. Study 2's use of Gibbs sampling with people is a sensible way to elicit preferences with limited participants, and the paper is transparent about its small, homogeneous, UK-based sample. The polarizing moderation effects for extraversion and neuroticism, if real, are genuinely interesting.\n\nNow the soft spots, in decreasing order of severity.\n\nFirst, the mismatch. In Section 2.1, the ArenaPref significant styles are listed as Richness, Complexity, Friendliness, and ChatbotArena as Richness, Presentation, Figurativeness. Table 3 shows ArenaPref significant Richness, Presentation, Figurativeness, and ChatbotArena Richness, Complexity, Friendliness. That is a complete swap of two rows. Someone has to reconcile that before any result can be trusted.\n\nSecond, the style manipulation. The prompts in Appendix B.3 change more than the intended axis. Richness L3 adds tangents, Persuasiveness L3 adds strong emotional appeal, Friendliness L3 shifts register. No manipulation check or annotator agreement is reported, and the same model family generates and measures the style. So the moderator coefficients attributed to 'figurativeness' or 'authoritativeness' may be absorbing content and length differences. The stress-test worry is on target.\n\nThird, the inferential base: 162 valid samples out of a planned 600, with the filtering criteria unexplained; no multiple-comparison correction across dozens of tests; and no cluster-robust errors for the within-subject design. Some of the significant interactions will be false positives.\n\nNone of this is fatal. The authors are transparent about limitations, and the pilot design is a reasonable starting point. The citation pattern is fine.\n\nRecommendation: this deserves a serious referee, not a desk reject. It is a preliminary result for the HCI/LLM evaluation community. If the authors fix the table/prose mismatch, report the filtering, add a manipulation check, and use proper multiple-comparison corrections, it becomes a useful short paper. I would not cite it as established evidence in the meantime.","headline":"A small, honest preliminary study whose novel trait-moderation result is undermined by likely content confounds in the style manipulations and by a prose/table mismatch in Study 1.","tokens_in":10928,"tokens_out":4242,"would_cite":false,"duration_ms":38183,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language style shifts LLM preferences, and user traits decide how","keywords":["user preference","language style","personality traits","LLM interaction","preference alignment","moderation analysis","personalization","human-AI interaction"],"falsifier":"Have independent human raters score the nine style features on a random sample of the exact responses used in both studies. If the automatic measures disagree with the human ratings, or if raters cannot tell which intensity level a style-transfer prompt intended, then re-estimating the regressions with validated human-rated styles would be the test of whether the trait-dependent style effects survive.","tokens_in":10057,"feed_emoji":"💬","tokens_out":8131,"duration_ms":64544,"temperature":0.7,"pith_summary":"This preliminary study asks why users prefer one LLM response over another in open-ended conversation. The authors argue that information accuracy is not the only driver: the LLM's language style—richness, presentation, complexity, figurativeness, friendliness, interactiveness, authoritativeness, persuasiveness, and active voice—systematically shifts user preference, but the direction of each style effect depends on who the user is. In three existing preference datasets, at least three styles are significant per population, with different styles mattering in different datasets. An experiment with ten UK users then shows that Big Five personality traits and trust in LLMs moderate these style effects, implying that trait-aware personalization might be possible, though the small, homogeneous sample limits the strength of the conclusion.","feed_headline":"Language style shifts LLM preferences, and user traits decide how","feed_subtitle":"Two studies show the same style can help or hurt depending on the user's Big Five personality and trust in LLMs.","key_machinery":"The carrying mechanism is a nine-feature style measurement pipeline combined with logistic preference regression. Style features are scored by rule-based NLP counts (e.g., part-of-speech frequencies for richness, markdown styling patterns for presentation, readability scores for complexity), by zero-shot classification prompting a large language model for figurativeness, friendliness, interactiveness, and persuasiveness, and by a neural classifier for authoritativeness. In Study 1, the difference in each style feature between two candidate responses enters a logit model of the user's binary preference; in Study 2, trait-by-style interaction terms are added, and preferences are elicited through Gibbs sampling with people, where users iteratively adjust style sliders until the response matches their taste.","core_discovery":"The paper's central claim is that an LLM's language style influences user preference, but the influence is population-specific and moderated by individual traits. Study 1 uses binary preference regression on the ArenaPref, MultiPref, and ChatbotArena datasets, finding, for example, that richer responses raise preference odds by 88.6% in ArenaPref and 68.3% in ChatbotArena, while MultiPref users prefer presentation, complexity, interactiveness, and persuasiveness but are less likely to prefer authoritativeness. Study 2's experiment shows trait-dependent reversals: for users high in neuroticism, figurativeness and active voice increase preference, while for users high in extraversion they decrease it; agreeableness, openness, and trust also shift which styles matter. The authors read this as evidence that style effects are not monolithic and that the user's own traits are part of the mechanism.","pith_inferences":["A natural next test is whether users' preferred styles track their own writing style or their perception of the model's social role; the paper does not identify the psychological mechanism behind the moderation.","The polarizing effects observed for extraversion and neuroticism suggest that a joint moderation model over all five traits might reveal non-additive preferences, which the independent per-trait analyses cannot capture.","If the style measures were replaced with validated, content-matched manipulations, the same design could separate stylistic persuasion from content changes, directly informing the misinformation risk the authors flag."],"forward_implications":["Different user populations respond to different styles: richness drives preference in ArenaPref and ChatbotArena, while MultiPref users favor presentation, complexity, interactiveness, and persuasiveness instead.","Individual traits can reverse a style's effect: figurativeness and active voice help for high-neuroticism users but hurt for high-extraversion users.","Trait-aware style personalization is in principle feasible: knowing a user's Big Five profile and trust level could predict which styles to emphasize.","Because style effects vary by population, preference-alignment pipelines that ignore user traits may systematically overfit the majority population's stylistic taste.","The same styles that raise preference could also raise acceptance of hallucinated or misinformed content, since persuasion and style are intertwined in the measured features."],"supporting_citations":[{"why":"Supplies the ArenaPref dataset of real user-LLM preference pairs used in Study 1's regression.","marker":"[2]"},{"why":"Provides the MultiPref dataset, another preference-alignment corpus analyzed in Study 1.","marker":"[19]"},{"why":"Provides the ChatbotArena preference dataset used as the third Study 1 population.","marker":"[29]"},{"why":"Supplies the 'Gibbs sampling with people' method used to elicit each user's preference-maximizing style.","marker":"[9]"},{"why":"Supplies the zero-shot style transfer recipe used to generate style-varying LLM responses for Study 2 stimuli.","marker":"[23]"},{"why":"Provides the 10-item Big-Five personality measure used to assess user traits.","marker":"[7]"},{"why":"Provides the ChatGPT Trust Scale, adapted to measure general trust toward LLMs.","marker":"[3]"},{"why":"Provides the pairwise choice regression method underlying the binary preference models.","marker":"[21]"}],"fun_headline_variants":["LLM style sways users, but personality flips the effect","User traits decide if LLM style helps or hurts preference","Same LLM tone, different reactions: it's your traits","Language style impacts LLM preference, but not uniformly","Personality moderates how language style shapes LLM liking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument rests on the assumption that the automatic style measurements and the style-transfer prompts isolate each language style on its own, leaving response content and accuracy unchanged; if that assumption fails, the regression coefficients cannot be attributed to specific styles.","fun_headline_variants_meta":{"raw":{"variants":["LLM style sways users, but personality flips the effect","User traits decide if LLM style helps or hurts preference","Same LLM tone, different reactions: it's your traits","Language style impacts LLM preference, but not uniformly","Personality moderates how language style shapes LLM liking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1329,"prompt_tokens":987,"completion_tokens":342,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":259}},"tokens_in":603,"tokens_out":342,"duration_ms":3499,"temperature":1.0,"reasoning_tokens":259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:50:14.815370+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent human raters score the nine style features on a random sample of the exact responses used in both studies. If the automatic measures disagree with the human ratings, or if raters cannot tell which intensity level a style-transfer prompt intended, then re-estimating the regressions with validated human-rated styles would be the test of whether the trait-dependent style effects survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 'Gibbs sampling with people' method used to elicit each user's preference-maximizing style."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the 10-item Big-Five personality measure used to assess user traits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ChatGPT Trust Scale, adapted to measure general trust toward LLMs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pairwise choice regression method underlying the binary preference models."}],"review_version":1}