{"id":"d7368442-eea4-4f40-a446-03989162416a","arxiv_id":"2608.11164","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In conversational information seeking, assistant personality effects are task-dependent, and trust depends on both task style fit and user-assistant personality compatibility.","lead":"A controlled study with 26 people using three LLM assistants found that the best assistant personality depends on the task: extraverted for travel planning, neutral for shopping, and conscientious for health advice. The result suggests LLM-powered search assistants should adapt their style to context rather than use a single global persona.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task ecology is perfectly confounded with task content: only one scenario per ecology, so the headline RQ3 interaction cannot be attributed to exploratory/comparative/verification-sensitive demands.","rationale":"I read the paper in good faith: it is a careful within-subject design with counterbalancing, a fixed base model, honest reporting of null results, and an explicit limitations section. The behavioral distinctness of the three prompts is well supported, and the exit-questionnaire preferences for style choice/adaptation are plausible. My concern is not about data fabrication or statistical sloppiness, but about the inference from 'task' to 'ecology'. The reader's weakest assumption points in the same direction, but the problem is sharper: even if the three tasks are perfect exemplars of their ecologies, one exemplar cannot support a general claim about ecologies, because any observed difference is equally consistent with domain-specific content. This is load-bearing because the paper's contribution—'context-sensitive design variable'—is exactly the generalization across task types. A second concern is that the single significant interaction (p=.037) is fragile, but correcting p-values would not solve the confound; adding more participants would not either. I therefore keep the conditional verdict: the claim is plausible and worth testing, but the ecological interpretation should not be accepted on the current design.","tokens_in":16935,"tokens_out":5365,"duration_ms":55119,"concrete_test":"Run a follow-up study with at least two task scenarios per ecology, e.g., exploratory: Turkey travel and home-renovation planning; comparative: smartphone and laptop purchase; verification-sensitive: detox/cleanse claim and investment-return claim. Keep the same three assistant prompts and counterbalancing. Fit a mixed-effects model with task scenario as a random effect nested in ecology and test the Ecology × Assistant interaction on trust/delegation. If the interaction survives with scenario as a random effect and the Table 6 ordering reproduces within each ecology pair, the ecological attribution is supported; if the ordering varies within ecology pairs or the interaction disappears, the original RQ3 result is task-content-specific rather than ecology-general.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that assistant personality is a context-sensitive design variable depends on RQ3's task-by-assistant interaction. In Table 2, each ecology is instantiated by exactly one task scenario: travel=exploratory, smartphone shopping=comparative, health/diet=verification-sensitive. This makes 'task ecology' inseparable from 'task domain/content'. The significant interaction F(2,46)=3.55, p=.037 could therefore reflect topic-specific factors—health-topic caution, shopping price precision, travel openness—rather than the abstract information-seeking ecology the paper generalizes to. The paper provides no independent validity check for the ecology labels (e.g., rater judgments or multiple task instances per ecology), and the only other quantitative support is a marginal overall composite interaction (p=.051). The assistant conditions also differ substantially in verbosity (Titan 146 words/turn vs Europa 90), so task-specific tolerance for verbosity could create the observed preference pattern without implying personality-ecology fit. This is a construct-validity/identifiability problem, not just a sample-size issue: no re-analysis of the existing single-scenario data can separate ecology from content.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a within-subject user study (N=26) examining how prompt-defined assistant personality (extraverted Titan, conscientious Europa, neutral Neptune) affects behavior and evaluations across three information-seeking tasks (travel planning, smartphone shopping, health/diet). It claims that: assistant styles produce behaviorally distinct interaction patterns; participant personality alone does not robustly predict behavioral traces; participant–assistant compatibility, operationalized as a trait-balance score, predicts trust and delegation; and a task-by-assistant interaction on trust and delegation shows different preferred styles per task. The authors conclude that assistant personality should be treated as a context-sensitive interaction design variable rather than a globally optimizable system property.","tokens_in":17034,"tokens_out":6870,"duration_ms":65270,"significance":"If the claims hold, the paper makes a useful empirical contribution to conversational IR and HCI by showing that assistant style effectiveness varies by task and by user–assistant fit, and by providing concrete behavioral trace evidence alongside questionnaire data. The design has real strengths: a fixed base model across conditions, pre-study prompt validation, counterbalanced within-subject assignment, multiple data sources, and a generally honest limitations section. The main limitation is that the headline ecology-level interpretation is confounded with task content, and the inferential base is thin, resting on a small number of uncorrected tests. The contribution is publishable as an exploratory study, but the scope of the conclusions should be narrowed or supported by additional evidence.","major_comments":[{"comment":"The central RQ3 claim that assistant personality effectiveness varies by information-seeking ecology is not identifiable from the current design, because each ecology is instantiated by exactly one task scenario. The significant Task × Assistant interaction (F(2,46)=3.55, p=.037) may reflect topic-specific factors such as health-topic caution, shopping price sensitivity, travel openness, or task-specific tolerance for the large verbosity differences between conditions (Titan 146 words/turn, Neptune 115, Europa 90). The paper provides no independent validity check of the ecology labels and no second task instance per ecology, so no re-analysis of the existing data can separate exploratory/comparative/verification-sensitive demands from travel/shopping/health content. This is load-bearing for RQ3 and for the conclusion that assistant personality is a context-sensitive design variable. The claims should be reframed as task-specific, or the design should be extended with multiple scenarios per ecology and a manipulation check of perceived task demands.","section":"Section 5.4, Table 2"},{"comment":"The headline interaction is a single uncorrected test. The RQ1 analyses are described as Holm-corrected, but no correction is reported for the RQ3 interaction or for the RQ2 compatibility slope. With at least two primary questionnaire constructs (trust/delegation and the overall composite) plus the behavioral traces, p=.037 would not survive a Holm correction across even two outcomes (threshold .025), and the composite interaction is already marginal at p=.051. Given the paper's own stated commitment to conservative interpretation, the RQ3 interaction should be reported with its correction status and described as suggestive rather than confirmatory. The phrase 'the strongest supported effect' overstates the evidence.","section":"Section 5.4"},{"comment":"The RQ2 compatibility slope (b=0.648, F(1,22)=6.38, p=.019, partial eta-squared=.225) is based on a researcher-defined trait-balance score, z(Extraversion)-z(Conscientiousness), and a small regression. Because this contrast is defined after the two trait-linked conditions are known and is not accompanied by robustness checks (e.g., alternative operationalizations, raw-score versions, or checks for influential observations), it should be interpreted as exploratory. The conclusion that compatibility matters for trust and delegation requires replication in a larger sample before design implications are drawn.","section":"Section 5.3"},{"comment":"The behavioral distinctness of the assistant conditions is entangled with verbosity and interaction pacing. Titan produced 146 words per turn versus Europa's 90, and Europa produced more turns and a higher user word share. Because the tasks differ in how much verbosity is tolerable, the observed preference pattern (Titan best in travel, Neptune best in shopping, Europa best in health) could in part be a task-specific response to verbosity or pacing rather than to extraversion or conscientiousness as such. The paper acknowledges the trace differences descriptively but does not model response length as a covariate or otherwise disentangle the personality construct from its surface realization.","section":"Section 5.1 and 5.4"}],"minor_comments":[{"comment":"The sentence 'However, effect of user personality in such interactions are not explicitly studied yet' has a subject-verb agreement error and should be revised to 'effects ... have not been explicitly studied.'","section":"Section 2.1"},{"comment":"The sentence 'The prompt validation profiles in Figure 2 indicate that the trait-linked prompts produced visibly different profiles; the live trace data further showed differences...' repeats the preceding sentence almost verbatim; one of the two instances should be removed.","section":"Section 4.2"},{"comment":"The table header 'Name shown Condition' is awkward; consider splitting it into separate columns or rewording it as 'Name shown to participants' and 'Intended interaction style.'","section":"Table 1"},{"comment":"The median-split preference figures are dense and the addition of a 'Tie' category in Figure 6 makes the percentages hard to read; a note clarifying that these are descriptive triangulation, as the text states, would help readers interpret the bars.","section":"Figures 5 and 6"}],"recommendation":"major_revision","confidential_remarks":"This is an honest, well-structured exploratory study with a genuinely useful empirical contribution. My main concern is the ecological-level interpretation of the task-by-assistant interaction, which is confounded with task content; I agree with the stress-test concern that the existing single-scenario data cannot separate the two. I recommend major revision rather than rejection: the authors should either narrow the claims to the specific task domains or add follow-up evidence such as multiple scenarios per ecology and manipulation checks. The paper is within scope for an IR/HCI venue and should not be dismissed on sample size alone, but the central claim needs to be more carefully delimited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is a small but well-run user study asking whether assistant personality, user personality, and task type jointly shape conversational information seeking. The headline finding is that no single assistant style wins: a task-by-assistant interaction on trust/delegation (p=.037) with Titan best for exploratory travel, Neptune for comparative shopping, Europa for verification-sensitive health. The companion compatibility slope (p=.019) suggests users trust an assistant whose expressed style matches their own E-C balance. If these hold, they are genuinely useful for CIS design—style is a context-sensitive variable, not a global default.\n\nWhat the paper does well: the design is careful. Counterbalancing via Graeco-Latin square, blind planet names, delayed personality reveal, task instructions kept out of prompts, a fixed base model varying only the system prompt. The authors are honest about their own limits: they label exploratory effects as directional, apply Holm correction in some places, and list limitations including the convenience sample and the absence of factual verification. The null result for RQ1 (personality not a direct predictor of behavior) is reported cleanly and is itself informative.\n\nThe soft spots are real but not fatal. The most serious is the task-ecology confound: each information-seeking ecology is instantiated by exactly one task scenario, so 'exploratory' is just 'travel', 'comparative' is just 'phone shopping', and 'verification-sensitive' is just 'health'. You cannot separate ecology from content with this design; the significant interaction might reflect topic-specific caution or pricing pressure rather than abstract ecological demands. Also, the two stylized assistants differ wildly in verbosity (146 vs 90 words/turn), so the effect could be partly verbosity, not personality. The headline interaction is a single test at p=.037 without correction; the composite interaction is only p=.051. The compatibility slope comes from 26 participants and one researcher-defined index. No prompts or analysis scripts are provided. None of these are disqualifying for an exploratory study, but they point to needed replications with multiple tasks per ecology and corrected reporting.\n\nFor whom: anyone working on personality-aware conversational search or assistant personalization will want to read this; it gives concrete design hypotheses and a sensible experimental template. It deserves serious peer review, with revisions that address the confounding and reporting issues. I'd rather see it in the record with corrections than desk-rejected.\n\nBest,","headline":"A careful small-N study yielding a plausible but statistically thin context-dependence claim; the task-ecology confound is the main reason for caution, but the paper deserves peer review.","tokens_in":17644,"tokens_out":2754,"would_cite":true,"duration_ms":23972,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No single assistant personality is optimal; task context shapes which style users trust.","keywords":["conversational information seeking","assistant personality","Big Five personality","trust and delegation","task-style fit","user study","LLM prompting","personalization"],"falsifier":"Run a replication in which the same task materials are presented under different task-ecology framings, for example the identical smartphone scenario labelled as explore-options versus decide-quickly, and test whether the task-by-assistant interaction on trust and delegation follows the label or the material; if it follows the label, or disappears when task difficulty and domain familiarity are entered as covariates, the paper's task-style-fit interpretation is weakened.","tokens_in":16616,"feed_emoji":"🤖","tokens_out":9681,"duration_ms":80180,"temperature":0.7,"pith_summary":"This paper tries to establish that the personality an LLM assistant expresses is a context-sensitive design variable rather than a globally optimal setting. In a within-subject study, twenty-six participants each completed three information-seeking tasks with three prompt-defined styles of the same base assistant: extraverted, conscientious, and neutral. The clearest result was a task-by-assistant interaction on trust and delegation, with different styles winning in different task ecologies: extraverted in exploratory travel planning, neutral in comparative smartphone shopping, and conscientious in verification-sensitive health information. Participant personality on its own did not reliably predict conversational behaviour, but compatibility between the participant's extraversion-conscientiousness profile and the assistant's style was associated with trust and delegation. The authors conclude that style choice or adaptation, rather than a fixed persona, is the more promising direction for conversational search design.","feed_headline":"Best chatbot style depends on the task, not the model","feed_subtitle":"In controlled trials, the extraverted, conscientious, and neutral assistants each earned the most trust in a different task.","key_machinery":"The central object is the expressed assistant personality, instantiated by prompting a single base LLM into three styles: Titan (extraverted), Europa (conscientious), and Neptune (neutral baseline). The second object is the task ecology, operationalised by three scenarios classified as exploratory travel planning, comparative smartphone shopping, and verification-sensitive health advice. The key measurement that carries the argument is the trust and delegation construct, which showed the supported Task-by-assistant interaction, together with a compatibility score defined as $z(\\text{Extraversion}) - z(\\text{Conscientiousness})$ from Big Five measurements. Prompt validation and behavioural traces (assistant response length, user word share, and turn count) serve as checks that the three styles were behaviourally distinct in the live interactions.","core_discovery":"The central claim is that assistant personality is a situational interaction parameter: the same expressed style can be trusted in one task and rejected in another, so there is no universally best persona. The strongest supported effect was a significant Task-by-assistant interaction on trust and delegation ($F(2,46)=3.55$, $p=.037$, partial $\\eta^2=.134$), with Titan (extraverted) rated highest in travel, Neptune (neutral) in shopping, and Europa (conscientious) in health. The same ordering appeared descriptively for the overall post-interaction composite but with weaker statistical support ($p=.051$). Separately, the compatibility slope between the participant's trait balance and assistant style was significant for trust and delegation ($b=0.648$, $F(1,22)=6.38$, $p=.019$, partial $\\eta^2=.225$), while broad Big Five traits alone did not reliably predict surface behaviour such as turn count, word share, or question frequency. The paper therefore frames assistant personality as a task- and user-sensitive interactional design variable rather than a globally optimisable system property.","pith_inferences":["Editorial extension: if the task-by-assistant interaction replicates, an assistant that infers or asks about the task type and adjusts its expressed style could raise trust without retraining the model; the paper's exit-questionnaire results on adaptation preference are consistent with this but do not test it.","Editorial extension: the null result for direct personality effects suggests that self-reported Big Five traits are weak proxies for observable conversational behaviour, so future work might measure task-specific interaction preferences or implicit behavioural signatures instead.","Editorial extension: the three task scenarios differ in stakes, difficulty, and participant familiarity as well as ecology, so a replication that manipulates those dimensions separately would clarify whether the interaction is genuinely about task-style fit."],"forward_implications":["Assistant personality should be evaluated per task rather than once per system, because the same style can be top-rated in one ecology and bottom-rated in another.","Trust and delegation are the outcomes most sensitive to style fit, so they should be measured separately from general liking in conversational search evaluations.","Prompt-level style differences measurably change interaction structure, such as turn count, user word share, and response length, even when the underlying model and task materials are fixed.","Participants rated both explicit style choice and automatic style adaptation well above the midpoint, and the personality reveal did not substantially change preferences, supporting configurable or adaptive personas.","Participant-assistant compatibility acts as a trust-related moderator rather than a broad improvement to interaction quality."],"supporting_citations":[{"why":"Defines the conversational search action space that motivates treating assistant style as part of the interaction interface.","marker":"[37]"},{"why":"Establishes conversational information seeking as a turn-based sequence, the framing behind the behavioural trace measures.","marker":"[50]"},{"why":"Shows that interaction style can be measured and that style mismatch affects perceived effort, the basis for the compatibility analysis.","marker":"[44]"},{"why":"Supplies evidence that users align to an agent's expressed style, which the task-fit findings extend.","marker":"[45]"},{"why":"Distinguishes exploratory search from lookup and comparison, supporting the three task ecologies.","marker":"[34]"},{"why":"Documents that health information seeking weights trustworthiness and evidence, motivating the verification-sensitive task.","marker":"[43]"},{"why":"Provides the within-subject evaluation and counterbalancing method used to avoid task-assistant confounds.","marker":"[29]"},{"why":"Supplies the Big Five trait taxonomy used to measure participant personality and frame the two trait-linked assistant styles.","marker":"[28]"}],"fun_headline_variants":["No best chatbot persona: task determines trust","Chatbot style: trust varies by task, not model","Personality in AI: one size doesn't fit all tasks","Task, not model, decides best chatbot personality","Trust in chatbots: it's the task that matters"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three hand-written scenarios are valid representatives of three distinct information-seeking ecologies, exploratory, comparative, and verification-sensitive, so that the task-by-assistant trust interaction reflects style fit rather than task difficulty, domain familiarity, or another correlated property.","fun_headline_variants_meta":{"raw":{"variants":["No best chatbot persona: task determines trust","Chatbot style: trust varies by task, not model","Personality in AI: one size doesn't fit all tasks","Task, not model, decides best chatbot personality","Trust in chatbots: it's the task that matters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1586,"prompt_tokens":1032,"completion_tokens":554,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":478}},"tokens_in":648,"tokens_out":554,"duration_ms":4900,"temperature":1.0,"reasoning_tokens":478,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:54:04.481900+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a replication in which the same task materials are presented under different task-ecology framings, for example the identical smartphone scenario labelled as explore-options versus decide-quickly, and test whether the task-by-assistant interaction on trust and delegation follows the label or the material; if it follows the label, or disappears when task difficulty and domain familiarity are entered as covariates, the paper's task-style-fit interpretation is weakened.","supporting_citations":[{"cited_title":"Trippas, Jeff Dalton, and Filip Radlinski","cited_arxiv_id":null,"evidence_quote":"Establishes conversational information seeking as a turn-based sequence, the framing behind the behavioural trace measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that interaction style can be measured and that style mismatch affects perceived effort, the basis for the compatibility analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Distinguishes exploratory search from lookup and comparison, supporting the three task ecologies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents that health information seeking weights trustworthiness and evidence, motivating the verification-sensitive task."},{"cited_title":"John and Sanjay Srivastava","cited_arxiv_id":null,"evidence_quote":"Supplies the Big Five trait taxonomy used to measure participant personality and frame the two trait-linked assistant styles."}],"review_version":1}