Pith. sign in

REVIEW 4 major objections 4 minor 56 references

Role of Personality in Conversational Information Seeking

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read No single assistant personality is optimal; task context shapes which style users trust.

desk verdict A careful small-N study yielding a plausible but statistically thin context-dependence claim; the task-ecology confound is the main reason for caution, but the paper deserves peer review. read the letter →

arxiv 2608.11164 v1 pith:6VOPQKUQ submitted 2026-08-11 cs.IR

classification cs.IR
keywords conversationalinformationseekingassistantpersonalityBigFivetrustanddelegationtask-stylefituserstudyLLMpromptingpersonalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the personality an LLM assistant expresses is a context-sensitive design variable rather than a globally optimal setting. In a within-subject study, twenty-six participants each completed three information-seeking tasks with three prompt-defined styles of the same base assistant: extraverted, conscientious, and neutral. The clearest result was a task-by-assistant interaction on trust and delegation, with different styles winning in different task ecologies: extraverted in exploratory travel planning, neutral in comparative smartphone shopping, and conscientious in verification-sensitive health information. Participant personality on its own did not reliably predict conversational behaviour, but compatibility between the participant's extraversion-conscientiousness profile and the assistant's style was associated with trust and delegation. The authors conclude that style choice or adaptation, rather than a fixed persona, is the more promising direction for conversational search design.

What carries the argument

The central object is the expressed assistant personality, instantiated by prompting a single base LLM into three styles: Titan (extraverted), Europa (conscientious), and Neptune (neutral baseline). The second object is the task ecology, operationalised by three scenarios classified as exploratory travel planning, comparative smartphone shopping, and verification-sensitive health advice. The key measurement that carries the argument is the trust and delegation construct, which showed the supported Task-by-assistant interaction, together with a compatibility score defined as $z(\text{Extraversion}) - z(\text{Conscientiousness})$ from Big Five measurements. Prompt validation and behavioural traces (assistant response length, user word share, and turn count) serve as checks that the three styles were behaviourally distinct in the live interactions.

What would settle it

Run a replication in which the same task materials are presented under different task-ecology framings, for example the identical smartphone scenario labelled as explore-options versus decide-quickly, and test whether the task-by-assistant interaction on trust and delegation follows the label or the material; if it follows the label, or disappears when task difficulty and domain familiarity are entered as covariates, the paper's task-style-fit interpretation is weakened.

Watch

Extended reading notes

Core claim

The central claim is that assistant personality is a situational interaction parameter: the same expressed style can be trusted in one task and rejected in another, so there is no universally best persona. The strongest supported effect was a significant Task-by-assistant interaction on trust and delegation ($F(2,46)=3.55$, $p=.037$, partial $\eta^2=.134$), with Titan (extraverted) rated highest in travel, Neptune (neutral) in shopping, and Europa (conscientious) in health. The same ordering appeared descriptively for the overall post-interaction composite but with weaker statistical support ($p=.051$). Separately, the compatibility slope between the participant's trait balance and assistant style was significant for trust and delegation ($b=0.648$, $F(1,22)=6.38$, $p=.019$, partial $\eta^2=.225$), while broad Big Five traits alone did not reliably predict surface behaviour such as turn count, word share, or question frequency. The paper therefore frames assistant personality as a task- and user-sensitive interactional design variable rather than a globally optimisable system property.

Load-bearing premise

The load-bearing premise is that the three hand-written scenarios are valid representatives of three distinct information-seeking ecologies, exploratory, comparative, and verification-sensitive, so that the task-by-assistant trust interaction reflects style fit rather than task difficulty, domain familiarity, or another correlated property.

Editorial extensions

If this is right

  • Assistant personality should be evaluated per task rather than once per system, because the same style can be top-rated in one ecology and bottom-rated in another.
  • Trust and delegation are the outcomes most sensitive to style fit, so they should be measured separately from general liking in conversational search evaluations.
  • Prompt-level style differences measurably change interaction structure, such as turn count, user word share, and response length, even when the underlying model and task materials are fixed.
  • Participants rated both explicit style choice and automatic style adaptation well above the midpoint, and the personality reveal did not substantially change preferences, supporting configurable or adaptive personas.
  • Participant-assistant compatibility acts as a trust-related moderator rather than a broad improvement to interaction quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: if the task-by-assistant interaction replicates, an assistant that infers or asks about the task type and adjusts its expressed style could raise trust without retraining the model; the paper's exit-questionnaire results on adaptation preference are consistent with this but do not test it.
  • Editorial extension: the null result for direct personality effects suggests that self-reported Big Five traits are weak proxies for observable conversational behaviour, so future work might measure task-specific interaction preferences or implicit behavioural signatures instead.
  • Editorial extension: the three task scenarios differ in stakes, difficulty, and participant familiarity as well as ecology, so a replication that manipulates those dimensions separately would clarify whether the interaction is genuinely about task-style fit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper reports a within-subject user study (N=26) examining how prompt-defined assistant personality (extraverted Titan, conscientious Europa, neutral Neptune) affects behavior and evaluations across three information-seeking tasks (travel planning, smartphone shopping, health/diet). It claims that: assistant styles produce behaviorally distinct interaction patterns; participant personality alone does not robustly predict behavioral traces; participant–assistant compatibility, operationalized as a trait-balance score, predicts trust and delegation; and a task-by-assistant interaction on trust and delegation shows different preferred styles per task. The authors conclude that assistant personality should be treated as a context-sensitive interaction design variable rather than a globally optimizable system property.

Significance. If the claims hold, the paper makes a useful empirical contribution to conversational IR and HCI by showing that assistant style effectiveness varies by task and by user–assistant fit, and by providing concrete behavioral trace evidence alongside questionnaire data. The design has real strengths: a fixed base model across conditions, pre-study prompt validation, counterbalanced within-subject assignment, multiple data sources, and a generally honest limitations section. The main limitation is that the headline ecology-level interpretation is confounded with task content, and the inferential base is thin, resting on a small number of uncorrected tests. The contribution is publishable as an exploratory study, but the scope of the conclusions should be narrowed or supported by additional evidence.

major comments (4)
  1. [Section 5.4, Table 2] The central RQ3 claim that assistant personality effectiveness varies by information-seeking ecology is not identifiable from the current design, because each ecology is instantiated by exactly one task scenario. The significant Task × Assistant interaction (F(2,46)=3.55, p=.037) may reflect topic-specific factors such as health-topic caution, shopping price sensitivity, travel openness, or task-specific tolerance for the large verbosity differences between conditions (Titan 146 words/turn, Neptune 115, Europa 90). The paper provides no independent validity check of the ecology labels and no second task instance per ecology, so no re-analysis of the existing data can separate exploratory/comparative/verification-sensitive demands from travel/shopping/health content. This is load-bearing for RQ3 and for the conclusion that assistant personality is a context-sensitive design variable. The claims should be reframed as task-specific, or the design should be extended with multiple scenarios per ecology and a manipulation check of perceived task demands.
  2. [Section 5.4] The headline interaction is a single uncorrected test. The RQ1 analyses are described as Holm-corrected, but no correction is reported for the RQ3 interaction or for the RQ2 compatibility slope. With at least two primary questionnaire constructs (trust/delegation and the overall composite) plus the behavioral traces, p=.037 would not survive a Holm correction across even two outcomes (threshold .025), and the composite interaction is already marginal at p=.051. Given the paper's own stated commitment to conservative interpretation, the RQ3 interaction should be reported with its correction status and described as suggestive rather than confirmatory. The phrase 'the strongest supported effect' overstates the evidence.
  3. [Section 5.3] The RQ2 compatibility slope (b=0.648, F(1,22)=6.38, p=.019, partial eta-squared=.225) is based on a researcher-defined trait-balance score, z(Extraversion)-z(Conscientiousness), and a small regression. Because this contrast is defined after the two trait-linked conditions are known and is not accompanied by robustness checks (e.g., alternative operationalizations, raw-score versions, or checks for influential observations), it should be interpreted as exploratory. The conclusion that compatibility matters for trust and delegation requires replication in a larger sample before design implications are drawn.
  4. [Section 5.1 and 5.4] The behavioral distinctness of the assistant conditions is entangled with verbosity and interaction pacing. Titan produced 146 words per turn versus Europa's 90, and Europa produced more turns and a higher user word share. Because the tasks differ in how much verbosity is tolerable, the observed preference pattern (Titan best in travel, Neptune best in shopping, Europa best in health) could in part be a task-specific response to verbosity or pacing rather than to extraversion or conscientiousness as such. The paper acknowledges the trace differences descriptively but does not model response length as a covariate or otherwise disentangle the personality construct from its surface realization.
minor comments (4)
  1. [Section 2.1] The sentence 'However, effect of user personality in such interactions are not explicitly studied yet' has a subject-verb agreement error and should be revised to 'effects ... have not been explicitly studied.'
  2. [Section 4.2] The sentence 'The prompt validation profiles in Figure 2 indicate that the trait-linked prompts produced visibly different profiles; the live trace data further showed differences...' repeats the preceding sentence almost verbatim; one of the two instances should be removed.
  3. [Table 1] The table header 'Name shown Condition' is awkward; consider splitting it into separate columns or rewording it as 'Name shown to participants' and 'Intended interaction style.'
  4. [Figures 5 and 6] The median-split preference figures are dense and the addition of a 'Tie' category in Figure 6 makes the percentages hard to read; a note clarifying that these are descriptive triangulation, as the text states, would help readers interpret the bars.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's central task-by-assistant interaction and compatibility results are empirical findings on measured outcomes, not consequences of the way inputs were defined.

full rationale

This is an empirical within-subject user study rather than a derivation chain. The central RQ3 claim—a Task × Assistant interaction on trust and delegation (F(2,46)=3.55, p=.037)—is an inferential result computed from measured post-interaction ratings, and the descriptive ordering (Titan strongest for Travel, Neptune for Shopping, Europa for Health) comes directly from the Table 6 means. The task-ecology labels (exploratory, comparative, verification-sensitive) are design assumptions, but they are not used to construct the outcome variable: trust/delegation is participant-reported and can in principle go against any task label, so the interaction is not forced by definition. Similarly, the RQ2 compatibility score z(Extraversion)−z(Conscientiousness) is a researcher-chosen predictor, but it is measured from BFI scores and regressed against the Titan−Europa trust difference; the reported slope (b=0.648, F(1,22)=6.38, p=.019) is an empirical estimate that could have been null or opposite, so it does not reduce to a fitted input renamed as a prediction. The manipulation validation is likewise empirical: prompt profiles and behavioural traces (response length, user word share, turn count) were used to confirm that the conditions differed, not to define the outcome. The paper does contain self-citations (e.g., refs [3], [12–17], [22], [36], [47], [49], [52], [53]), but these appear in related work and future directions and are not load-bearing for the empirical claims; no argument for the main findings reduces to a self-citation or to an imported uniqueness theorem. The construct-validity concern that each task ecology is instantiated by only one scenario is a genuine threat to generalization, but it is a correctness risk rather than circularity. Accordingly, no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests primarily on the a priori task-ecology mapping and the validity of prompt-based style manipulation. No new physical or conceptual entities are introduced; the compatibility score is a derived variable, not an invented entity. The statistical analysis assumes standard test assumptions that are not checked.

free parameters (1)
  • trait-balance compatibility score = z(Extraversion) - z(Conscientiousness)
    Constructed from participant BFI scores to align with the two trait-linked assistant conditions; used as the predictor in the RQ2 regression. It is a researcher-chosen contrast that could affect the reported compatibility slope.
assumptions (4)
  • domain assumption The Big Five framework is a valid way to measure participant personality.
    The study uses BFI scores as the measure of participant personality and constructs compatibility from two of the five traits; if Big Five does not capture relevant individual differences, the compatibility results are undermined (Section 3.4).
  • domain assumption System prompts can reliably produce distinct, stable assistant interaction styles in GPT-4.1.
    All assistant effects rest on the assumption that the Titan, Europa, and Neptune conditions differed only in expressed style, validated only by pre-study prompt inspection and descriptive traces (Sections 3.3, 4.2).
  • ad hoc to paper The travel, smartphone, and health tasks instantiate distinct information-seeking ecologies (exploratory, comparative, verification-sensitive).
    The task labels are assigned a priori in Table 2 and are not measured or validated during the study; RQ3 conclusions depend on this mapping.
  • standard math Within-subject ANOVA/regression assumptions (normality, sphericity, independence) hold for the F-tests.
    The paper does not report assumption checks for the F-tests in Sections 5.3 and 5.4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Role of Personality in Conversational Information Seeking." pith.science (2026). https://pith.science/paper/6VOPQKUQ

@misc{pith2026260811164,
  author       = {Pith},
  title        = {Pith review of: Role of Personality in Conversational Information Seeking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6VOPQKUQ}},
  note         = {Machine review of arXiv:2608.11164}
}
read the original abstract

Large language models (LLMs) are increasingly used for information seeking, where users find, compare, and evaluate information through dialogue. In this role, the assistant does more than retrieve or generate content: it shapes how users articulate constraints, ask follow-up questions, verify claims, and decide when an answer is sufficient for action. Yet little is known about how user personality, assistant personality, and task context jointly influence these interactions. We examine personality as a controllable variable in conversational information seeking and study its effects on user behaviour and interaction quality. We conducted a controlled within-subject study in which assistant personality and task type were experimentally varied, while participant personality was measured using Big Five scores. Twenty-six participants each completed three information-seeking tasks under three assistant personality conditions: extraverted, conscientious, and neutral. Tasks covered exploratory travel planning, comparative smartphone shopping, and verification-sensitive health and diet information seeking. Data included conversation logs, behavioural traces, post-interaction questionnaires, an exit questionnaire, and Big Five measures. The assistant conditions were behaviourally distinct: the extraverted assistant produced longer turns, the conscientious assistant elicited higher user word share and more turns, and the neutral baseline fell between them. The strongest effect was a task-by-assistant interaction on trust and delegation, with preferred styles varying by task. No global winner emerged, but participants strongly preferred style choice or adaptation. These findings position assistant personality as a context-sensitive interactional design variable rather than a globally optimisable system property.

Figures

Figures reproduced from arXiv: 2608.11164 by the authors.

Figure 1
Figure 1. Custom interface used during the live study. Panel A shows an example initial interaction with the conscientious [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Prompt validation profiles for the three assistant [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Spearman correlations between Big Five traits and [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Compatibility slope for the Titan–Europa dif [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Assistant preference distribution by median-split Big Five groups for the overall post-interaction composite. Bars [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Assistant preference distribution by median-split Big Five groups for trust and delegation. Bars show the percentage [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 30 canonical work pages

  1. [1]

    Hosam Al-Samarraie, Atef Eldenfria, and Husameddin Dawoud. 2017. The im- pact of personality traits on users’ information-seeking behavior.Information Processing & Management53, 1 (2017), 237–247. doi:10.1016/j.ipm.2016.08.004

  2. [2]

    Bruce Croft

    Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, and W. Bruce Croft

  3. [4]

    Hyeonjeong Byeon, Uran Oh, and Gary Hsieh. 2026. Understanding the Effects of Conversational Agent Personality on the Credibility of LLM-Based Conver- sational Search. InProceedings of the 2026 Conference on Human Information Interaction and Retrieval (CHIIR ’26). Association for Computing Machinery, New York, NY, USA, 222–230. doi:10.1145/3786304.3788843

  4. [5]

    Na Cai, Shuhong Gao, and Jinzhe Yan. 2024. How the communication style of chatbots influences consumers’ satisfaction, trust, and engagement in the context of service failure.Humanities and Social Sciences Communications11, 1 (2024), 1–11

  5. [6]

    Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman

    Aaron Chatterji, Thomas Cunningham, David J. Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025.How People Use ChatGPT. NBER Working Paper 34255. National Bureau of Economic Research, Cambridge, MA. https://www.nber.org/papers/w34255

  6. [7]

    Clark and Susan E

    Herbert H. Clark and Susan E. Brennan. 1991. Grounding in Communication. In Perspectives on Socially Shared Cognition, Lauren B. Resnick, John M. Levine, and Stephanie D. Teasley (Eds.). American Psychological Association, Washington, DC, 127–149. doi:10.1037/10096-006

  7. [8]

    Paul T Costa Jr and Robert R McCrae. 1995. Domains and facets: Hierarchical personality assessment using the Revised NEO Personality Inventory.Journal of personality assessment64, 1 (1995), 21–50

  8. [9]

    Trippas, and Hamed Zamani

    Jeffrey Dalton, Sophie Fischer, Paul Owoicho, Filip Radlinski, Federico Rossetto, Johanne R. Trippas, and Hamed Zamani. 2022. Conversational Information Seeking: Theory and Application. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval(Madrid, Spain)(SIGIR ’22). Association for Computing Mach...

Show all 56 references
  1. [10]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020. CAsT 2020: The Con- versational Assistance Track Overview. InText Retrieval Conference. https: //api.semanticscholar.org/CorpusID:214735659

  2. [11]

    Jeffrey Dalton, Chenyan Xiong, and Jamie Callan. 2020. TREC CAsT 2019: The Conversational Assistance Track Overview.CoRRabs/2003.13624 (2020). arXiv:2003.13624 https://arxiv.org/abs/2003.13624

  3. [12]

    Jose, and Zhaochun Ren

    Junchen Fu, Xuri Ge, Alexandros Karatzoglou, Ioannis Arapakis, Suzan Ver- berne, Joemon M. Jose, and Zhaochun Ren. 2026. Differentiable Semantic ID for Generative Recommendation. InProceedings of the 49th International ACM SIGIR Conference on Research and Development in Inform...

  4. [13]

    Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Jie Wang, and Joemon M Jose. 2024. IISAN: Efficiently adapting multimodal repre- sentation for sequential recommendation with decoupled PEFT. InProceedings of the 47th International ACM SIGIR Conference on...

  5. [14]

    Junchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Kaiwen Zheng, Yongxin Ni, and Joemon M Jose Joemon. 2025. Efficient and effective adaptation of multimodal foundation models in sequential recommendation.IEEE Transactions on Knowledge and Data Engineering(2025)

  6. [15]

    Junchen Fu, Yongxin Ni, Joemon M Jose, Ioannis Arapakis, Kaiwen Zheng, Youhua Li, and Xuri Ge. 2025. Crossan: Towards efficient and effective adaptation of multiple multimodal foundation models for sequential recommendation.arXiv preprint arXiv:2504.10307(2025)

  7. [16]

    Junchen Fu, Fajie Yuan, Yu Song, Zheng Yuan, Mingyue Cheng, Shenghui Cheng, Jiaqi Zhang, Jie Wang, and Yunzhu Pan. 2024. Exploring adapter-based transfer learning for recommender systems: Empirical studies and practical insights. In Proceedings of the 17th ACM international co...

  8. [17]

    Junchen Fu, Kaiwen Zheng, Ioannis Arapakis, Wenhao Deng, Xin Xin, Joemon M Jose, and Xuri Ge. 2026. Stream-aware Side Adaptation for Large Pre-trained Multimodal Embedding Models in Sequential Recommendation.arXiv preprint arXiv:2607.10909(2026)

  9. [18]

    Lucie Galland, Catherine Pelachaud, and Florian Pecune. 2022. Adapting conversa- tional strategies to co-optimize agent’s task performance and user’s engagement. InProceedings of the 22nd ACM International Conference on Intelligent Virtual Agents(Faro, Portugal)(IV A ’22). Ass...

  10. [19]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Gen- eration for Large Language Models: A Survey.arXiv preprint arXiv:2312.10997 (2023). https://arxiv.org/abs/2312.10997

  11. [20]

    Description of Personality

    Lewis R. Goldberg. 1990. An Alternative “Description of Personality”: The Big-Five Factor Structure.Journal of Personality and Social Psychology59, 6 (1990), 1216–1229. https://projects.ori.org/lrg/pdfs_papers/goldberg.big-five- factorsstructure.jpsp.1990.pdf

  12. [21]

    Bin Han, Deuksin Kwon, and Jonathan Gratch. 2026. Personality Expres- sion Across Contexts: Linguistic and Behavioral Variation in LLM Agents. arXiv:2602.01063 [cs.CL] https://arxiv.org/abs/2602.01063

  13. [22]

    Yaoqin He, Junchen Fu, Kaiwen Zheng, Songpei Xu, Fuhai Chen, Jie Li, Joe- mon M Jose, and Xuri Ge. 2025. Double-filter: Efficient fine-tuning of pre-trained vision-language models via patch&layer filtering. InForty-second International Conference on Machine Learning

  14. [23]

    Orland Hoeber. 2025. Design Principles for Exploratory Search Interfaces. In Proceedings of the 2025 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR ’25). Association for Computing Machinery, New York, NY, USA, 12–22. doi:10.1145/3698204.3716443

  15. [24]

    Sture Holm. 1979. A Simple Sequentially Rejective Multiple Test Procedure. Scandinavian Journal of Statistics6, 2 (1979), 65–70. http://www.jstor.org/stable/ 4615733

  16. [25]

    Yuta Imasaka and Hideo Joho. 2024. Effect of LLM’s Personality Traits on Query Generation. InProceedings of SIGIR-AP 2024. doi:10.1145/3673791.3698433

  17. [26]

    Shakyani Jayasiriwardene, Hongyu Zhou, Weiwei Jiang, Benjamin Tag, Nicholas Koemel, Matthew Ahmadi, Jorge Goncalves, Emmanuel Stamatakis, Anusha Withana, and Zhanna Sarsenbayeva. 2026. From Fixed to Flexible: Shaping AI Personality in Context-Sensitive Interaction. arXiv:2601....

  18. [27]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.ACM computing surveys55, 12 (2023), 1–38

  19. [28]

    John and Sanjay Srivastava

    Oliver P. John and Sanjay Srivastava. 1999. The Big Five Trait Taxonomy: History, Measurement, and Theoretical Perspectives. InHandbook of Personality: Theory and Research(2 ed.), Lawrence A. Pervin and Oliver P. John (Eds.). Guilford Press, 11 CIKM ’26, November 7–11, 2026, R...

  20. [29]

    Diane Kelly. 2009. Methods for Evaluating Interactive Information Re- trieval Systems with Users.Foundations and Trends in Information Retrieval 3, 1-2 (05 2009), 1–224. arXiv:https://www.emerald.com/ftinr/article-pdf/3/1- 2/1/11047868/1500000012en.pdf doi:10.1561/1500000012

  21. [30]

    Sanna Kumpulainen. 2017. Task-based information searching: research methods. Encyclopedia of library and information sciences(2017), 4526–4536

  22. [31]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing s...

  23. [32]

    Kai-Hui Liang, Weiyan Shi, Yoo Jung Oh, Hao-Chuan Wang, Jingwen Zhang, and Zhou Yu. 2024. Dialoging Resonance in Human-Chatbot Conversation: How Users Perceive and Reciprocate Recommendation Chatbot’s Self-Disclosure Strategy.Proc. ACM Hum.-Comput. Interact.8, CSCW1, Article 2...

  24. [33]

    François Mairesse and Marilyn Walker. 2008. Trainable generation of big-five personality styles through data-driven parameter estimation. InProceedings of ACL-08: HLT. 165–173

  25. [34]

    Gary Marchionini. 2006. Exploratory search: from finding to understanding. Commun. ACM49, 4 (2006), 41–46

  26. [35]

    Clifford Nass and Youngme Moon. 2000. Machines and Mindlessness: Social Responses to Computers.Journal of Social Issues56, 1 (2000), 81–103. doi:10. 1111/0022-4537.00153

  27. [36]

    Anastasiia Potiagalova, Joemon M Jose, Benjamin R Cowan, and Gareth JF Jones

  28. [37]

    Filip Radlinski and Nick Craswell. 2017. A Theoretical Framework for Conver- sational Search. InProceedings of the 2017 Conference on Human Information Interaction and Retrieval (CHIIR ’17). Association for Computing Machinery, New York, NY, USA, 117–126. doi:10.1145/3020165.3020183

  29. [38]

    1996.The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Places

    Byron Reeves and Clifford Nass. 1996.The Media Equation: How People Treat Computers, Television, and New Media Like Real People and Places. CSLI Pub- lications and Cambridge University Press. https://web.stanford.edu/group/ cslipublications/cslipublications/site/1575860538.shtml

  30. [39]

    Thomas Schmidt and Christian Wolff. 2016. Personality and information be- havior in web search.Proceedings of the Association for Information Science and Technology53, 1 (2016), 1–6

  31. [40]

    Gregory Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. 2023. Personality Traits in Large Language Models.arXiv preprint arXiv:2307.00184 (2023). https://arxiv.org/abs/2307.00184

  32. [41]

    Chirag Shah. 2025. From Prompt Engineering to Prompt Science with Humans in the Loop.Commun. ACM68, 6 (June 2025), 54–61. doi:10.1145/3709599

  33. [42]

    Weiwei Sun, Keyi Kong, Xinyu Ma, Shuaiqiang Wang, Dawei Yin, Maarten de Rijke, Zhaochun Ren, and Yiming Yang. 2026. ZeroGR: A Generalizable and Scalable Framework for Zero-Shot Generative Retrieval. InThe Fourteenth Inter- national Conference on Learning Representations. https...

  34. [43]

    Yalin Sun, Yan Zhang, Jacek Gwizdka, and Ciaran B Trace. 2019. Consumer evaluation of the quality of online health information: systematic literature review of relevant criteria and indicators.Journal of medical Internet research21, 5 (2019), e12522

  35. [44]

    Paul Thomas, Mary Czerwinski, Daniel McDuff, Nick Craswell, and Gloria Mark

  36. [45]

    Paul Thomas, Daniel McDuff, Mary Czerwinski, and Nick Craswell. 2020. Expres- sions of Style in Information-Seeking Conversation with an Embodied Agent. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20...

  37. [46]

    Sarah Theres Völkel, Ramona Schoedel, Lale Kaya, and Sven Mayer. 2022. User Perceptions of Extraversion in Chatbots after Repeated Use. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems(New Orleans, LA, USA)(CHI ’22). Association for Computing Mach...

  38. [47]

    Ryen White, Ian Ruthven, and Joemon M. Jose. 2002. Finding relevant documents using top ranking sentences: an evaluation of two alternative schemes. InSIGIR 2002: Proceedings of the 25th Annual International ACM SIGIR Conference on Research and Development in Information Retri...

  39. [48]

    White and Resa A

    Ryen W. White and Resa A. Roth. 2009.Exploratory Search: Beyond the Query- Response Paradigm. Morgan & Claypool Publishers. doi:10.1007/978-3-031- 02260-9

  40. [49]

    Yu Ye, Junchen Fu, Yu Song, Kaiwen Zheng, and Joemon M Jose. 2026. Are multimodal embeddings truly beneficial for recommendation? A deep dive into whole vs. individual modalities. InEuropean Conference on Information Retrieval. Springer, 66–81

  41. [50]

    Trippas, Jeff Dalton, and Filip Radlinski

    Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023. Conver- sational information seeking.Foundations and Trends®in Information Retrieval 17, 3–4 (2023), 244–456

  42. [51]

    Binquan Zhang, Li Zhang, Haoyuan Zhang, Fang Liu, Song Wang, Bo Shen, An Fu, and Lin Shi. 2025. Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild. arXiv:2512.10493 [cs.SE] https: //arxiv.org/abs/2512.10493

  43. [52]

    Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M Jose, and Xuri Ge

  44. [53]

    Ziyi Zhuang, Hongji Li, Junchen Fu, Jiacheng Liu, Joemon M Jose, Youhua Li, and Yongxin Ni. 2025. Frequency-Decoupled distillation for efficient multimodal recommendation. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 4571–4581. 12

  45. [2018]

    InProceedings of the 2018 Conference on Human Information Interaction and Retrieval (CHIIR ’18)

    Style and Alignment in Information-Seeking Conversation. InProceedings of the 2018 Conference on Human Information Interaction and Retrieval (CHIIR ’18). 42–51. doi:10.1145/3176349.3176388

  46. [2019]

    InProceedings of the 42nd International ACM SIGIR Confer- ence on Research and Development in Information Retrieval(Paris, France)(SI- GIR’19)

    Asking Clarifying Questions in Open-Domain Information-Seeking Conversations. InProceedings of the 42nd International ACM SIGIR Confer- ence on Research and Development in Information Retrieval(Paris, France)(SI- GIR’19). Association for Computing Machinery, New York, NY, USA,...

  47. [2025]

    In2025 International Conference on Content-Based Multimedia Indexing (CBMI)

    A Comparative Study of Conversational and Conventional Search Methods for Image Retrieval. In2025 International Conference on Content-Based Multimedia Indexing (CBMI). IEEE, 1–7

  48. [2026]

    Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?arXiv preprint arXiv:2607.12787(2026)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.