REVIEW 3 major objections 5 minor 72 references
Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that most users' mental models of generative conversational search are too abstract to interpret any single search instance, and that adding transparency cues does not repair this gap, instead stabilizing expectation…
desk verdict Qualitative mental model work is solid and worth reading; the transparency experiment is fatally confounded by fixed condition order, so treat RQ2 conclusions with caution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study is carried by the mental-model construct, organized using a four-part framework -- components, functions, attributes, and feelings -- and elicited through semi-structured interviews, think-aloud protocols, pre/post expectation-violation ratings, and logs of query repairs. The transparency manipulation is implemented as four retrieval-augmented generation (RAG) chat interfaces that differ only in which of three textual vectors are shown: source attribution, query transformation, and a yes/no faithfulness flag. The working mechanism is that these vectors are meant to act as interaction cues that users can fold into their mental models, but the paper observes that a cue only teaches when the user already has enough of a model to interpret it.
What would settle it
Run the same four conditions with the order counterbalanced across participants: if the downward satisfaction trend and the failure of transparency to improve learning disappear when order is randomized, the RQ2 conclusions are artifacts of the fixed sequence rather than real effects.
Extended reading notes
Core claim
On the paper's own terms, the central finding is that most users have a globally plausible picture of generative conversational search -- an LLM trained on internet data that takes a natural-language query and produces a response -- but cannot apply that picture to interpret a particular result. That leaves room for contradictory beliefs (the same participants sometimes described the system as an unstructured corpus and sometimes as a searched database), for heavy reliance on imagined limitations, and for volatile trust rather than stable over- or under-trust. The transparency manipulation did not reliably teach: source links were noticed and liked but did not convey online retrieval unless users already knew the system could search online; query transformations were treated as reusable search queries rather than evidence of conversation history; and even the most-transparent condition did not improve learning for users with incomplete mental models. Expectation violations stabilized after the first transparency addition, but satisfaction trended downward, which the authors interpret as transparency exposing system errors and limitations rather than repairing understanding.
Load-bearing premise
The load-bearing assumption is that the four interfaces differ only in the transparency vectors and that the fixed order of conditions, always from baseline to most-transparent, did not itself cause the observed drift in satisfaction and expectations.
Editorial extensions
If this is right
- Source links alone will not teach users that a chatbot can retrieve live web results; users may see them as a verification aid or even as possibly fake.
- Showing query transformations and faithfulness flags can lower satisfaction when they surface system errors, even while stabilizing the user's expectations.
- Users with abstract mental models compensate by building hybrid workflows, so designs that make switching between chat and web search seamless would support behavior users already exhibit.
- Mental-model incompleteness, not just interface opacity, should be treated as the trust problem; interventions such as onboarding or pre-interaction explanation may be needed.
- The same interface can produce over-trust in some users and under-trust in others, so studies of conversational search should measure trust volatility rather than only average trust.
Reading between the lines
- Editorial inference: a longitudinal version of this study might find that transparency pays off only after repeated exposure, so the single-session satisfaction decline observed here may understate the long-term value of the cues.
- Editorial inference: a natural next experiment would hold the interface constant and vary a short tutorial explaining what sources, query transformations, and faithfulness mean; if mental-model accuracy rises, the bottleneck is missing knowledge rather than missing cues.
- Editorial inference: transparency could be made adaptive -- for example, revealing query transformations only after the user demonstrates an understanding of what a query is -- which would directly test the paper's claim that cues require prior model completeness.
- Editorial inference: the hybrid-workflow result suggests trust in conversational search is best modeled as a choice between two complementary tools rather than as a property of the chat system alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a within-subject, mixed-methods study with 16 participants who each completed four search tasks using four RAG-based conversational search interfaces that differ only in the number of transparency vectors (Baseline, Least Transparent, Transparent, Most Transparent). The paper's RQ1 asks what mental models users have of generative conversational search, and RQ2 asks how interface transparency affects mental models, expectations, and satisfaction. The qualitative analysis (interviews, think-aloud, content analysis) finds that most participants hold abstract, incomplete, and sometimes contradictory mental models, describing the system as a 'black box' while also referencing databases and online sources, and that users compensate through hybrid web-conversational workflows. The quantitative and comparative aspects of RQ2 report descriptive satisfaction means (5.00, 4.60, 3.38, 4.33) and expectation-violation deltas across interfaces, and conclude that transparency may reduce satisfaction while stabilizing expectation violation, and that it helps interpretability only when mental models are already more complete. The paper frames these findings as design-relevant contributions to conversational search and trust calibration.
Significance. The RQ1 contribution is valuable and timely: it provides a rich, systematically analyzed account of mental models of generative conversational search, a topic that is underexplored despite the rapid adoption of these systems. The study's use of multiple elicitation techniques (interviews, think-aloud, self-reports) and its qualitative insights into contradictory mental models and hybrid search workflows are a credible foundation for future design work. The paper also offers plausible, falsifiable hypotheses (H1–H3) for future validation. However, the RQ2 conclusions about the causal effects of transparency are seriously undermined by the fixed condition order, which perfectly confounds interface condition with session position. If the RQ2 claims are softened to exploratory, descriptive observations and the confound is explicitly acknowledged, the paper's central qualitative contribution remains sound. As it stands, the transparency-effect claims go beyond what the design and analysis can support.
major comments (3)
- [Section 3.5.1] The fixed condition order (always baseline to most-transparent) makes transparency perfectly confounded with session position for every participant, task order randomization notwithstanding. Consequently, all RQ2 comparisons in Section 4.3—the declining satisfaction means, the apparent stabilization of expectation violation, and the observation that transparency did not aid learning—are equally explainable by fatigue, learning, or growing criticalness across the session. Section 7 does not acknowledge this confound. The authors should either present RQ2 explicitly as an exploratory, order-sensitive description with a prominent limitation, or, if they wish to retain causal language, provide evidence that order effects are negligible, which the current design cannot supply.
- [Section 4.3.2] The central RQ2 quantitative claims rest on descriptive means and standard deviations only (e.g., satisfaction 5.00, 4.60, 3.38, 4.33; SDs 1.41, 2.26, 2.36, 1.95) with no inferential statistics. Even setting aside the absence of significance tests, the pre-task expectation measure was administered only once before any interaction, so the expectation-violation delta for the transparent and most-transparent interfaces compares post-task ratings against a baseline formed before the participant had used any of the study's interfaces. This mixes updated expectations with the transparency manipulation and further prevents causal attribution. These analyses should be reframed as descriptive observations, not as evidence of a causal effect of transparency.
- [Section 4.3.1] The claim that transparency vectors 'improved system interpretability only when mental models were more complete' is based on the observation that only N=3 participants with richer prior models connected source attribution to online retrieval. This comparison is order-sensitive because those participants had already experienced the baseline and least-transparent interfaces by the time they encountered the transparent conditions, so the connection could have been learned within the session rather than imported from prior knowledge. The manuscript needs to either present a systematic comparison of participants by mental-model completeness that accounts for the fixed order, or substantially weaken the causal wording of this finding.
minor comments (5)
- [Section 1] There is a typo in the Introduction: 'which are are neural-net-based' should read 'which are neural-net-based'.
- [Section 5.1] In the sentence beginning 'Indeed, users invested cognitive effort...', the phrase 'hybrid web-CA search paradigmsAs' contains a missing space between 'paradigms' and 'As'.
- [Section 3.5.1] The phrase '[withdrawn for review]' for the ethics committee approval number is a placeholder; it should be replaced with the actual approval reference or a note explaining how to obtain it.
- [Section 2.1] The word 'outwith' (e.g., 'Outwith the conversational domain') is regionally specific; consider using 'outside' for broader accessibility.
- [Figure 6] The parenthetical note in the caption ('Where the median line is “missing” in a box plot, then the median line coincides with one of the quartile lines') is awkwardly placed; consider moving it to the main text or a footnote.
Circularity Check
No circular derivation; the paper's claims are empirical inferences from interviews, think-aloud, surveys, and logs, not consequences of definitions or fitted parameters.
full rationale
This is an empirical HCI study, not a derivation chain. The central claims—that mental models are too abstract to support interpreting individual search instances and that transparency vectors do not close this gap—are supported by thematic analysis of interviews, think-aloud transcripts, conversation logs, and exploratory survey statistics. The classification of mental-model completeness is an analytic interpretation of the same interview data that is later used to explain transparency effects, but this is inductive reasoning rather than a self-definitional equivalence: the paper does not define 'abstract mental model' as 'does not benefit from transparency,' nor are its quantitative satisfaction and expectation-violation outcomes fitted to their own explanatory labels. The cited prior work, including the authors' own [16] and [45], is background literature on conversational agents and appropriate trust; it is not invoked as a uniqueness theorem or as a proof of the present empirical results. No fitted parameter is renamed as a prediction, and no known result is re-derived from its own assumptions. One genuine validity concern—the fixed condition order described in Section 3.5.1, which perfectly confounds interface transparency with session position—is a methodological threat to the RQ2 comparisons, but it does not make the paper circular, because the conclusions still rest on observed patterns rather than on a constructed equivalence. Therefore no circularity step is identified, and the appropriate score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Verbal reports (think-aloud and interviews) are a valid window on latent mental models.
- domain assumption The four prototype interfaces vary only in the three transparency vectors and together represent generative conversational search.
- domain assumption The llamaindex faithfulness flag correctly indicates whether the response matches source documents.
- domain assumption Zhang's components/functions/attributes/feelings framework transfers to generative conversational search mental models.
Cite this review
Pith. "Pith review of Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency." pith.science (2026). https://pith.science/paper/3NCEQYIV
@misc{pith2026250603807,
author = {Pith},
title = {Pith review of: Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency},
year = {2026},
howpublished = {\url{https://pith.science/paper/3NCEQYIV}},
note = {Machine review of arXiv:2506.03807}
}
read the original abstract
The experience and adoption of conversational search is tied to the accuracy and completeness of users' mental models -- their internal frameworks for understanding and predicting system behaviour. Thus, understanding these models can reveal areas for design interventions. Transparency is one such intervention which can improve system interpretability and enable mental model alignment. While past research has explored mental models of search engines, those of generative conversational search remain underexplored, even while the popularity of these systems soars. To address this, we conducted a study with 16 participants, who performed 4 search tasks using 4 conversational interfaces of varying transparency levels. Our analysis revealed that most user mental models were too abstract to support users in explaining individual search instances. These results suggest that 1) mental models may pose a barrier to appropriate trust in conversational search, and 2) hybrid web-conversational search is a promising novel direction for future search interface design.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Hervé Abdi and Lynne J Williams. 2010. Principal component analysis. Wiley interdisciplinary reviews: computational statistics 2, 4 (2010), 433–459
2010
-
[2]
Alexandre Arnold, Gérard Dupont, Catherine Kobus, François Lancelot, and Ying-Hsang Liu. 2020. Perceived Usefulness of Conversational Agents Predicts Search Performance in Aerospace Domain. In Proceedings of the 2nd Conference on Conversational User Interfaces (CUI ’20) . Association for Computing Machinery, New York, NY, USA, 1–3. https://doi.org/10.1145...
-
[3]
Zahra Ashktorab, Mohit Jain, Q. Vera Liao, and Justin D. Weisz. 2019. Resilient Chatbots: Repair Strategy Preferences for Conversational Breakdowns. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1–12. https://doi.org/10.1145/3290605.3300484
arXiv 2019
-
[4]
Leif Azzopardi, Mateusz Dubiel, Martin Halvey, and Jeffery Dalton. 2024. A Conceptual Framework for Conversational Search and Recommendation: Conceptualizing Agent-Human Interactions During the Conversational Search Process. (April 2024). https://doi.org/10.48550/arXiv.2404.08630 arXiv:2404.08630 [cs]
work page Pith review arXiv doi:10.48550/arxiv.2404.08630 2024
-
[5]
Jeff A Bauhs and Nancy J Cooke. 1994. Is knowing more really better? Effects of system development information in human-expert system interactions. In Conference Companion on Human Factors in Computing Systems . 99–100
work page 1994
-
[6]
Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Lierni Sestorain Saralegui, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, and Kellie Webster. 2023. Att...
-
[7]
T. Boren and J. Ramey. 2000. Thinking aloud: reconciling theory and practice. IEEE Transactions on Professional Communication 43, 3 (Sept. 2000), 261–278. https://doi.org/10.1109/47.867942
-
[8]
Virginia Braun and Victoria Clarke. 2021. One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative Research in Psychology 18, 3 (July 2021), 328–352. https://doi.org/10.1080/14780887.2020.1769238
arXiv 2021
Show all 72 references
-
[9]
V Braun and V Clarke. 2021. Thematic analysis: a practical guide [eBook version]. SAGE moradi H, vaezi A. lessons learned from Korea: COVID-19 pandemic 41 (2021), 873–4
2021
-
[10]
Carroll and Judith Reitman Olson
John M. Carroll and Judith Reitman Olson. 1988. Chapter 2 - Mental Models in Human-Computer Interaction . North-Holland, Amsterdam, 45–65. https://doi.org/10.1016/b978-0-444-70536-5.50007-5
1988 doi
-
[11]
Victoria Clarke and Virginia Braun. 2013. Successful qualitative research: A practical guide for beginners . Sage publications ltd, London, UK. 1–400 pages
2013
- [12]
-
[13]
de Visser, Marieke M
Ewart J. de Visser, Marieke M. M. Peeters, Malte F. Jung, Spencer Kohn, Tyler H. Shaw, Richard Pak, and Mark A. Neerincx. 2020. Towards a Theory of Longitudinal Trust Calibration in Human–Robot Teams. International Journal of Social Robotics 12, 2 (May 2020), 459–478. https: /...
2020 doi
-
[14]
Smit Desai and Michael Twidale. 2022. Is Alexa like a computer? A search engine? A friend? A silly child? Yes.. InProceedings of the 4th Conference on Conversational User Interfaces (CUI ’22) . Association for Computing Machinery, New York, NY, USA, 1–4. https://doi.org/10.114...
2022
- [15]
- [16]
-
[17]
Ghosh, J
S. Ghosh, J. Gogoi, and K. Chua. 2023. Exploring the economics of conversational search sessions. Aslib Journal of Information Management 76, 4 (2023), 613-628 pages. https://doi.org/10.1108/AJIM-08-2022-0368
2023 doi
-
[18]
Mark Grimes, Ryan M
G. Mark Grimes, Ryan M. Schuetzler, and Justin Scott Giboney. 2021. Mental models and expectation violations in conversational AI interactions. Decision Support Systems 144 (May 2021), 113515. https://doi.org/10.1016/j.dss.2021.113515
2021
-
[19]
Monique Harrison and Philip Hernandez. 2022. Supporting Interviews with Technology: How Software Integration Can Benefit Participants and Interviewers. (April 2022). https://doi.org/10.2139/ssrn.4113477 Preprint
2022 doi
-
[20]
Harwood and Tony Garry
Tracy G. Harwood and Tony Garry. 2003. An Overview of Content Analysis. The Marketing Review 3, 4 (Dec. 2003), 479–498. https://doi.org/10. 1362/146934703771910080
2003
-
[21]
Hernandez-Bocanegra and Jürgen Ziegler
Diana C. Hernandez-Bocanegra and Jürgen Ziegler. 2021. Conversational review-based explanations for recommender systems: Exploring users’ query behavior. In Proceedings of the 3rd Conference on Conversational User Interfaces (CUI ’21) . Association for Computing Machinery, New...
2021
-
[22]
Hoffman, Timothy Miller, Gary Klein, Shane T
Robert R. Hoffman, Timothy Miller, Gary Klein, Shane T. Mueller, and William J. Clancey. 2023. Increasing the Value of XAI for Users: A Psychological Perspective. KI - Künstliche Intelligenz 37, 2 (Dec. 2023), 237–247. https://doi.org/10.1007/s13218-023-00806-9
2023 doi
- [23]
-
[24]
Holtgraves, S.J
T.M. Holtgraves, S.J. Ross, C.R. Weywadt, and T.L. Han. 2007. Perceiving artificial social agents. Computers in Human Behavior 23, 5 (Sept. 2007), 2163–2174. https://doi.org/10.1016/j.chb.2006.02.017 Transparency and Models of Conversational Search 21
2007 doi
-
[25]
Tom Janssen. 2024. Why search engines and chatbots are becoming more alike. https://www.universiteitleiden.nl/en/news/2024/05/why-search- engines-and-chatbots-are-becoming-more-alike
2024
-
[26]
Craig J Johnson, Mustafa Demir, Nathan J McNeese, Jamie C Gorman, Alexandra T Wolff, and Nancy J Cooke. 2021. The impact of training on human–autonomy team communications and trust calibration. Human factors (2021), 00187208211047323
2021
-
[27]
Carolina Centeio Jorge, Siddharth Mehrotra, Catholijn M Jonker, and Myrthe L Tielman. 2021. Trust should correspond to Trustworthiness: a Formalization of Appropriate Mutual Trust in Human-Agent Teams. Proceedings of the 22nd International Workshop on Trust in Agent Societies ...
2021
-
[28]
Hyunhoon Jung, Changhoon Oh, Gilhwan Hwang, Cindy Yoonjung Oh, Joonhwan Lee, and Bongwon Suh. 2019. Tell Me More: Understanding User Interaction of Smart Speaker News Powered by Conversational Search. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computi...
2019
-
[29]
Cecilia Katzeff. 1990. System demands on mental models for a fulltext database. International Journal of Man-Machine Studies 32, 5 (May 1990), 483–509. https://doi.org/10.1016/S0020-7373(05)80031-8
1990 doi
-
[30]
Siddartha Khastgir, Stewart Birrell, Gunwant Dhadyalla, and Paul Jennings. 2018. Calibrating trust through knowledge: Introducing the concept of informed safety for automation in vehicles. Transportation Research Part C: Emerging Technologies 96 (Nov. 2018), 290–303. https://d...
2018 doi
-
[31]
Michael Khoo and Catherine Hall. 2012. What Would ‘Google’ Do? Users’ Mental Models of a Digital Library Search Engine. In Theory and Practice of Digital Libraries , Panayiotis Zaphiris, George Buchanan, Edie Rasmussen, and Fernando Loizides (Eds.). Springer, Berlin, Heidelber...
2012 doi
-
[32]
Johannes Kiesel, Damiano Spina, Henning Wachsmuth, and Benno Stein. 2021. The Meant, the Said, and the Understood: Conversational Argument Search and Cognitive Biases. In Proceedings of the 3rd Conference on Conversational User Interfaces (CUI ’21) . Association for Computing ...
2021
- [33]
-
[34]
Linda Kopitz and Linda Kopitz. 2021. Alexa, Affect, and the Algorithmic Imaginary: Addressing Privacy and Security Concerns Through Emotional Advertising. Screen 6, 1 (June 2021), 1–17. https://doi.org/10.3167/screen.2021.060103
2021
-
[35]
Todd Kulesza, Margaret Burnett, Weng-Keen Wong, and Simone Stumpf. 2015. Principles of Explanatory Debugging to Personalize Interactive Machine Learning. In Proceedings of the 20th International Conference on Intelligent User Interfaces (IUI ’15) . Association for Computing Ma...
2015
-
[36]
Todd Kulesza, Simone Stumpf, Margaret Burnett, and Irwin Kwan. 2012. Tell me more? the effects of mental model soundness on personalizing an intelligent agent. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’12) . Association for Computing M...
2012
-
[37]
Weronika Łajewska, Damiano Spina, Johanne Trippas, and Krisztian Balog. 2024. Explainability for Transparent Conversational Information-Seeking. (May 2024). https://doi.org/10.1145/3626772.3657768 Preprint
2024
-
[38]
John D Lee and Neville Moray. 1994. Trust, self-confidence, and operators’ adaptation to automation. International journal of human-computer studies 40, 1 (1994), 153–184
1994
-
[39]
Lee and Katrina A
John D. Lee and Katrina A. See. 2004. Trust in automation: designing for appropriate reliance. Human Factors 46, 1 (2004), 50–80. https: //doi.org/10.1518/hfes.46.1.50_30392
2004 doi
-
[40]
Sunok Lee, Minji Cho, and Sangsu Lee. 2020. What If Conversational Agents Became Invisible? Comparing Users’ Mental Models According to Physical Entity of AI Speaker. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 4, 3 (Sept. 2020), 88:1–88...
2020 doi
-
[41]
Vera Liao, Werner Geyer, Michael Muller, and Yasaman Khazaen
Q. Vera Liao, Werner Geyer, Michael Muller, and Yasaman Khazaen. 2020.Conversational Interfaces for Information Search . Springer International Publishing, Cham, 267–287. https://doi.org/10.1007/978-3-030-38825-6_13
2020 doi
-
[42]
Satisfaction with Failure
Mengyang Liu, Yiqun Liu, Jiaxin Mao, Cheng Luo, Min Zhang, and Shaoping Ma. 2018. "Satisfaction with Failure" or "Unsatisfied Success": Investigating the Relationship between Search Success and User Satisfaction. In Proceedings of the 2018 World Wide Web Conference (Lyon, Fran...
2018
-
[43]
Nelson Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating Verifiability in Generative Search Engines. InFindings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 7001–7025. https://doi.org/10.18653/v1/2023.fi...
2023 doi
-
[44]
Like Having a Really Bad PA
Ewa Luger and Abigail Sellen. 2016. “Like Having a Really Bad PA”: The Gulf between User Expectation and Experience of Conversational Agents. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems . ACM, San Jose California USA, 5286–5297. https: //doi...
2016
-
[45]
Jonker, and Myrthe L
Siddharth Mehrotra, Chadha Degachi, Oleksandra Vereschak, Catholijn M. Jonker, and Myrthe L. Tielman. 2024. A Systematic Review on Fostering Appropriate Trust in Human-AI Interaction: Trends, Opportunities and Challenges. ACM J. Responsib. Comput. 1, 4, Article 26 (Nov. 2024),...
2024 doi
-
[46]
Sifiso Mlilo and Andrew Thatcher. 2011. Mental Models: Have Users’ Mental Models of Web Search Engines Improved in the Last Ten Years?. In Engineering Psychology and Cognitive Ergonomics , Don Harris (Ed.). Springer, Berlin, Heidelberg, 243–253. https://doi.org/10.1007/978-3-6...
2011 doi
-
[47]
Isabela Motta and Manuela Quaresma. 2022. Exploring the Opinions of Experts in Conversational Design: A Study on Users’ Mental Models of Voice Assistants. In Human-Computer Interaction. User Experience and Behavior , Masaaki Kurosu (Ed.). Springer International Publishing, Cha...
2022 doi
-
[48]
Isabela Motta and Manuela Quaresma. 2023. Increasing Transparency to Design Inclusive Conversational Agents (CAs): Perspectives and Open Issues. In Proceedings of the 5th International Conference on Conversational User Interfaces (CUI ’23) . Association for Computing Machinery...
2023
-
[49]
Jack Muramatsu and Wanda Pratt. 2001. Transparent Queries: investigation users’ mental models of search engines. In Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval (SIGIR ’01) . Association for Computing Ma...
2001
-
[50]
Don Norman. 1988. The Design of Everyday Things
1988
-
[51]
Donald A Norman. 2014. Some observations on mental models. In Mental models. Psychology Press, 15–22
2014
-
[52]
O’Brien and S
L. O’Brien and S. Wilson. 2023. Talking About Thinking Aloud: Perspectives from Interactive Think- Aloud Practitioners. Journal of User Experience 18, 33 (May 2023), 113–132
2023
-
[53]
Andrea Papenmeier, Alexander Frummet, and Dagmar Kern. 2022. “Mhm... ” – Conversational Strategies For Product Search Assistants. InProceedings of the 2022 Conference on Human Information Interaction and Retrieval (CHIIR ’22) . Association for Computing Machinery, New York, NY...
2022
-
[54]
Raja Parasuraman and Victor Riley. 1997. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors 39, 2 (June 1997), 230–253. https://doi.org/10.1518/001872097778543886
1997 doi
-
[55]
Elizabeth Phillips, Scott Ososky, Janna Grove, and Florian Jentsch. 2011. From Tools to Teammates: Toward the Development of Appropriate Mental Models for Intelligent Robots. Proceedings of the Human Factors and Ergonomics Society Annual Meeting 55, 1 (Sept. 2011), 1491–1495. ...
2011 doi
-
[56]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving Language Understanding by Generative Pre-Training. (2018). https://paperswithcode.com/paper/improving-language-understanding-by Preprint
2018
-
[57]
Filip Radlinski and Nick Craswell. 2017. A Theoretical Framework for Conversational Search. In Proceedings of the 2017 Conference on Conference Human Information Interaction and Retrieval (CHIIR ’17) . Association for Computing Machinery, New York, NY, USA, 117–126. https://do...
2017
- [58]
-
[59]
Vahid Sadiri Javadi, Johanne R Trippas, and Lucie Flek. 2024. Unveiling Information Through Narrative In Conversational Information Seeking. In Proceedings of the 6th ACM Conference on Conversational User Interfaces (CUI ’24) . Association for Computing Machinery, New York, NY...
2024
-
[60]
David Schuster Scott Ososky. 2013. Building Appropriate Trust in Human-Robot Teams. 2013 AAAI spring symposium series 7 (2013). https: //aaai.org/papers/05784-5784-building-appropriate-trust-in-human-robot-teams/
2013
- [61]
-
[62]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval Augmentation Reduces Hallucination in Conversation. arXiv:2104.07567 [cs.CL]
2021 arXiv
- [63]
-
[64]
Nancy Staggers and Anthony F. Norcio. 1993. Mental models: concepts for human-computer interaction research. International Journal of Man-machine studies 38, 4 (1993), 587–605
1993
-
[65]
Paul Thomas, Bodo Billerbeck, Nick Craswell, and Ryen W. White. 2020. Investigating Searchers’ Mental Models to Inform Search Explanations. ACM Transactions on Information Systems 38, 1 (Jan. 2020), 1–25. https://doi.org/10.1145/3371390
2020 doi
-
[66]
Z. Xing, X. Yuan, D. Wu, Y. Huang, and J. Mostafa. 2020. Understanding Voice Search Behavior: Review and Synthesis of Research. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 12424 LNCS (2020...
2020 doi
-
[67]
Trippas, Jeff Dalton, and Filip Radlinski
Hamed Zamani, Johanne R. Trippas, Jeff Dalton, and Filip Radlinski. 2023. Conversational Information Seeking. Foundations and Trends in Information Retrieval 17, 3–4 (Aug. 2023), 244–456. https://doi.org/10.1561/1500000081
2023 doi
-
[68]
Ines Zelch, Matthias Hagen, and Martin Potthast. 2024. A User Study on the Acceptance of Native Advertising in Generative IR. In Proceedings of the 2024 Conference on Human Information Interaction and Retrieval (Sheffield, United Kingdom) (CHIIR ’24). Association for Computing...
2024
-
[69]
Yan Zhang. 2008. Undergraduate students’ mental models of the Web as an information retrieval system. Journal of the American Society for Information Science and Technology 59, 13 (2008), 2087–2098. https://doi.org/10.1002/asi.20915
2008 doi
-
[70]
Bruce Croft
Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W. Bruce Croft. 2018. Towards Conversational Search and Recommendation: System Ask, User Respond. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management (CIKM ’18) . Association for Com...
2018
-
[71]
Vera Liao, and Rachel K
Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* ’20). Association for C...
2020
-
[72]
It’s a Fair Game
Zhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024. “It’s a Fair Game”, or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents. In Proceedings of the...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.