{"id":"d852d27e-e69d-4fae-8a4e-2a3d15eae300","arxiv_id":"2506.03807","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Users of generative conversational search mostly hold abstract, incomplete mental models, and added interface transparency did not reliably improve those models or satisfaction.","lead":"This study of 16 people found that most users of AI-powered conversational search have vague, and sometimes contradictory, mental models of how the system works. It also suggests that adding extra transparency details, such as source links or rewritten queries, did not clearly improve understanding and may have lowered satisfaction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed condition order confounds every RQ2 comparison: transparency is perfectly correlated with session position, so the satisfaction decline and the 'no learning' pattern could be sequence effects rather than transparency effects.","rationale":"The paper's first contribution (RQ1) is supported by detailed qualitative data and is largely independent of the order confound. The second contribution (RQ2) is the load-bearing part of the paper's headline claim about transparency. Because condition order and transparency are perfectly confounded, the causal language in Section 5 ('exposing search mechanisms through transparency vectors does not support learning') cannot be inferred from this design. This is an internal validity problem, not a disagreement with field consensus, and it is implicitly acknowledged by the authors' own decision to 'gradually introduce interventions' but is not listed as a limitation. The reader's weakest assumption identified the same issue, so I agree with that assessment. The qualitative mental-model explication retains independent value, and the transparency claims are recoverable with a re-designed or re-analyzed study, so the appropriate recommendation is to request revision rather than reject the paper. This leaves the reader's conditional verdict unchanged.","tokens_in":19728,"tokens_out":5441,"duration_ms":57468,"concrete_test":"Conduct a pre-registered replication (or a supplementary second cohort) with the same four RAG interfaces and approximately the same N, but with condition order counterbalanced in a Latin square and a pre-task expectation rating immediately before each condition. If the monotonic satisfaction decline and the 'only richer mental models benefit from sources' pattern do not reproduce under counterbalancing, the original RQ2 conclusions are sequence effects rather than transparency effects. If they do reproduce, the fixed-order concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.5.1 states that task order was randomized but condition order was not: participants always moved from baseline to most-transparent. Thus interface transparency and session position are perfectly confounded for all quantitative and comparative qualitative RQ2 analyses. The satisfaction means (5.00, 4.60, 3.38, 4.33) and the apparent stabilization of expectation violation from the least-transparent condition onward are exactly what one would expect from fatigue, learning, or growing criticalness, and the transcripts contain direct evidence of such a trend (e.g., P2 becoming 'a little bit sceptical' after an irrelevant link in a later condition). Because the pre-task expectation survey was administered only once before any interaction, post-task deltas for later interfaces also mix updated expectations with the manipulation. The claim that only N=3 participants with richer prior models connected sources to online retrieval is likewise order-sensitive: those participants had already seen the system operate by the time they reached the transparent conditions. Section 7 does not acknowledge this confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a within-subject, mixed-methods study with 16 participants who each completed four search tasks using four RAG-based conversational search interfaces that differ only in the number of transparency vectors (Baseline, Least Transparent, Transparent, Most Transparent). The paper's RQ1 asks what mental models users have of generative conversational search, and RQ2 asks how interface transparency affects mental models, expectations, and satisfaction. The qualitative analysis (interviews, think-aloud, content analysis) finds that most participants hold abstract, incomplete, and sometimes contradictory mental models, describing the system as a 'black box' while also referencing databases and online sources, and that users compensate through hybrid web-conversational workflows. The quantitative and comparative aspects of RQ2 report descriptive satisfaction means (5.00, 4.60, 3.38, 4.33) and expectation-violation deltas across interfaces, and conclude that transparency may reduce satisfaction while stabilizing expectation violation, and that it helps interpretability only when mental models are already more complete. The paper frames these findings as design-relevant contributions to conversational search and trust calibration.","tokens_in":19923,"tokens_out":3813,"duration_ms":37929,"significance":"The RQ1 contribution is valuable and timely: it provides a rich, systematically analyzed account of mental models of generative conversational search, a topic that is underexplored despite the rapid adoption of these systems. The study's use of multiple elicitation techniques (interviews, think-aloud, self-reports) and its qualitative insights into contradictory mental models and hybrid search workflows are a credible foundation for future design work. The paper also offers plausible, falsifiable hypotheses (H1–H3) for future validation. However, the RQ2 conclusions about the causal effects of transparency are seriously undermined by the fixed condition order, which perfectly confounds interface condition with session position. If the RQ2 claims are softened to exploratory, descriptive observations and the confound is explicitly acknowledged, the paper's central qualitative contribution remains sound. As it stands, the transparency-effect claims go beyond what the design and analysis can support.","major_comments":[{"comment":"The fixed condition order (always baseline to most-transparent) makes transparency perfectly confounded with session position for every participant, task order randomization notwithstanding. Consequently, all RQ2 comparisons in Section 4.3—the declining satisfaction means, the apparent stabilization of expectation violation, and the observation that transparency did not aid learning—are equally explainable by fatigue, learning, or growing criticalness across the session. Section 7 does not acknowledge this confound. The authors should either present RQ2 explicitly as an exploratory, order-sensitive description with a prominent limitation, or, if they wish to retain causal language, provide evidence that order effects are negligible, which the current design cannot supply.","section":"Section 3.5.1"},{"comment":"The central RQ2 quantitative claims rest on descriptive means and standard deviations only (e.g., satisfaction 5.00, 4.60, 3.38, 4.33; SDs 1.41, 2.26, 2.36, 1.95) with no inferential statistics. Even setting aside the absence of significance tests, the pre-task expectation measure was administered only once before any interaction, so the expectation-violation delta for the transparent and most-transparent interfaces compares post-task ratings against a baseline formed before the participant had used any of the study's interfaces. This mixes updated expectations with the transparency manipulation and further prevents causal attribution. These analyses should be reframed as descriptive observations, not as evidence of a causal effect of transparency.","section":"Section 4.3.2"},{"comment":"The claim that transparency vectors 'improved system interpretability only when mental models were more complete' is based on the observation that only N=3 participants with richer prior models connected source attribution to online retrieval. This comparison is order-sensitive because those participants had already experienced the baseline and least-transparent interfaces by the time they encountered the transparent conditions, so the connection could have been learned within the session rather than imported from prior knowledge. The manuscript needs to either present a systematic comparison of participants by mental-model completeness that accounts for the fixed order, or substantially weaken the causal wording of this finding.","section":"Section 4.3.1"}],"minor_comments":[{"comment":"There is a typo in the Introduction: 'which are are neural-net-based' should read 'which are neural-net-based'.","section":"Section 1"},{"comment":"In the sentence beginning 'Indeed, users invested cognitive effort...', the phrase 'hybrid web-CA search paradigmsAs' contains a missing space between 'paradigms' and 'As'.","section":"Section 5.1"},{"comment":"The phrase '[withdrawn for review]' for the ethics committee approval number is a placeholder; it should be replaced with the actual approval reference or a note explaining how to obtain it.","section":"Section 3.5.1"},{"comment":"The word 'outwith' (e.g., 'Outwith the conversational domain') is regionally specific; consider using 'outside' for broader accessibility.","section":"Section 2.1"},{"comment":"The parenthetical note in the caption ('Where the median line is “missing” in a box plot, then the median line coincides with one of the quartile lines') is awkwardly placed; consider moving it to the main text or a footnote.","section":"Figure 6"}],"recommendation":"major_revision","confidential_remarks":"The RQ1 mental-model findings are a solid qualitative contribution, but the RQ2 causal claims about transparency are not supported by the design because the fixed condition order is a perfect confound. I would advise the editor that the manuscript can be made publishable if the authors substantially reframe RQ2 as exploratory and descriptive, explicitly acknowledge the order confound in the limitations, and remove or heavily qualify causal statements such as 'transparency drove down satisfaction' and 'stabilised expectation violation'. The descriptive observations can still serve as motivation for future counterbalanced studies. I do not see this as a reject; the qualitative contribution and the design implications based on RQ1 are valuable. However, the current version overclaims in ways that would mislead readers, so the revision must be substantive rather than cosmetic."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the qualitative mental model work is the real contribution; the transparency experiment is confounded by the fixed condition order, so don't let RQ2 drive the citation.\n\nThe paper does something genuinely new: it elicits mental models of generative conversational search from 16 experienced users, using interviews, think-aloud, and expectation-violation surveys. The RQ1 analysis is rich and believable. The finding that users' models are too abstract to explain individual search instances, leading to contradictory notions (corpus vs database) and hybrid web-CA workflows, is well supported by the quotes and the thematic coding. The authors also connect their work to prior mental model research on web search and voice assistants. That part deserves to be published.\n\nThe soft spot is exactly what the stress-test flags. Section 3.5.1 says condition order was fixed from baseline to most-transparent. That means transparency is perfectly correlated with session position. The satisfaction means (5.00, 4.60, 3.38, 4.33) and the pattern of expectation violation could be fatigue or growing skepticism, and the transcript itself contains evidence of that (e.g., P2 becoming 'a little bit sceptical' after an irrelevant link in a later condition). The paper's claim that transparency 'drove down satisfaction' is not supported. It also weakens the claim that only participants with richer mental models connected sources to online search, because those participants had already seen the system operate in earlier conditions. The Limitations section does not mention this confound, which is a genuine omission.\n\nI'd also note that the RQ2 survey analysis is purely descriptive; no inferential statistics. That's fine for an exploratory study, but the prose overstates what the data can support. The clustering analysis is a nice idea, but the cluster interpretation is post-hoc and not robust with n=16.\n\nThe citation pattern is appropriate; they cite the relevant prior work on mental models and conversational search. The ecological validity choice (real search API results) is a strength but also introduces variability, which they acknowledge. Sample is small and homogeneous, but they state that.\n\nBottom line: a solid qualitative paper with an overreaching quantitative claim. A serious referee should engage with it; the fix is to either counterbalance condition order in a follow-up or to reframe RQ2 as exploratory and explicitly discuss the ordering confound. I'd accept it for review with the expectation of substantial revision. For my own work, I'd cite the RQ1 findings on mental models, but not the transparency conclusions.","headline":"Qualitative mental model work is solid and worth reading; the transparency experiment is fatally confounded by fixed condition order, so treat RQ2 conclusions with caution.","tokens_in":20421,"tokens_out":2185,"would_cite":true,"duration_ms":20645,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that most users' mental models of generative conversational search are too abstract to interpret any single search instance, and that adding transparency cues does not repair this gap, instead stabilizing expectation…","keywords":["mental models","conversational search","generative conversational agents","interface transparency","expectation violation","search satisfaction","retrieval augmented generation","trust calibration"],"falsifier":"Run the same four conditions with the order counterbalanced across participants: if the downward satisfaction trend and the failure of transparency to improve learning disappear when order is randomized, the RQ2 conclusions are artifacts of the fixed sequence rather than real effects.","tokens_in":19559,"feed_emoji":"💬","tokens_out":6002,"duration_ms":57588,"temperature":0.7,"pith_summary":"The paper tries to establish that people who regularly use generative conversational search -- chat-based AI tools that retrieve and synthesize information -- typically have mental models that are too abstract to explain any one answer the system gives, and that adding transparency widgets to the interface does not close that gap. It reports a study of 16 experienced users who completed four search tasks on four retrieval-augmented chat interfaces, from no transparency cues to three cues: source links, query transformations, and a faithfulness flag. The authors argue that mental models, rather than interface opacity alone, are the main barrier to appropriate trust, and that users' already-common habit of combining chat search with web search is a promising direction for future design.","feed_headline":"Most users can't explain a single AI chat search answer","feed_subtitle":"Small study: transparency widgets like source links did not close the gap; people fell back on hybrid web-and-chat workflows.","key_machinery":"The study is carried by the mental-model construct, organized using a four-part framework -- components, functions, attributes, and feelings -- and elicited through semi-structured interviews, think-aloud protocols, pre/post expectation-violation ratings, and logs of query repairs. The transparency manipulation is implemented as four retrieval-augmented generation (RAG) chat interfaces that differ only in which of three textual vectors are shown: source attribution, query transformation, and a yes/no faithfulness flag. The working mechanism is that these vectors are meant to act as interaction cues that users can fold into their mental models, but the paper observes that a cue only teaches when the user already has enough of a model to interpret it.","core_discovery":"On the paper's own terms, the central finding is that most users have a globally plausible picture of generative conversational search -- an LLM trained on internet data that takes a natural-language query and produces a response -- but cannot apply that picture to interpret a particular result. That leaves room for contradictory beliefs (the same participants sometimes described the system as an unstructured corpus and sometimes as a searched database), for heavy reliance on imagined limitations, and for volatile trust rather than stable over- or under-trust. The transparency manipulation did not reliably teach: source links were noticed and liked but did not convey online retrieval unless users already knew the system could search online; query transformations were treated as reusable search queries rather than evidence of conversation history; and even the most-transparent condition did not improve learning for users with incomplete mental models. Expectation violations stabilized after the first transparency addition, but satisfaction trended downward, which the authors interpret as transparency exposing system errors and limitations rather than repairing understanding.","pith_inferences":["Editorial inference: a longitudinal version of this study might find that transparency pays off only after repeated exposure, so the single-session satisfaction decline observed here may understate the long-term value of the cues.","Editorial inference: a natural next experiment would hold the interface constant and vary a short tutorial explaining what sources, query transformations, and faithfulness mean; if mental-model accuracy rises, the bottleneck is missing knowledge rather than missing cues.","Editorial inference: transparency could be made adaptive -- for example, revealing query transformations only after the user demonstrates an understanding of what a query is -- which would directly test the paper's claim that cues require prior model completeness.","Editorial inference: the hybrid-workflow result suggests trust in conversational search is best modeled as a choice between two complementary tools rather than as a property of the chat system alone."],"forward_implications":["Source links alone will not teach users that a chatbot can retrieve live web results; users may see them as a verification aid or even as possibly fake.","Showing query transformations and faithfulness flags can lower satisfaction when they surface system errors, even while stabilizing the user's expectations.","Users with abstract mental models compensate by building hybrid workflows, so designs that make switching between chat and web search seamless would support behavior users already exhibit.","Mental-model incompleteness, not just interface opacity, should be treated as the trust problem; interventions such as onboarding or pre-interaction explanation may be needed.","The same interface can produce over-trust in some users and under-trust in others, so studies of conversational search should measure trust volatility rather than only average trust."],"supporting_citations":[{"why":"Supplies the four-part mental-model framework (components, functions, attributes, feelings) used to structure the interviews and analysis.","marker":"[69]"},{"why":"Established that showing query transformations improves mental models of web search; the study adapts this transparency vector to conversational agents.","marker":"[49]"},{"why":"Explored source attribution, confidence scores, and limitations in conversational search explanations; the least-transparent and faithfulness designs build on it.","marker":"[37]"},{"why":"Showed that confidence highlighting reduces overreliance in LLM-based search, providing the baseline expectation for what transparency should achieve.","marker":"[63]"},{"why":"Documented how opacity in voice assistants widens the gap between user expectations and system reality, motivating the transparency manipulation.","marker":"[44]"},{"why":"Reported that LLM-based conversational agents elicit inaccurate and incomplete mental models; the direct predecessor for studying these models in search.","marker":"[72]"},{"why":"Provided the expectation-violation self-report method used before and after each search task.","marker":"[18]"},{"why":"Defined key properties of conversational search including system revealment, framing the transparency and memory-related design discussion.","marker":"[57]"}],"fun_headline_variants":["Users can't explain why AI chat gave that answer","Source links don't fix AI chat confusion","Mental models of AI search are too vague to predict results","Users fall back to hybrid search when AI chat is opaque","Confusion about AI chat undermines trust in answers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the four interfaces differ only in the transparency vectors and that the fixed order of conditions, always from baseline to most-transparent, did not itself cause the observed drift in satisfaction and expectations.","fun_headline_variants_meta":{"raw":{"variants":["Users can't explain why AI chat gave that answer","Source links don't fix AI chat confusion","Mental models of AI search are too vague to predict results","Users fall back to hybrid search when AI chat is opaque","Confusion about AI chat undermines trust in answers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1462,"prompt_tokens":890,"completion_tokens":572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":496}},"tokens_in":506,"tokens_out":572,"duration_ms":5744,"temperature":1.0,"reasoning_tokens":496,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:54:38.552986+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four conditions with the order counterbalanced across participants: if the downward satisfaction trend and the failure of transparency to improve learning disappear when order is randomized, the RQ2 conclusions are artifacts of the fixed sequence rather than real effects.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established that showing query transformations improves mental models of web search; the study adapts this transparency vector to conversational agents."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defined key properties of conversational search including system revealment, framing the transparency and memory-related design discussion."}],"review_version":1}