{"id":"5f72cb9d-d50e-4492-9e2c-febd42857b89","arxiv_id":"2608.06091","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"In a laptop-buying chatbot study, novices rated the format combining technical specs, performance categories, and plain-language explanations as more helpful and appropriate than specs alone, while expert ratings did not differ across formats.","lead":"A user study with 251 participants tested how a chatbot sales advisor should present laptop specs to novices versus experts. Novices felt that plain-language explanations helped them learn and judged the combined format most appropriate, while experts were indifferent to the added detail.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"TE/TCE bundle explanatory sentences with plain-language terminology, so the 'explanations are crucial' mechanism is confounded and the central design guidelines are not uniquely supported.","rationale":"The reader's weakest assumption is exactly the confound I identify: TE/TCE change both explanations and terminology, so the mechanism is not cleanly identified. This is load-bearing because the paper's headline guidelines—particularly 'use TCE by default' and 'avoid standalone categories'—are explicitly justified by the claim that explanations are the crucial component. If plain-language rewording alone drives the novice benefits, the guidelines are overfitted to the specific stimulus set. I do not see this as a fatal flaw: the descriptive results are credible, the statistical analysis is appropriate with BH correction, effect sizes are reported, and open materials are available. The confound is addressable with a follow-up experiment or a targeted additional condition, which is why I recommend keeping the reader's CONDITIONAL verdict rather than escalating to rejection. I also considered the expert null-result issue (no significant differences used to infer no harm), but that is a weaker concern because the paper's wording is 'measurably' and the descriptive claim about experts is appropriately cautious; the terminology confound is the more central threat to the paper's unique contribution. My concrete test is designed to settle whether the explanation mechanism holds or whether terminology simplification is sufficient.","tokens_in":18456,"tokens_out":2784,"duration_ms":26090,"concrete_test":"Run a preregistered follow-up with a crossed 2x2 design holding performance categories constant: (1) technical terms, no explanation; (2) plain terms, no explanation; (3) technical terms, with explanatory sentence; (4) plain terms, with explanatory sentence, using the same Cleo script and outcome measures. If cell (2) yields perceived-learning and helpfulness ratings statistically indistinguishable from cells (3) and (4), while cell (3) does not differ from cell (1), the explanatory sentence is not the active ingredient and the guidelines must be revised. A lower-cost alternative: add a single 'plain terms only' condition to the existing OSF materials and compare it against the existing TE and TCE data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central mechanism claim is that attribute explanations, not merely simpler wording, drive novices' benefits and that 'explanations are crucial to understand and benefit from performance categories.' This is not identifiable from the reported design. In Section 3.2.3, TE is described as adding a plain-language explanatory sentence AND replacing technical abbreviations with less specialized terms; Table 1 shows the same for TCE (e.g., 'RAM' appears as 'RAM (working memory)'). Thus every novice comparison that isolates explanations (TE vs. T, TCE vs. TC) simultaneously varies two factors: presence of an explanatory sentence and presence of plain-language definitions. The observed gains in perceived learning and helpfulness, and the lower Attribute Confusion in qualitative data, could be driven entirely by the simplified terminology rather than by the explanatory content. If so, the design guideline 'use TCE by default' and the specific injunction to 'avoid standalone categories' would not be supported by the data; a simpler package of plain terms plus categories might suffice. This confound does not undermine the descriptive finding that the TE/TCE packages outperformed T/TC for novices, but it does undermine the causal attribution that is central to the paper's contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects online experiment (n = 251) on how a rule-based text-based product advisor ('Cleo') should present laptop attribute information to users with different self-reported domain knowledge. Four presentation formats are compared: technical information only (T), technical information plus performance categories (TC), technical information plus attribute explanations (TE), and a combined format (TCE). Novices (self-rated knowledge 1–4) and experts (5–7) rated perceived appropriateness of information quantity, perceived learning, relevance, trust, and helpfulness, and answered one open question. The main findings are that novices rate TE and TCE as more helpful and more conducive to perceived learning than T and TC, novices rate TCE as more appropriate in information quantity than T and TC, and experts show no significant differences across conditions. The authors derive four design guidelines: use TCE by default, keep a single inclusive interface, avoid standalone categories, and support user agency and personalization.","tokens_in":18596,"tokens_out":3481,"duration_ms":30211,"significance":"If the results hold, the paper offers practically useful guidance for conversational commerce: simple, scripted additions to technical product information can improve novices' perceived learning and information-appropriateness without measurably harming experts. The study has notable strengths: the expertise split threshold was set a priori using an archival dataset, stratified random assignment was used, materials and anonymized data are available on OSF, and the statistical analysis is appropriate for ordinal Likert data, using Kruskal-Wallis tests with Dunn's post-hoc tests, Benjamini-Hochberg correction, Mann-Whitney U tests, and Cliff's delta effect sizes. The central caveat is that the causal interpretation — that explanations specifically, rather than simpler terminology, drive novices' benefits — is not identifiable from the reported design, and several design guidelines depend on that interpretation.","major_comments":[{"comment":"The TE and TCE conditions change two factors at once relative to T and TC: they add a plain-language explanatory sentence, and they replace technical abbreviations with less specialized terms (for example, 'RAM' becomes 'RAM (working memory)' in Table 1). Consequently, every contrast that is used to attribute novices' higher perceived learning or helpfulness to explanations (TE vs. T, TCE vs. TC) also varies terminology simplicity. The observed benefits could be driven entirely by the simplified wording, which would undermine the mechanistic claim in the abstract and in Section 6 that 'explanations are crucial to understand and benefit from performance categories,' as well as design guideline 3 ('avoid standalone categories'). To support the causal attribution, the authors would need an additional condition that provides plain-language terminology without explanatory sentences, or a condition that holds terminology constant while varying only the explanatory content. As it stands, the manuscript should either add such a condition or substantially reframe the conclusions to describe the benefits of the whole TE/TCE package rather than of explanations specifically.","section":"§3.2.3 and Table 1"},{"comment":"The claim that the attribute-recommendation accuracy measure 'ensured that the perceived plausibility of Cleo's recommendations did not confound participants' evaluations' is not supported by the reported analysis. The section only reports overall means and standard deviations for novices and experts, with no statistical comparison across the four conditions or between the two expertise groups, and the 66 'I don't know' responses are excluded without justification. Without showing that accuracy ratings did not differ by condition (or at least by condition within expertise group), the assertion that accuracy was not a confound is not established. Report inferential tests for accuracy ratings across conditions and groups, and explain how the 'I don't know' responses were handled.","section":"§5.1.6"},{"comment":"The evidence for the specific guideline 'avoid standalone categories' is weaker than the presentation suggests. While novices rated TCE higher than TC in appropriateness, the direct comparison between TC and T on appropriateness was not significant (p_adj > .05), and most other TC-versus-T comparisons were not significant. The qualitative 'Attribute Confusion' difference between TC (60%) and TCE (34%) is also affected by the terminology confound described above, since TCE uses 'RAM (working memory)' while TC uses 'RAM.' The guideline is therefore reasonable as an exploratory design suggestion but should not be stated as a firm empirical conclusion without disentangling the terminology effect.","section":"§5.1.1 and §5.2"}],"minor_comments":[{"comment":"The caption mentions 'black and red brackets' but does not define in the caption which test each color corresponds to; please add a one-line explanation, e.g., black = Dunn's post-hoc, red = Mann-Whitney U.","section":"Figure 2"},{"comment":"Cliff's delta is attributed to reference [63], which is a paper on machine learning and software development; the canonical source (Cliff, 1993) should be cited instead or additionally.","section":"References"},{"comment":"The quote from P17 contains the bracketed pronoun '[they were]' inside what appears to be a direct quotation; either use the original wording or use a standard ellipsis/bracketing convention for editorial insertions.","section":"§5.2"},{"comment":"The sentence describing the archival dataset says the threshold was 'specified before data collection' but does not state whether the archival dataset was used to set the split only, not to fit any outcome measure; clarifying this in the main text would preempt concerns about circularity.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The paper fits CHIIR's scope and the empirical work is generally well executed. The main issue is the confound between explanatory sentences and simplified terminology, which is load-bearing for the paper's central mechanism claim and for several design guidelines. I do not see this as grounds for rejection, because the descriptive results are useful and the claims can be reframed or supplemented with a follow-up experiment. I would encourage the editor to require either an additional disentangling condition or a substantial softening of the causal and guideline claims before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a solid empirical study of knowledge-level adaptation in a text-based product advisor. If you need one takeaway: the descriptive results are believable, but the paper's central mechanism claim—that attribute explanations drive novices' benefit—is confounded with simpler wording, so accept the data, treat the guidelines as tentative.\n\nWhat's new: prior chatbot personalization work focused on lexical/personality alignment; knowledge-level adaptation had prototypes but not much user evaluation. Here you get a clean between-subjects manipulation of four presentation formats (T, TC, TE, TCE) with 251 participants, an a priori novice/expert split based on an archival dataset, and appropriate non-parametric statistics with BH correction and effect sizes. The qualitative analysis is also careful: independent coders, codebook, negotiated agreement. Open data on OSF is a plus.\n\nSoft spots:\n\n1. The TE and TCE conditions change two things at once: they add an explanatory sentence and replace technical abbreviations with everyday synonyms (e.g., RAM becomes RAM (working memory)). So the observed gains over T/TC could come entirely from the simpler wording. The paper's claim that 'explanations are crucial to understand and benefit from performance categories' isn't identifiable from this design. This is a real confound, and it directly weakens guideline 3 (avoid standalone categories). The descriptive finding that the TE/TCE packages beat T/TC for novices is fine; the causal story is not.\n\n2. The accuracy check is described as 'ensuring' that recommendation plausibility didn't confound evaluations, but the paper only reports overall means (3.83 novices, 3.94 experts), not a test across conditions. That's an overstatement; minor.\n\n3. The expert non-detriment conclusion is inferred from null results. The authors are appropriately cautious in the discussion, so I'd call this minor.\n\nThe limitations section is honest about self-reported knowledge and the laptop-only domain. I'd rather see a 2x2 design (categories on/off, explanations on/off) to disentangle the mechanism, but that's a follow-up.\n\nWho it's for: the CHIIR crowd, conversational commerce researchers, and designers of inclusive product advisors. Practitioners can use the guidelines, but should treat the 'explanations, not just plain language' part as provisional.\n\nIt deserves serious peer review. I'd accept it conditionally: require the authors to either soften the causal claims or provide a re-analysis that acknowledges the confound, and fix the accuracy-check claim.","headline":"Useful empirical comparison of knowledge-level adaptation formats; descriptive results hold, but the 'explanations are crucial' mechanism is confounded with simpler wording, so the causal claims need softening.","tokens_in":19181,"tokens_out":2399,"would_cite":true,"duration_ms":18753,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Text-based product advisors should present technical specifications together with performance categories and plain-language explanations, because this combination helps novices learn and feel supported while leaving experts' experience…","keywords":["conversational commerce","product search","digital assistant","informed decision-making","personalization","domain knowledge","knowledge-level adaptation","user perception"],"falsifier":"Run the same laptop-advisor study with a crossed design: explanations on or off crossed with plain or technical terminology. If novices show the same benefit whenever terminology is plain, with no extra effect from the explanatory sentence, the paper's central mechanism fails; if the explanation effect persists even with technical terms, the claim survives.","tokens_in":18192,"feed_emoji":"🛒","tokens_out":6248,"duration_ms":47497,"temperature":0.7,"pith_summary":"This paper asks how a text-based product advisor should present complex technical information to shoppers with different levels of domain knowledge. In a chatbot-assisted laptop search experiment with 251 participants, it compares four formats: technical specifications alone, with performance categories, with attribute explanations, or with both. The central finding is that novices perceive formats containing attribute explanations (TE and TCE) as more helpful and report more learning, and they find the combined format (TCE) more appropriate in information quantity than plain specs or categories alone. Experts show no significant differences across any format, which the authors read as evidence that adding novice-friendly explanations does not harm experts. If correct, this supports a single inclusive interface that gives every user technical details, categories, and explanations.","feed_headline":"Chatbot explanations help novice shoppers, leave experts unharmed","feed_subtitle":"In a 251-person laptop-search test, adding plain-language explanations raised novices' learning and ratings without expert backlash.","key_machinery":"The key machinery is the four-condition information presentation format implemented in a rule-based chatbot (Cleo) for laptop searching: technical information only (T), plus performance categories (TC), plus attribute explanations (TE), or both (TCE). Performance categories are plain-language bands such as 'entry-level to mid-range configuration' for RAM; attribute explanations are one-sentence descriptions that also replace technical abbreviations with everyday terms (e.g., 'RAM (working memory)'). This manipulation isolates the type and amount of supplementary information presented with each attribute recommendation, and the comparison of novice and expert ratings across the four formats is what carries the argument.","core_discovery":"The paper's central claim is that knowledge-level accommodation in a text-based product advisor works best when technical attribute recommendations are supplemented with both performance categories and plain-language attribute explanations, and that this accommodation can be offered to everyone rather than hidden from experts. In the authors' data, novices rated the explanation-bearing conditions (TE and TCE) significantly more helpful than categories alone (TC) and reported significantly higher perceived learning than in T or TC; novices also rated TCE more appropriate than T and TC in information quantity. Experts exhibited no significant differences across T, TC, TE, and TCE on any dependent measure, which the authors interpret as showing that the extra information did not detract from expert experience. Qualitative responses reinforce the account: attribute confusion was novices' dominant concern (60% in T and TC, falling to 34% in TCE), while experts mainly asked for additional attributes such as price and brand rather than objecting to the supplements.","pith_inferences":["Extension: The TE condition changes two variables at once (it adds explanations and replaces technical abbreviations with everyday terms), so the paper's attribution of novice gains to 'explanations' is not fully isolated; a crossed design separating wording from explanation could test this.","Extension: The same presentation-format logic could be adapted to other technical consumer domains (smartphones, cameras, software plans) and to LLM-based advisors, where these formats could be used as prompt templates with explanation depth scaled by user signals.","Extension: The experts' indifference may be specific to short, sequential attribute messages; in denser or more interactive interfaces the expertise reversal effect might reappear, so 'inclusive by default' should be rechecked for information-heavy designs.","Extension: The 'perceived learning' measure is a self-report; an objective knowledge test before and after the interaction would show whether the higher perceived learning corresponds to actual comprehension gains."],"forward_implications":["A text-based product advisor for technical products should default to combining technical specifications, performance categories, and attribute explanations (TCE), since novices rate it more appropriate and more helpful.","Performance categories should not be presented alone: novices found TC no better than plain specs and rated it less appropriate than TCE, so explanations appear necessary for categories to be useful.","There is no empirical reason to strip supplementary information for experts in a single interface, because experts' perceptions did not differ across conditions.","In this domain, explanations primarily serve novices' learning and comprehension; designers should pair any category label with a short explanation of what the attribute does.","User agency and personalization remain wanted by both groups, so advisors should offer options and tailor to the stated use case rather than only controlling information type."],"supporting_citations":[{"why":"Supplies the consumer-expertise theory that motivates performance categories and attribute explanations as support for novices.","marker":"[2]"},{"why":"Provides communication accommodation theory, the theoretical basis for adapting chatbot language to user knowledge.","marker":"[20]"},{"why":"Defines the expertise reversal effect that the hypotheses about experts' possible aversion to extra information build on.","marker":"[51]"},{"why":"Documents novice versus expert search-behavior differences that motivate supporting low-knowledge consumers in product search.","marker":"[45]"},{"why":"Presents a conversational prototype for recipient design based on inferred user knowledge, which this study's user evaluation extends.","marker":"[4]"},{"why":"Shows language models can adapt output complexity to user familiarity, the capability the paper argues needs user-facing evaluation in commerce.","marker":"[57]"},{"why":"Provides a prior user study where tailoring dialogue to comprehension level improved understanding, the closest empirical precedent.","marker":"[22]"}],"fun_headline_variants":["Explanations help novice shoppers, experts unaffected","Add plain-talk explanations: novices learn, pros fine","Chatbot advice with explanations wins novices, spares experts","Inclusive product bot: explain attributes, novices gain, pros stay","One chatbot interface with explanations aids novices, not experts' loss"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes that the explanatory sentences, and not the simpler wording that came with them, are what helped novices.","fun_headline_variants_meta":{"raw":{"variants":["Explanations help novice shoppers, experts unaffected","Add plain-talk explanations: novices learn, pros fine","Chatbot advice with explanations wins novices, spares experts","Inclusive product bot: explain attributes, novices gain, pros stay","One chatbot interface with explanations aids novices, not experts' loss"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1337,"prompt_tokens":986,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":266}},"tokens_in":602,"tokens_out":351,"duration_ms":3911,"temperature":1.0,"reasoning_tokens":266,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T15:04:31.557323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same laptop-advisor study with a crossed design: explanations on or off crossed with plain or technical terminology. If novices show the same benefit whenever terminology is plain, with no extra effect from the explanatory sentence, the paper's central mechanism fails; if the explanation effect persists even with technical terms, the claim survives.","supporting_citations":[{"cited_title":"Alba and J","cited_arxiv_id":null,"evidence_quote":"Supplies the consumer-expertise theory that motivates performance categories and attribute explanations as support for novices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides communication accommodation theory, the theoretical basis for adapting chatbot language to user knowledge."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the expertise reversal effect that the hypotheses about experts' possible aversion to extra information build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents novice versus expert search-behavior differences that motivate supporting low-knowledge consumers in product search."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows language models can adapt output complexity to user familiarity, the capability the paper argues needs user-facing evaluation in commerce."},{"cited_title":"Goldsmith","cited_arxiv_id":null,"evidence_quote":"Provides a prior user study where tailoring dialogue to comprehension level improved understanding, the closest empirical precedent."}],"review_version":1}