{"id":"6935e94e-2e35-4b71-ae6e-b37d2734f4bc","arxiv_id":"2502.06549","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A 58-participant survey found users favor short text explanations and that app-specific knowledge, confidence, and demographics are weak predictors of explanation preferences.","lead":"This paper surveys 58 software users and finds they mostly prefer short, moderately detailed text explanations, with app-specific knowledge and demographics showing little direct influence on that preference. The authors suggest these weak correlations complicate the idea of automatically tailoring explanations to user expertise.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'only weakly related' null is underpowered: with n≈55, non-significant correlations are compatible with moderate true effects; equivalence bounds and confidence intervals are needed and are not reported.","rationale":"The reader's weakest-assumption concern about construct validity is legitimate and important. My stress test identifies a closely related but more decisive issue: even if the measures were valid, the study's sample size and lack of equivalence testing mean the central null is underpowered. Non-significant correlations with n≈55 are compatible with moderate true associations, so the conclusion that app-specific knowledge is 'only weakly related' is not statistically justified. The authors deserve credit for transparently reporting threats to validity and for making the data and code publicly available; the paper also provides reasonable descriptive evidence that users prefer moderately detailed short-text explanations. But the load-bearing claim about weak relationships is used to advise against knowledge-adaptive explanation systems, and that advice is not supported by the reported hypothesis tests alone. The existing CONDITIONAL verdict is appropriate: the paper should be accepted only with a mandatory reanalysis reporting confidence intervals and equivalence tests, and with the scale-assumption issue for Eform addressed. If the reanalysis shows that the data cannot rule out moderate effects, the central contribution would shift from a null finding to an inconclusive pilot study, which would lower the paper's practical value for adaptive explanation design.","tokens_in":12330,"tokens_out":4429,"duration_ms":45185,"concrete_test":"Re-analyze the Zenodo dataset: report 95% bootstrap or Fisher-z confidence intervals for every Spearman r and chi-square effect size in Table 5, and run a TOST equivalence test with equivalence bounds of |r|=0.20 (small-to-moderate). If the confidence intervals for H2.30/H2.40 (objective knowledge vs. detail level) exclude |r|>=0.30 and the TOST is significant, the 'weakly related' claim is supported. If the intervals include ±0.30 or the TOST is not significant, the study only supports 'not detected,' not 'weakly related.' Additionally, run a known-groups validation of the knowledge quiz and preference scales (e.g., expert vs. novice users); if the groups do not separate on the measures, the null may reflect measurement artifact rather than a true absence of relationship.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Section 5.1, RQ1) is that app-specific knowledge is 'only weakly related' to preferred explanation form or detail level. This is a null claim, but the study treats failure to reject as evidence of absence. With n=57 and a Bonferroni-corrected alpha of 0.0125, Spearman correlations up to roughly |r|=0.33 are not significant, and the 95% confidence interval for an observed r of 0 is approximately ±0.27. The reported non-significant correlations (e.g., H2.30: r=-0.07; H2.40: r=-0.03; H3.30: r=-0.03) therefore have wide intervals that include moderate true effects. The paper does not report confidence intervals or equivalence tests (e.g., TOST), so 'weakly related' is an overinterpretation. The problem is compounded by the outcome measures: preferred detail level and form are single-item, hypothetical preference scales developed without pilot validation, and Eform is declared nominal but is analyzed with Spearman in Table 5, mixing scale assumptions. The construct-validity threat acknowledged in Section 5.4 is not merely a limitation; it directly undermines the central null because measurement attenuation could produce exactly the weak correlations observed. A significant moderate correlation in Office form (H2.10, r=-0.49) is dismissed as not generalizable, but no statistical bound is provided to justify the dismissal. The data and code are public, which makes these checks feasible, but as reported the central null is not statistically established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an online survey (n=58) investigating whether users' app-specific knowledge, software confidence, and demographics relate to their preferred explanation detail level and form for Office and Browser software. The authors report that participants generally prefer moderately detailed short-text explanations, that subjective and objective app-specific knowledge correlate strongly, and that most hypothesized relationships with explanation preferences are non-significant. The central claim is that app-specific knowledge is only weakly related to preferred explanation form/detail, so adaptive explanation systems should not rely heavily on knowledge or demographics. The paper includes public data/code and discusses threats to validity.","tokens_in":12646,"tokens_out":3270,"duration_ms":29457,"significance":"If the central claim were statistically established, the result would be practically useful for requirements engineering and explainability design, particularly for deciding whether user knowledge or demographic data can drive adaptive explanations. The paper's strengths are its concrete survey instrument, two app categories, public data/code, and transparent discussion of validity threats. However, the significance is currently limited by the small convenience sample, unvalidated self-built measures, and inferential issues that make the 'weakly related' null claim unsupported as stated.","major_comments":[{"comment":"The treatment of H4 is internally contradictory. The text states that 'H4 was not rejected, as its p-value is 0.0042, thus falling below the required significance level for rejection,' but p=0.0042 is below 0.05 and would imply rejection. Moreover, this p-value does not appear in Table 5, where all reported demographic p-values are much larger (e.g., 0.44, 0.14, 0.95). This inconsistency directly undermines the RQ3 conclusion and the abstract's claim about demographic influences, and it must be corrected with the actual aggregate test and a consistent interpretation.","section":"Section 4.3 / Table 5 / H4"},{"comment":"The central 'only weakly related' conclusion is an overinterpretation of non-significant correlations. With n≈53–57 and a Bonferroni-corrected alpha of 0.0125, Spearman correlations up to roughly |r|=0.33 can be non-significant, and the 95% confidence interval for an observed r of 0 spans approximately ±0.27. No confidence intervals or equivalence tests (e.g., TOST) are reported, so failure to reject does not establish 'no relationship' or 'weak relationship.' The significant Office-form correlation (H2.10, r=-0.49) is dismissed as not generalizable without a statistical bound, yet it is the strongest evidence against the paper's weak-relationship claim. The authors should report effect-size confidence intervals and/or equivalence bounds, or explicitly reframe the conclusion as exploratory.","section":"Section 5.1 / Table 5"},{"comment":"The construct-validity threat acknowledged in Section 5.4 is directly load-bearing for the central null. The preferred detail level and form are single-item hypothetical preference scales, the objective knowledge quiz was developed by two researchers without third-party review, and participants did not interact with real explanations. Table 2 declares Eform as nominal, yet Table 5 analyzes it with Spearman correlation coefficients, mixing scale assumptions. These measurement issues can attenuate or distort correlations, so the observed weak/null pattern may be a methodological artifact rather than a true absence of relationship. The authors should either provide validity evidence (e.g., pilot data, item analysis, or comparison with observed behavior) or temper the central claim to an exploratory finding.","section":"Section 5.4 / Table 2"},{"comment":"The abstract and conclusion claim that demographic aspects (like gender) influence app-specific knowledge and that app-specific knowledge correlates with application confidence, suggesting a possible mediated relationship. However, Table 5 contains no tests of demographics against app-specific knowledge or confidence; it only tests demographics against explanation form and detail level. These claims are not supported by any reported analysis and should be removed or substantiated with the missing correlation/mediation analyses.","section":"Abstract / Section 6"}],"minor_comments":[{"comment":"The sentence 'This study investigates factors influencing users' preferred the level of detail' contains a typo ('preferred the level' should be 'preferred level').","section":"Abstract"},{"comment":"The notation H10, H20, H30, H40 is easily confused with subhypotheses like H1.10, H2.10, etc.; consider renaming the aggregate hypotheses (e.g., H1, H2, H3, H4) to improve readability.","section":"Section 3.4 / Table 3"},{"comment":"The text reports a p-value of 0.00009 for H2.10, while Table 5 lists p=0.00; the table should use a consistent number of decimal places or a p<0.001 notation.","section":"Section 4.3 / Table 5"},{"comment":"The interpretation that the strong subjective-objective knowledge correlation validates the knowledge construct is overstated; a strong correlation between two self-report or quiz-based measures could reflect shared method variance, and this should be acknowledged.","section":"Section 5.2"},{"comment":"The threats-to-validity section is thoughtful, but it would be strengthened by explicitly stating that the sample size of 58 and the wide confidence intervals prevent the conclusion that non-significant findings constitute evidence of absence.","section":"Section 5.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about several threats to validity, and the public data/code are a plus. The core issue is that the central null claim is not statistically established under the current analysis, and the H4 contradiction suggests a lapse in result reporting. Because the missing analyses (confidence intervals, equivalence tests, and the demographic-knowledge tests) are feasible with the public data, a major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a modest survey paper with one genuinely new dataset and a transparent write-up, but the central 'weakly related' conclusion is not statistically supported, and the abstract overreaches. Worth reading as an example of how to (and how not to) report null results.\n\nWhat's actually new: the combination of objective and subjective app-specific knowledge against preferred explanation detail and form. The public data and code on Zenodo are a real plus. The H1 result (subjective and objective knowledge correlate around r=0.6) is a nice validation of self-assessment. The preference for moderately detailed, short text explanations replicates earlier work. The threats-to-validity section is honest and does not hide the main weaknesses.\n\nThe soft spots, in order of severity:\n\n1. The null is underpowered. With n≈55 and Bonferroni-correlated alpha around 0.0125, non-significant correlations up to |r|=0.33 are plausible. The 95% CI for r=0 is roughly ±0.27. So 'no effect' or 'weakly related' in Table 5 and RQ1 is an overinterpretation. The paper reports no confidence intervals and no equivalence tests. The one significant moderate correlation (H2.10, r=-0.49 in Office form) is dismissed as 'not generalizable' without any statistical bound. The honest conclusion is 'mixed' or 'inconclusive,' not 'weakly related.'\n\n2. The abstract and Section 6 claim demographic influence on knowledge and confidence, and mention a 'possible mediated relationship,' but no mediation analysis was performed. Those claims should be removed or tested.\n\n3. Internal contradiction: Section 4.3 says H4 was 'not rejected, as its p-value is 0.0042, thus falling below the required significance level for rejection.' That sentence is backwards. The aggregate p-value is not in Table 5, and the text later says H4 was not rejected. Needs fixing.\n\n4. Scale mismatch: Eform is declared nominal in Table 2 but analyzed with Spearman in Table 5. If the form scale is ordinal, say so; if nominal, use chi-square. This affects several tested hypotheses.\n\n5. Construct validity is a real threat, not just a footnote: the self-built quiz and single-item preference scales were not piloted or third-party reviewed. Measurement attenuation could produce exactly the nulls observed. That said, the authors acknowledge this clearly.\n\nThe paper is honest and the data are public, which makes the missing checks feasible. I would not cite the null as evidence, but the dataset and the H1 validation are potentially useful for requirements engineering researchers.\n\nRecommendation: deserves a serious referee, but expect major revision. The authors should report confidence intervals or equivalence bounds, fix the overclaims, clarify the Eform scale, and correct the H4 paragraph. I'd like to see a revised version.\n\nBest,\n[You]","headline":"Useful new dataset and honest limitations, but the central null is underpowered and the abstract overclaims; a revised version could be solid.","tokens_in":13138,"tokens_out":3050,"would_cite":false,"duration_ms":28030,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Users' app knowledge only weakly shapes which software explanations they prefer.","keywords":["explainability","software explanations","app-specific knowledge","explanation preferences","level of detail","explanation format","user survey","adaptive explanations"],"falsifier":"Present users with real explanation texts that vary in length and format, let them choose or use them in a working application, and compare choices with objective knowledge scores; if users with high app knowledge consistently choose shorter or more technical explanations while novices choose longer tutorial-style ones, the paper's claim of only a weak relationship would be overturned.","tokens_in":12167,"feed_emoji":"📊","tokens_out":8256,"duration_ms":67913,"temperature":0.7,"pith_summary":"Software is becoming harder to understand, and one proposed answer is to adapt in-app explanations to each user. This paper asks whether a user's app-specific knowledge can predict how detailed an explanation they want and in what form, so that help content could be tailored automatically. Based on an online survey of 58 users of office and browser software, the authors find that objective knowledge correlates strongly with self-assessed knowledge but only weakly with preferred explanation form and not at all with preferred detail level. The most common choice was short text explanations at moderate detail. The paper concludes that explanation preferences are largely subjective, so adaptive explanation systems and user personas should not treat knowledge scores or demographics as reliable predictors of explanation needs.","feed_headline":"App knowledge barely predicts which software explanations users want","feed_subtitle":"Survey of 58 users finds knowledge scores and demographics alone can't predict explanation needs.","key_machinery":"The argument runs on three measurement instruments developed for the survey: a five-level explanation detail scale with written examples, a five-level explanation form scale, and an objective six-question quiz per app category whose scores are mapped to five knowledge levels. Relationships are tested with rank correlation for ordinal or metric variables and chi-square tests for nominal variables, using a multiple-comparison correction to limit false positives. The strong correlation between self-assessed and objective knowledge acts as an internal validity check for the knowledge construct, while the office-versus-browser comparison tests whether effects generalize across application categories.","core_discovery":"The paper's central claim is that app-specific knowledge, measured by an objective quiz, is only weakly related to the explanation form and detail level users prefer. In its survey of 58 participants across two app categories, office software and browsers, self-assessed knowledge matched objectively quizzed knowledge strongly, but this knowledge did not predict desired explanation detail in either category. Knowledge did correlate moderately, negatively, with preferred explanation form in office software, meaning more knowledgeable users leaned toward simpler forms, while the same effect did not appear for browsers. Confidence in using software behaved similarly, and demographic factors showed no significant relationship with either form or detail. The authors interpret the overall pattern as evidence that explanation preferences are shaped more by individual subjectivity than by knowledge, and that default explanations should be short, moderately detailed texts.","pith_inferences":["Beyond the paper: stated preferences may not match actual behavior, so an A/B test in a real application could reveal stronger knowledge effects than the hypothetical scales did.","Beyond the paper: the authors leave open whether task stakes, rather than user traits, drive explanation needs; comparing high-stakes and routine tasks is a direct test of that possibility.","Beyond the paper: if the weak correlations hold, adaptive explanation systems may need to learn preferences from interaction history instead of static user profiles.","Beyond the paper: the strong subjective-objective knowledge correlation suggests the null result is not simply a flawed knowledge quiz, leaving the unvalidated preference scales as the main measurement threat."],"forward_implications":["Adaptive explanation systems should not infer a user's preferred detail level from app-knowledge scores, because the survey found no correlation with detail level in office or browser software.","Self-assessed knowledge can serve as a practical proxy for objective knowledge in user studies, given the strong correlation the survey found between the two.","Default help content should favor short text at moderate detail, since that was the most common preference in both app categories.","Persona-based requirements analysis should treat app-specific knowledge and confidence as weak, non-deterministic inputs rather than fixed predictors.","Because browser users leaned toward even less detail than office users, explanation defaults may need to vary by application category."],"supporting_citations":[{"why":"Supplies the premise that lengthy explanations increase cognitive load, motivating the study's focus on preferred detail and form.","marker":"[6]"},{"why":"Earlier work found that users with different prior knowledge require different explanations; this study tests how knowledge relates to preferences.","marker":"[9]"},{"why":"Provides the survey design principles that the study follows in constructing the instrument.","marker":"[17]"},{"why":"Previous finding that explanation preferences are highly subjective and weakly tied to mood, which this study extends to app knowledge.","marker":"[24]"},{"why":"Personas for software explainability, the user-centered method the findings are meant to inform.","marker":"[26]"},{"why":"Taxonomy and automated detection of explanation needs in app reviews, establishing that explanation needs vary individually.","marker":"[28]"},{"why":"Supplies the multiple-comparison correction applied when testing the hypothesis families.","marker":"[14]"}],"fun_headline_variants":["App knowledge fails to predict software explanation needs","Users' app knowledge doesn't drive explanation preferences","Why app expertise may not dictate explanation format","Knowledge-poor link: app know-how and explanation choices","Survey finds app knowledge weakly shapes explanation wants"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes its self-built quiz and hypothetical preference scales measure real app-specific knowledge and real explanation preferences, even though the authors note the scales were designed by two researchers without third-party review and participants never interacted with actual explanations.","fun_headline_variants_meta":{"raw":{"variants":["App knowledge fails to predict software explanation needs","Users' app knowledge doesn't drive explanation preferences","Why app expertise may not dictate explanation format","Knowledge-poor link: app know-how and explanation choices","Survey finds app knowledge weakly shapes explanation wants"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000137,"raw_usage":{"total_tokens":1147,"prompt_tokens":936,"completion_tokens":211,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":141}},"tokens_in":552,"tokens_out":211,"duration_ms":4071,"temperature":1.0,"reasoning_tokens":141,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:04:34.342398+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Present users with real explanation texts that vary in length and format, let them choose or use them in a working application, and compare choices with objective knowledge scores; if users with high app knowledge consistently choose shorter or more technical explanations while novices choose longer tutorial-style ones, the paper's claim of only a weak relationship would be overturned.","supporting_citations":[{"cited_title":"REJ 25(4) (2020)","cited_arxiv_id":null,"evidence_quote":"Supplies the premise that lengthy explanations increase cognitive load, motivating the study's focus on preferred detail and form."},{"cited_title":"In: EASE’23","cited_arxiv_id":null,"evidence_quote":"Earlier work found that users with different prior knowledge require different explanations; this study tests how knowledge relates to preferences."},{"cited_title":"SIGSOFT Softw","cited_arxiv_id":null,"evidence_quote":"Provides the survey design principles that the study follows in constructing the instrument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Previous finding that explanation preferences are highly subjective and weakly tied to mood, which this study extends to app knowledge."},{"cited_title":"In: HCI-COLLAB’21","cited_arxiv_id":null,"evidence_quote":"Personas for software explainability, the user-centered method the findings are meant to inform."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Taxonomy and automated detection of explanation needs in app reviews, establishing that explanation needs vary individually."},{"cited_title":"Springer, New York, NY, USA (2013)","cited_arxiv_id":null,"evidence_quote":"Supplies the multiple-comparison correction applied when testing the hypothesis families."}],"review_version":1}