REVIEW 4 major objections 5 minor 30 references
How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Users' app knowledge only weakly shapes which software explanations they prefer.
desk verdict Useful new dataset and honest limitations, but the central null is underpowered and the abstract overclaims; a revised version could be solid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on three measurement instruments developed for the survey: a five-level explanation detail scale with written examples, a five-level explanation form scale, and an objective six-question quiz per app category whose scores are mapped to five knowledge levels. Relationships are tested with rank correlation for ordinal or metric variables and chi-square tests for nominal variables, using a multiple-comparison correction to limit false positives. The strong correlation between self-assessed and objective knowledge acts as an internal validity check for the knowledge construct, while the office-versus-browser comparison tests whether effects generalize across application categories.
What would settle it
Present users with real explanation texts that vary in length and format, let them choose or use them in a working application, and compare choices with objective knowledge scores; if users with high app knowledge consistently choose shorter or more technical explanations while novices choose longer tutorial-style ones, the paper's claim of only a weak relationship would be overturned.
Extended reading notes
Core claim
The paper's central claim is that app-specific knowledge, measured by an objective quiz, is only weakly related to the explanation form and detail level users prefer. In its survey of 58 participants across two app categories, office software and browsers, self-assessed knowledge matched objectively quizzed knowledge strongly, but this knowledge did not predict desired explanation detail in either category. Knowledge did correlate moderately, negatively, with preferred explanation form in office software, meaning more knowledgeable users leaned toward simpler forms, while the same effect did not appear for browsers. Confidence in using software behaved similarly, and demographic factors showed no significant relationship with either form or detail. The authors interpret the overall pattern as evidence that explanation preferences are shaped more by individual subjectivity than by knowledge, and that default explanations should be short, moderately detailed texts.
Load-bearing premise
The study assumes its self-built quiz and hypothetical preference scales measure real app-specific knowledge and real explanation preferences, even though the authors note the scales were designed by two researchers without third-party review and participants never interacted with actual explanations.
Editorial extensions
If this is right
- Adaptive explanation systems should not infer a user's preferred detail level from app-knowledge scores, because the survey found no correlation with detail level in office or browser software.
- Self-assessed knowledge can serve as a practical proxy for objective knowledge in user studies, given the strong correlation the survey found between the two.
- Default help content should favor short text at moderate detail, since that was the most common preference in both app categories.
- Persona-based requirements analysis should treat app-specific knowledge and confidence as weak, non-deterministic inputs rather than fixed predictors.
- Because browser users leaned toward even less detail than office users, explanation defaults may need to vary by application category.
Reading between the lines
- Beyond the paper: stated preferences may not match actual behavior, so an A/B test in a real application could reveal stronger knowledge effects than the hypothetical scales did.
- Beyond the paper: the authors leave open whether task stakes, rather than user traits, drive explanation needs; comparing high-stakes and routine tasks is a direct test of that possibility.
- Beyond the paper: if the weak correlations hold, adaptive explanation systems may need to learn preferences from interaction history instead of static user profiles.
- Beyond the paper: the strong subjective-objective knowledge correlation suggests the null result is not simply a flawed knowledge quiz, leaving the unvalidated preference scales as the main measurement threat.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an online survey (n=58) investigating whether users' app-specific knowledge, software confidence, and demographics relate to their preferred explanation detail level and form for Office and Browser software. The authors report that participants generally prefer moderately detailed short-text explanations, that subjective and objective app-specific knowledge correlate strongly, and that most hypothesized relationships with explanation preferences are non-significant. The central claim is that app-specific knowledge is only weakly related to preferred explanation form/detail, so adaptive explanation systems should not rely heavily on knowledge or demographics. The paper includes public data/code and discusses threats to validity.
Significance. If the central claim were statistically established, the result would be practically useful for requirements engineering and explainability design, particularly for deciding whether user knowledge or demographic data can drive adaptive explanations. The paper's strengths are its concrete survey instrument, two app categories, public data/code, and transparent discussion of validity threats. However, the significance is currently limited by the small convenience sample, unvalidated self-built measures, and inferential issues that make the 'weakly related' null claim unsupported as stated.
major comments (4)
- [Section 4.3 / Table 5 / H4] The treatment of H4 is internally contradictory. The text states that 'H4 was not rejected, as its p-value is 0.0042, thus falling below the required significance level for rejection,' but p=0.0042 is below 0.05 and would imply rejection. Moreover, this p-value does not appear in Table 5, where all reported demographic p-values are much larger (e.g., 0.44, 0.14, 0.95). This inconsistency directly undermines the RQ3 conclusion and the abstract's claim about demographic influences, and it must be corrected with the actual aggregate test and a consistent interpretation.
- [Section 5.1 / Table 5] The central 'only weakly related' conclusion is an overinterpretation of non-significant correlations. With n≈53–57 and a Bonferroni-corrected alpha of 0.0125, Spearman correlations up to roughly |r|=0.33 can be non-significant, and the 95% confidence interval for an observed r of 0 spans approximately ±0.27. No confidence intervals or equivalence tests (e.g., TOST) are reported, so failure to reject does not establish 'no relationship' or 'weak relationship.' The significant Office-form correlation (H2.10, r=-0.49) is dismissed as not generalizable without a statistical bound, yet it is the strongest evidence against the paper's weak-relationship claim. The authors should report effect-size confidence intervals and/or equivalence bounds, or explicitly reframe the conclusion as exploratory.
- [Section 5.4 / Table 2] The construct-validity threat acknowledged in Section 5.4 is directly load-bearing for the central null. The preferred detail level and form are single-item hypothetical preference scales, the objective knowledge quiz was developed by two researchers without third-party review, and participants did not interact with real explanations. Table 2 declares Eform as nominal, yet Table 5 analyzes it with Spearman correlation coefficients, mixing scale assumptions. These measurement issues can attenuate or distort correlations, so the observed weak/null pattern may be a methodological artifact rather than a true absence of relationship. The authors should either provide validity evidence (e.g., pilot data, item analysis, or comparison with observed behavior) or temper the central claim to an exploratory finding.
- [Abstract / Section 6] The abstract and conclusion claim that demographic aspects (like gender) influence app-specific knowledge and that app-specific knowledge correlates with application confidence, suggesting a possible mediated relationship. However, Table 5 contains no tests of demographics against app-specific knowledge or confidence; it only tests demographics against explanation form and detail level. These claims are not supported by any reported analysis and should be removed or substantiated with the missing correlation/mediation analyses.
minor comments (5)
- [Abstract] The sentence 'This study investigates factors influencing users' preferred the level of detail' contains a typo ('preferred the level' should be 'preferred level').
- [Section 3.4 / Table 3] The notation H10, H20, H30, H40 is easily confused with subhypotheses like H1.10, H2.10, etc.; consider renaming the aggregate hypotheses (e.g., H1, H2, H3, H4) to improve readability.
- [Section 4.3 / Table 5] The text reports a p-value of 0.00009 for H2.10, while Table 5 lists p=0.00; the table should use a consistent number of decimal places or a p<0.001 notation.
- [Section 5.2] The interpretation that the strong subjective-objective knowledge correlation validates the knowledge construct is overstated; a strong correlation between two self-report or quiz-based measures could reflect shared method variance, and this should be acknowledged.
- [Section 5.4] The threats-to-validity section is thoughtful, but it would be strengthened by explicitly stating that the sample size of 58 and the wide confidence intervals prevent the conclusion that non-significant findings constitute evidence of absence.
Circularity Check
No significant circularity: empirical survey with external outcome measures and no fitted-input prediction chain.
full rationale
This paper is an empirical survey study with no mathematical derivation, fitted model, or uniqueness argument. The central claim—that app-specific knowledge is only weakly related to preferred explanation form and detail level (Section 5.1, RQ1)—is a null correlation result computed against external survey responses. The predictors (objective quiz score, self-assessed knowledge, confidence, demographics) and outcomes (preferred form and detail level) are independent measurements; neither is defined in terms of the other, and no parameter is fitted to the outcome and then renamed as a prediction. The H1 validation correlation between subjective and objective app knowledge is used as a manipulation check, not as an input to the preference correlations. Self-citations in the related work section are contextual and not load-bearing: for instance, [24] is a separate mood study and [23] is the dataset release. The paper explicitly acknowledges threats to construct validity, internal validity, and conclusion validity in Section 5.4, including the lack of third-party review of self-built scales, possible answer lookup, and Type II error risk from Bonferroni correction. These are legitimate methodological limitations, and the underpowered null interpretation could be criticized as statistically overconfident, but that is a correctness/statistical concern, not circularity. No quoted step reduces to its own input, so the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Objective knowledge category thresholds =
20% / 40% / 60% / 80%
assumptions (3)
- domain assumption The self-constructed quiz questions and ordinal scales are valid measures of app-specific knowledge and explanation preferences.
- domain assumption Participants' self-reported preferences for hypothetical explanation formats match their real-world preferences.
- standard math Nonparametric correlation tests on a sample of 53 to 57 participants can detect the relationships of interest.
Cite this review
Pith. "Pith review of How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?." pith.science (2026). https://pith.science/paper/P4LEYH5W
@misc{pith2026250206549,
author = {Pith},
title = {Pith review of: How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?},
year = {2026},
howpublished = {\url{https://pith.science/paper/P4LEYH5W}},
note = {Machine review of arXiv:2502.06549}
}
read the original abstract
Context and Motivation: Due to their increasing complexity, everyday software systems are becoming increasingly opaque for users. A frequently adopted method to address this difficulty is explainability, which aims to make systems more understandable and usable. Question/problem: However, explanations can also lead to unnecessary cognitive load. Therefore, adapting explanations to the actual needs of a user is a frequently faced challenge. Principal ideas/results: This study investigates factors influencing users' preferred the level of detail and the form of an explanation (e.g., short text or video tutorial) in software. We conducted an online survey with 58 participants to explore relationships between demographics, software usage, app-specific knowledge, as well as their preferred explanation form and level of detail. The results indicate that users prefer moderately detailed explanations in short text formats. Correlation analyses revealed no relationship between app-specific knowledge and the preferred level of detail of an explanation, but an influence of demographic aspects (like gender) on app-specific knowledge and its impact on application confidence were observed, pointing to a possible mediated relationship between knowledge and preferences for explanations. Contribution: Our results show that explanation preferences are weakly influenced by app-specific knowledge but shaped by demographic and psychological factors, supporting the development of adaptive explanation systems tailored to user expertise. These findings support requirements analysis processes by highlighting important factors that should be considered in user-centered methods such as personas.
Reference graph
Works this paper leans on
-
[1]
Andrade, H., Lwakatare, L.E., Crnkovic, I., Bosch, J.: So ftware challenges in het- erogeneous computing: A multiple case study in industry. In : SEAA (2019)
work page 2019
-
[2]
Antinyan, V.: Revealing the complexity of automotive sof tware. ESEC/FSE 2020
work page 2020
-
[3]
Brunotte, W., Specht, A., Chazette, L., Schneider, K.: Pr ivacy explanations–a means to end-user trust. JSS 195 (2023)
work page 2023
-
[4]
Buiten, M.C., Dennis, L.A., Schwammberger, M.: A vision o n what explanations of autonomous systems are of interest to lawyers. In: REW. IE EE (2023)
work page 2023
-
[5]
Chazette, L., Brunotte, W., Speith, T.: Exploring explai nability: a definition, a model, and a knowledge catalogue. In: RE. IEEE (2021)
work page 2021
-
[6]
Chazette, L., Schneider, K.: Explainability as a non-fun ctional requirement: chal- lenges and recommendations. REJ 25(4) (2020)
work page 2020
- [7]
-
[8]
In: REFSQ’24 (2 024) 16 Obaidi et al
Deters, H., Droste, J., Obaidi, M., Schneider, K.: How exp lainable is your system? towards a quality model for explainability. In: REFSQ’24 (2 024) 16 Obaidi et al
Show all 30 references
-
[9]
In: EASE’23
Deters, H., Droste, J., Schneider, K.: A means to what end? evaluating the ex- plainability of software systems using goal oriented heuri stics. In: EASE’23
-
[10]
In : RE4AI’24
Droste, J., Deters, H., Fuchs, R., Schneider, K.: Peekin g outside the black-box: Ai explainability requirements beyond interpretability. In : RE4AI’24
-
[11]
In: R E’24
Droste, J., Deters, H., Obaidi, M., Schneider, K.: Expla nations in everyday software systems: Towards a taxonomy for explainability needs. In: R E’24
-
[12]
Droste, J., Deters, H., Puglisi, J., Klünder, J.: Design ing end-user personas for explainability requirements using mixed methods research . In: REW. IEEE (2023)
2023
-
[13]
In: Shin, C.S., Di Bucchian- ico, G., Fukuda, S., Ghim, Y.G., Montagna, G., Carvalho, C
Gabbas, M., Ryu, Y., Park, J., Kim, K.: Understanding cha llenges of designing for complex users by adapting the existing framework. In: Shin, C.S., Di Bucchian- ico, G., Fukuda, S., Ghim, Y.G., Montagna, G., Carvalho, C. ( eds.) Advances in Industrial Design. Springer Interna...
2021
-
[14]
Springer, New York, NY, USA (2013)
Haynes, W.: Bonferroni Correction. Springer, New York, NY, USA (2013)
2013
-
[15]
Kästner, L., Langer, M., Lazar, V., Schomäcker, A., Spei th, T., Sterz, S.: On the relation of trust and explainability: Why to engineer for tr ustworthiness. REW’21
-
[16]
Kim, G., Yeo, D., Jo, T., Rus, D., Kim, S.: What and when to e xplain? on-road evaluation of explanations in highly automated vehicles. P roc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7(3) (sep 2023)
2023
-
[17]
SIGSOFT Softw
Kitchenham, B.A., Pfleeger, S.L.: Principles of survey r esearch part 2: designing a survey. SIGSOFT Softw. Eng. Notes 27(1), 18–20 (Jan 2002)
2002
-
[18]
In: RE’19
Köhl, M.A., Baum, K., Langer, M., Oster, D., Speith, T., B ohlender, D.: Explain- ability as a non-functional requirement. In: RE’19
-
[19]
Empirical Software Engineering 26(1) (2021)
Levy, O., Feitelson, D.: Understanding large-scale sof tware systems – structure and flows. Empirical Software Engineering 26(1) (2021)
2021
-
[20]
Compute r 45(08) (aug 2012)
Mens, T.: On the complexity of software systems. Compute r 45(08) (aug 2012)
2012
-
[21]
User Modeling an d User-Adapted In- teraction 27 (2017)
Nunes, I., Jannach, D.: A systematic review and taxonomy of explanations in decision support and recommender systems. User Modeling an d User-Adapted In- teraction 27 (2017)
2017
-
[22]
(Sep 2024)
Obaidi, M.: Dataset: Gold standard dataset for explaina bility need detection in app reviews. (Sep 2024). https://doi.org/10.5281/zenodo.11522828
2024 doi
-
[23]
https://doi.org/10.5281/zenodo.13175641
Obaidi, M.: Dataset: How Does Users’ App Knowledge Influe nce the Preferred Detail and Format of Software Explanations? (Nov 2024). https://doi.org/10.5281/zenodo.13175641
2024 doi
-
[24]
Obaidi, M., Droste, J., Deters, H., Herrmann, M., Klünde r, J., Schneider, K.: Do users’ explainability needs in software change with mood? I n: REFSQ’25 (2025)
2025
-
[25]
, Fischbach, J., Schneider, K.: Automating explanation need management in app reviews: A case study from the navigation app industry
Obaidi, M., Voß, N., Droste, J., Deters, H., Herrmann, M. , Fischbach, J., Schneider, K.: Automating explanation need management in app reviews: A case study from the navigation app industry. ICSE-SEIP’25 (2025)
2025
-
[26]
In: HCI-COLLAB’21
Ramos, H., Fonseca, M., Ponciano, L.: Modeling and evalu ating personas with software explainability requirements. In: HCI-COLLAB’21 . Springer
-
[27]
In: UMAP ’24
Sadeghi, M., Pöttgen, D., Ebel, P., Vogelsang, A.: Expla ining the unexplainable: The impact of misleading explanations on trust in unreliabl e predictions for hardly assessable tasks. In: UMAP ’24. Association for Computing M achinery
-
[28]
Unterbusch, M., Sadeghi, M., Fischbach, J., Obaidi, M., Vogelsang, A.: Explanation needs in app reviews: Taxonomy and automated detection. In: REW. IEEE (2023)
2023
-
[29]
i’d like an explanation for that!
Wiegand, G., Eiband, M., Haubelt, M., Hussmann, H.: “i’d like an explanation for that!”exploring reactions to unexpected autonomous drivi ng. In: MobileHCI’20
-
[30]
Springer (2012) Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ϯϱ ϰϬ ϰϱ ϱϬ Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ĂƚĂ ĂƚĂ
Wohlin, C., Runeson, P., Höst, M., Ohlsson, M.C., Regnel l, B., Wesslén, A.: Ex- perimentation in software engineering. Springer (2012) Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ϯϱ ϰϬ ϰϱ ϱϬ Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ĂƚĂ ĂƚĂ
2012
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.