Pith. sign in

REVIEW 4 major objections 5 minor 30 references

How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Users' app knowledge only weakly shapes which software explanations they prefer.

desk verdict Useful new dataset and honest limitations, but the central null is underpowered and the abstract overclaims; a revised version could be solid. read the letter →

arxiv 2502.06549 v1 pith:P4LEYH5W submitted 2025-02-10 cs.SE

classification cs.SE
keywords explainabilitysoftwareexplanationsapp-specificknowledgeexplanationpreferenceslevelofdetailformatusersurveyadaptive
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Software is becoming harder to understand, and one proposed answer is to adapt in-app explanations to each user. This paper asks whether a user's app-specific knowledge can predict how detailed an explanation they want and in what form, so that help content could be tailored automatically. Based on an online survey of 58 users of office and browser software, the authors find that objective knowledge correlates strongly with self-assessed knowledge but only weakly with preferred explanation form and not at all with preferred detail level. The most common choice was short text explanations at moderate detail. The paper concludes that explanation preferences are largely subjective, so adaptive explanation systems and user personas should not treat knowledge scores or demographics as reliable predictors of explanation needs.

What carries the argument

The argument runs on three measurement instruments developed for the survey: a five-level explanation detail scale with written examples, a five-level explanation form scale, and an objective six-question quiz per app category whose scores are mapped to five knowledge levels. Relationships are tested with rank correlation for ordinal or metric variables and chi-square tests for nominal variables, using a multiple-comparison correction to limit false positives. The strong correlation between self-assessed and objective knowledge acts as an internal validity check for the knowledge construct, while the office-versus-browser comparison tests whether effects generalize across application categories.

What would settle it

Present users with real explanation texts that vary in length and format, let them choose or use them in a working application, and compare choices with objective knowledge scores; if users with high app knowledge consistently choose shorter or more technical explanations while novices choose longer tutorial-style ones, the paper's claim of only a weak relationship would be overturned.

Watch

Extended reading notes

Core claim

The paper's central claim is that app-specific knowledge, measured by an objective quiz, is only weakly related to the explanation form and detail level users prefer. In its survey of 58 participants across two app categories, office software and browsers, self-assessed knowledge matched objectively quizzed knowledge strongly, but this knowledge did not predict desired explanation detail in either category. Knowledge did correlate moderately, negatively, with preferred explanation form in office software, meaning more knowledgeable users leaned toward simpler forms, while the same effect did not appear for browsers. Confidence in using software behaved similarly, and demographic factors showed no significant relationship with either form or detail. The authors interpret the overall pattern as evidence that explanation preferences are shaped more by individual subjectivity than by knowledge, and that default explanations should be short, moderately detailed texts.

Load-bearing premise

The study assumes its self-built quiz and hypothetical preference scales measure real app-specific knowledge and real explanation preferences, even though the authors note the scales were designed by two researchers without third-party review and participants never interacted with actual explanations.

Editorial extensions

If this is right

  • Adaptive explanation systems should not infer a user's preferred detail level from app-knowledge scores, because the survey found no correlation with detail level in office or browser software.
  • Self-assessed knowledge can serve as a practical proxy for objective knowledge in user studies, given the strong correlation the survey found between the two.
  • Default help content should favor short text at moderate detail, since that was the most common preference in both app categories.
  • Persona-based requirements analysis should treat app-specific knowledge and confidence as weak, non-deterministic inputs rather than fixed predictors.
  • Because browser users leaned toward even less detail than office users, explanation defaults may need to vary by application category.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: stated preferences may not match actual behavior, so an A/B test in a real application could reveal stronger knowledge effects than the hypothetical scales did.
  • Beyond the paper: the authors leave open whether task stakes, rather than user traits, drive explanation needs; comparing high-stakes and routine tasks is a direct test of that possibility.
  • Beyond the paper: if the weak correlations hold, adaptive explanation systems may need to learn preferences from interaction history instead of static user profiles.
  • Beyond the paper: the strong subjective-objective knowledge correlation suggests the null result is not simply a flawed knowledge quiz, leaving the unvalidated preference scales as the main measurement threat.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper reports an online survey (n=58) investigating whether users' app-specific knowledge, software confidence, and demographics relate to their preferred explanation detail level and form for Office and Browser software. The authors report that participants generally prefer moderately detailed short-text explanations, that subjective and objective app-specific knowledge correlate strongly, and that most hypothesized relationships with explanation preferences are non-significant. The central claim is that app-specific knowledge is only weakly related to preferred explanation form/detail, so adaptive explanation systems should not rely heavily on knowledge or demographics. The paper includes public data/code and discusses threats to validity.

Significance. If the central claim were statistically established, the result would be practically useful for requirements engineering and explainability design, particularly for deciding whether user knowledge or demographic data can drive adaptive explanations. The paper's strengths are its concrete survey instrument, two app categories, public data/code, and transparent discussion of validity threats. However, the significance is currently limited by the small convenience sample, unvalidated self-built measures, and inferential issues that make the 'weakly related' null claim unsupported as stated.

major comments (4)
  1. [Section 4.3 / Table 5 / H4] The treatment of H4 is internally contradictory. The text states that 'H4 was not rejected, as its p-value is 0.0042, thus falling below the required significance level for rejection,' but p=0.0042 is below 0.05 and would imply rejection. Moreover, this p-value does not appear in Table 5, where all reported demographic p-values are much larger (e.g., 0.44, 0.14, 0.95). This inconsistency directly undermines the RQ3 conclusion and the abstract's claim about demographic influences, and it must be corrected with the actual aggregate test and a consistent interpretation.
  2. [Section 5.1 / Table 5] The central 'only weakly related' conclusion is an overinterpretation of non-significant correlations. With n≈53–57 and a Bonferroni-corrected alpha of 0.0125, Spearman correlations up to roughly |r|=0.33 can be non-significant, and the 95% confidence interval for an observed r of 0 spans approximately ±0.27. No confidence intervals or equivalence tests (e.g., TOST) are reported, so failure to reject does not establish 'no relationship' or 'weak relationship.' The significant Office-form correlation (H2.10, r=-0.49) is dismissed as not generalizable without a statistical bound, yet it is the strongest evidence against the paper's weak-relationship claim. The authors should report effect-size confidence intervals and/or equivalence bounds, or explicitly reframe the conclusion as exploratory.
  3. [Section 5.4 / Table 2] The construct-validity threat acknowledged in Section 5.4 is directly load-bearing for the central null. The preferred detail level and form are single-item hypothetical preference scales, the objective knowledge quiz was developed by two researchers without third-party review, and participants did not interact with real explanations. Table 2 declares Eform as nominal, yet Table 5 analyzes it with Spearman correlation coefficients, mixing scale assumptions. These measurement issues can attenuate or distort correlations, so the observed weak/null pattern may be a methodological artifact rather than a true absence of relationship. The authors should either provide validity evidence (e.g., pilot data, item analysis, or comparison with observed behavior) or temper the central claim to an exploratory finding.
  4. [Abstract / Section 6] The abstract and conclusion claim that demographic aspects (like gender) influence app-specific knowledge and that app-specific knowledge correlates with application confidence, suggesting a possible mediated relationship. However, Table 5 contains no tests of demographics against app-specific knowledge or confidence; it only tests demographics against explanation form and detail level. These claims are not supported by any reported analysis and should be removed or substantiated with the missing correlation/mediation analyses.
minor comments (5)
  1. [Abstract] The sentence 'This study investigates factors influencing users' preferred the level of detail' contains a typo ('preferred the level' should be 'preferred level').
  2. [Section 3.4 / Table 3] The notation H10, H20, H30, H40 is easily confused with subhypotheses like H1.10, H2.10, etc.; consider renaming the aggregate hypotheses (e.g., H1, H2, H3, H4) to improve readability.
  3. [Section 4.3 / Table 5] The text reports a p-value of 0.00009 for H2.10, while Table 5 lists p=0.00; the table should use a consistent number of decimal places or a p<0.001 notation.
  4. [Section 5.2] The interpretation that the strong subjective-objective knowledge correlation validates the knowledge construct is overstated; a strong correlation between two self-report or quiz-based measures could reflect shared method variance, and this should be acknowledged.
  5. [Section 5.4] The threats-to-validity section is thoughtful, but it would be strengthened by explicitly stating that the sample size of 58 and the wide confidence intervals prevent the conclusion that non-significant findings constitute evidence of absence.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical survey with external outcome measures and no fitted-input prediction chain.

full rationale

This paper is an empirical survey study with no mathematical derivation, fitted model, or uniqueness argument. The central claim—that app-specific knowledge is only weakly related to preferred explanation form and detail level (Section 5.1, RQ1)—is a null correlation result computed against external survey responses. The predictors (objective quiz score, self-assessed knowledge, confidence, demographics) and outcomes (preferred form and detail level) are independent measurements; neither is defined in terms of the other, and no parameter is fitted to the outcome and then renamed as a prediction. The H1 validation correlation between subjective and objective app knowledge is used as a manipulation check, not as an input to the preference correlations. Self-citations in the related work section are contextual and not load-bearing: for instance, [24] is a separate mood study and [23] is the dataset release. The paper explicitly acknowledges threats to construct validity, internal validity, and conclusion validity in Section 5.4, including the lack of third-party review of self-built scales, possible answer lookup, and Type II error risk from Bonferroni correction. These are legitimate methodological limitations, and the underpowered null interpretation could be criticized as statistically overconfident, but that is a correctness/statistical concern, not circularity. No quoted step reduces to its own input, so the circularity score is 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claims rest on measurement validity and sample representativeness. There are no model-fitting parameters, but the objective knowledge thresholds are hand-chosen. The study does not introduce new theoretical entities.

free parameters (1)
  • Objective knowledge category thresholds = 20% / 40% / 60% / 80%
    Quiz scores are mapped to five ordinal knowledge levels using these hand-chosen cutoffs (Section 3.4). Different cutoffs could change the correlation between objective knowledge and preferences.
assumptions (3)
  • domain assumption The self-constructed quiz questions and ordinal scales are valid measures of app-specific knowledge and explanation preferences.
    Section 5.4 admits the metric was developed by two researchers without third-party review, and the survey lacks direct interaction with actual explanations.
  • domain assumption Participants' self-reported preferences for hypothetical explanation formats match their real-world preferences.
    Section 5.4 acknowledges 'the absence of direct interaction with software explanations in the survey introduces a risk of response bias'.
  • standard math Nonparametric correlation tests on a sample of 53 to 57 participants can detect the relationships of interest.
    Spearman and chi-square tests are standard, but the small sample yields low statistical power, making null results weak evidence of absence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?." pith.science (2026). https://pith.science/paper/P4LEYH5W

@misc{pith2026250206549,
  author       = {Pith},
  title        = {Pith review of: How Does Users' App Knowledge Influence the Preferred Level of Detail and Format of Software Explanations?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P4LEYH5W}},
  note         = {Machine review of arXiv:2502.06549}
}
read the original abstract

Context and Motivation: Due to their increasing complexity, everyday software systems are becoming increasingly opaque for users. A frequently adopted method to address this difficulty is explainability, which aims to make systems more understandable and usable. Question/problem: However, explanations can also lead to unnecessary cognitive load. Therefore, adapting explanations to the actual needs of a user is a frequently faced challenge. Principal ideas/results: This study investigates factors influencing users' preferred the level of detail and the form of an explanation (e.g., short text or video tutorial) in software. We conducted an online survey with 58 participants to explore relationships between demographics, software usage, app-specific knowledge, as well as their preferred explanation form and level of detail. The results indicate that users prefer moderately detailed explanations in short text formats. Correlation analyses revealed no relationship between app-specific knowledge and the preferred level of detail of an explanation, but an influence of demographic aspects (like gender) on app-specific knowledge and its impact on application confidence were observed, pointing to a possible mediated relationship between knowledge and preferences for explanations. Contribution: Our results show that explanation preferences are weakly influenced by app-specific knowledge but shaped by demographic and psychological factors, supporting the development of adaptive explanation systems tailored to user expertise. These findings support requirements analysis processes by highlighting important factors that should be considered in user-centered methods such as personas.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

30 extracted references · 29 canonical work pages

  1. [1]

    In : SEAA (2019)

    Andrade, H., Lwakatare, L.E., Crnkovic, I., Bosch, J.: So ftware challenges in het- erogeneous computing: A multiple case study in industry. In : SEAA (2019)

  2. [2]

    ESEC/FSE 2020

    Antinyan, V.: Revealing the complexity of automotive sof tware. ESEC/FSE 2020

  3. [3]

    JSS 195 (2023)

    Brunotte, W., Specht, A., Chazette, L., Schneider, K.: Pr ivacy explanations–a means to end-user trust. JSS 195 (2023)

  4. [4]

    Buiten, M.C., Dennis, L.A., Schwammberger, M.: A vision o n what explanations of autonomous systems are of interest to lawyers. In: REW. IE EE (2023)

  5. [5]

    Chazette, L., Brunotte, W., Speith, T.: Exploring explai nability: a definition, a model, and a knowledge catalogue. In: RE. IEEE (2021)

  6. [6]

    REJ 25(4) (2020)

    Chazette, L., Schneider, K.: Explainability as a non-fun ctional requirement: chal- lenges and recommendations. REJ 25(4) (2020)

  7. [7]

    In: RE W

    Deters, H., Droste, J., Fechner, M., Klünder, J.: Explana tions on demand-a tech- nique for eliciting the actual need for explanations. In: RE W. IEEE (2023)

  8. [8]

    In: REFSQ’24 (2 024) 16 Obaidi et al

    Deters, H., Droste, J., Obaidi, M., Schneider, K.: How exp lainable is your system? towards a quality model for explainability. In: REFSQ’24 (2 024) 16 Obaidi et al

Show all 30 references
  1. [9]

    In: EASE’23

    Deters, H., Droste, J., Schneider, K.: A means to what end? evaluating the ex- plainability of software systems using goal oriented heuri stics. In: EASE’23

  2. [10]

    In : RE4AI’24

    Droste, J., Deters, H., Fuchs, R., Schneider, K.: Peekin g outside the black-box: Ai explainability requirements beyond interpretability. In : RE4AI’24

  3. [11]

    In: R E’24

    Droste, J., Deters, H., Obaidi, M., Schneider, K.: Expla nations in everyday software systems: Towards a taxonomy for explainability needs. In: R E’24

  4. [12]

    Droste, J., Deters, H., Puglisi, J., Klünder, J.: Design ing end-user personas for explainability requirements using mixed methods research . In: REW. IEEE (2023)

  5. [13]

    In: Shin, C.S., Di Bucchian- ico, G., Fukuda, S., Ghim, Y.G., Montagna, G., Carvalho, C

    Gabbas, M., Ryu, Y., Park, J., Kim, K.: Understanding cha llenges of designing for complex users by adapting the existing framework. In: Shin, C.S., Di Bucchian- ico, G., Fukuda, S., Ghim, Y.G., Montagna, G., Carvalho, C. ( eds.) Advances in Industrial Design. Springer Interna...

  6. [14]

    Springer, New York, NY, USA (2013)

    Haynes, W.: Bonferroni Correction. Springer, New York, NY, USA (2013)

  7. [15]

    Kästner, L., Langer, M., Lazar, V., Schomäcker, A., Spei th, T., Sterz, S.: On the relation of trust and explainability: Why to engineer for tr ustworthiness. REW’21

  8. [16]

    Kim, G., Yeo, D., Jo, T., Rus, D., Kim, S.: What and when to e xplain? on-road evaluation of explanations in highly automated vehicles. P roc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7(3) (sep 2023)

  9. [17]

    SIGSOFT Softw

    Kitchenham, B.A., Pfleeger, S.L.: Principles of survey r esearch part 2: designing a survey. SIGSOFT Softw. Eng. Notes 27(1), 18–20 (Jan 2002)

  10. [18]

    In: RE’19

    Köhl, M.A., Baum, K., Langer, M., Oster, D., Speith, T., B ohlender, D.: Explain- ability as a non-functional requirement. In: RE’19

  11. [19]

    Empirical Software Engineering 26(1) (2021)

    Levy, O., Feitelson, D.: Understanding large-scale sof tware systems – structure and flows. Empirical Software Engineering 26(1) (2021)

  12. [20]

    Compute r 45(08) (aug 2012)

    Mens, T.: On the complexity of software systems. Compute r 45(08) (aug 2012)

  13. [21]

    User Modeling an d User-Adapted In- teraction 27 (2017)

    Nunes, I., Jannach, D.: A systematic review and taxonomy of explanations in decision support and recommender systems. User Modeling an d User-Adapted In- teraction 27 (2017)

  14. [22]

    (Sep 2024)

    Obaidi, M.: Dataset: Gold standard dataset for explaina bility need detection in app reviews. (Sep 2024). https://doi.org/10.5281/zenodo.11522828

  15. [23]

    https://doi.org/10.5281/zenodo.13175641

    Obaidi, M.: Dataset: How Does Users’ App Knowledge Influe nce the Preferred Detail and Format of Software Explanations? (Nov 2024). https://doi.org/10.5281/zenodo.13175641

  16. [24]

    Obaidi, M., Droste, J., Deters, H., Herrmann, M., Klünde r, J., Schneider, K.: Do users’ explainability needs in software change with mood? I n: REFSQ’25 (2025)

  17. [25]

    , Fischbach, J., Schneider, K.: Automating explanation need management in app reviews: A case study from the navigation app industry

    Obaidi, M., Voß, N., Droste, J., Deters, H., Herrmann, M. , Fischbach, J., Schneider, K.: Automating explanation need management in app reviews: A case study from the navigation app industry. ICSE-SEIP’25 (2025)

  18. [26]

    In: HCI-COLLAB’21

    Ramos, H., Fonseca, M., Ponciano, L.: Modeling and evalu ating personas with software explainability requirements. In: HCI-COLLAB’21 . Springer

  19. [27]

    In: UMAP ’24

    Sadeghi, M., Pöttgen, D., Ebel, P., Vogelsang, A.: Expla ining the unexplainable: The impact of misleading explanations on trust in unreliabl e predictions for hardly assessable tasks. In: UMAP ’24. Association for Computing M achinery

  20. [28]

    Unterbusch, M., Sadeghi, M., Fischbach, J., Obaidi, M., Vogelsang, A.: Explanation needs in app reviews: Taxonomy and automated detection. In: REW. IEEE (2023)

  21. [29]

    i’d like an explanation for that!

    Wiegand, G., Eiband, M., Haubelt, M., Hussmann, H.: “i’d like an explanation for that!”exploring reactions to unexpected autonomous drivi ng. In: MobileHCI’20

  22. [30]

    Springer (2012) Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ϯϱ ϰϬ ϰϱ ϱϬ Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ĂƚĂ ĂƚĂ

    Wohlin, C., Runeson, P., Höst, M., Ohlsson, M.C., Regnel l, B., Wesslén, A.: Ex- perimentation in software engineering. Springer (2012) Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ϯϱ ϰϬ ϰϱ ϱϬ Ϭ ϱ ϭϬ ϭϱ ϮϬ Ϯϱ ϯϬ ĂƚĂ ĂƚĂ

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.