REVIEW 4 major objections 4 minor 53 references
Explainers' Mental Representations of Explainees' Needs in Everyday Explanations
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Human explainers continuously update a partner model of the listener's knowledge and interests, and explainable AI should do the same.
desk verdict A careful qualitative study of explainers' partner models, with a real method caveat that tempers but doesn't kill the developmental story. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the 'partner model': the explainer's mental representation of the explainee, here decomposed into a 'technical model' (what the listener knows about the artifact) and a set of assumptions about interests. The paper maps both onto the dual nature theory of technological artifacts, which distinguishes Architecture (observable, measurable features such as pieces, rules, and structure) from Relevance (interpretable aspects such as purpose, meaning, and strategy). The argument is carried by the coded distribution of interview segments across these categories in different explanation phases—before, at the start, middle, and end of the explanation, and after it—showing how the partner model is refined and updated.
What would settle it
An experiment that probes explainers' assumptions in real time—for example, asking them immediately after each explanation move, or tracking where they look and what they say—would settle the question; if live probes show no Architecture-to-both trajectory in knowledge or no Relevance-to-both trajectory in interests, the recalled developmental pattern would be shown to be a post-hoc reconstruction rather than a real-time mental representation.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that explainers' partner models—their assumptions about what the explainee knows and cares about—are dynamic and dual-sided. Coding of pre-, post-, and video recall-interviews from nine explainers who explained the board game Quarto shows that assumed knowledge begins Architecture-centered and later incorporates Relevance, whereas assumed interests begin Relevance-centered and later incorporate Architecture. The patterns move from vague early assumptions to well-defined beliefs, and explainers frequently end the explanation even while believing the explainee still has gaps, treating some missing knowledge as unimportant or as requiring hands-on experience rather than more talk. These findings are offered as the empirical basis for a user-model component in adaptive explainable systems.
Load-bearing premise
The findings rest on the assumption that video recall-interviews give valid, undistorted access to what explainers actually thought during the explanation.
Editorial extensions
If this is right
- Adaptive explainable systems should maintain a partner model that tracks both the user's knowledge and the user's interests on the Architecture and Relevance dimensions, rather than a single measure of understanding.
- An explanation system could follow the decision rules the paper derives: start when knowledge is low and interest is high, continue until the demanded knowledge level is reached, change perspective when interests shift, re-explain when knowledge stagnates, and stop when knowledge is high and interest is low.
- Ending an explanation despite known knowledge gaps is a normal part of human explanation behavior, so XAI systems need not aim for complete coverage of every aspect.
- The observed sequence implies that explanations of technological artifacts should typically establish Architecture first, because assumptions about Relevance knowledge develop later on that foundation.
Reading between the lines
- If the two-dimensional partner model generalizes beyond the board-game case, a compact state representation—knowledge and interest values on Architecture and Relevance—could be a practical starting point for user models in XAI, though the paper itself stops at qualitative indicators.
- A natural test the paper does not run is whether explainers whose assumptions follow the observed trajectory produce explanations that explainees find more satisfying; linking the model to explanation quality would strengthen its practical relevance.
- For digital artifacts with many features, the partner model may need to track which subset of features a user cares about, not just the global balance of Architecture and Relevance, because not every feature will be relevant to every user.
- If the retrospective-recall worry is real, future work could validate the developmental pattern with live behavioral markers such as the timing of questions or gaze direction, which would also provide XAI with observable signals to monitor.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a qualitative study of explainers' mental representations of explainees' knowledge and interests during everyday explanations of a technological artifact (the board game Quarto). Nine explainers were interviewed before, during (via video recall), and after their explanations, and the transcripts were analyzed with deductive qualitative content analysis using the dual nature theory (Architecture vs. Relevance) as the coding lens. The central claims are that explainers' assumptions about explainees' knowledge begin centered on Architecture and develop toward both Architecture and Relevance, while assumptions about interests begin centered on Relevance and develop toward both; and that explainers often ended explanations despite perceiving residual knowledge gaps. The authors translate these findings into practical implications for user models in adaptive explainable systems.
Significance. If the developmental claims are valid, the paper provides a useful empirical starting point for XAI user modeling by identifying content dimensions (Architecture and Relevance) and temporal dynamics that an adaptive explainer should track. The study is carefully documented: the coding manual is described, inter-coder reliability is reported (Cohen's kappa = .75), and the analysis is transparent with detailed tables. The paper also explicitly limits its scope to first steps and acknowledges several methodological limitations. However, the central developmental claims currently rest on raw frequency counts that are not directly comparable across phases, and the video recall method measures retrospective reconstructions, not necessarily real-time mental states. The value of the contribution therefore depends on whether these evidentiary gaps can be resolved in revision.
major comments (4)
- [Section 4, Tables 3 and 5] The central developmental claims are based on raw code frequencies across phases that are not directly comparable. The number of video recall scenes differs by phase (15 start, 24 middle, 11 end), and the pre-interview probes knowledge and interests about board games in general, while the video recall and post-interviews probe knowledge and interests about Quarto specifically. The shift from 'Architecture-centered' to 'both Architecture and Relevance' may therefore reflect the change in measurement instrument rather than a true change in explainers' assumptions. The authors should either restrict comparisons to like-for-like categories (e.g., only Quarto-specific codes across VR-S, VR-M, VR-E, and Post) or report proportions per scene or per participant, and explicitly discuss the confound between phase and question focus.
- [Section 4, Answer to RQ1] The paper states that 'Many EXs weren't aware of this in the explanation itself but realized the knowledge gaps, particularly in the video recall-interview.' This admission directly undermines the claim that the video recall data reflect mental representations that were available to the explainer in real time to guide and adapt the explanation. The paper's framing as 'thereby enabling explainers to react to explainees' needs' is therefore not supported by the evidence. The authors should either reframe the findings as retrospective reconstructions of explainers' assumptions or provide convergent behavioral evidence (e.g., observable adaptation in the explanation itself) to support the real-time interpretation.
- [Section 3, Procedure and Section 4, Answer to RQ2] The interview questions differ systematically across phases: the pre-interview asks 'Which aspects of board games does the EE enjoy?' while the video recall asks 'What knowledge needs regarding the game did the EE have in that particular moment?' The paper notes that the phrase 'knowledge needs' was deliberately chosen to elicit richer answers about interests. This procedure confounds the phase being studied with the wording and focus of the elicitation question. The observed shift in interest frequencies (from Relevance in pre-interviews to Architecture in video recall-start) may be an artifact of the question change rather than a genuine developmental change in explainers' assumptions. The authors should analyze the data separately by elicitation question type or otherwise control for this confound.
- [Section 4, Table 3] For board-game knowledge, the proportion of Relevance codes is nearly identical in pre- and post-interviews (35/81 = 43.2% vs. 31/73 = 42.5%), while the overall shift toward 'both Architecture and Relevance' is driven entirely by Quarto-specific codes that were not elicited in the pre-interview. This indicates that the developmental finding may be a measurement artifact: the pre-interview simply did not ask about Quarto-specific knowledge. The authors should analyze board-game knowledge and Quarto knowledge separately and temper the abstract's claim that 'the assumed knowledge of explainees in the beginning is centered around Architecture and develops toward knowledge with regard to both Architecture and Relevance.'
minor comments (4)
- [Table 2] The percentages in the 'Total' row of Table 2 are inconsistent with the column totals: the table reports 612 (56.98%) and 379 (35.25%) for a total of 991 (100%), but 612/991 = 61.76% and 379/991 = 38.24%. The 56.98% and 35.25% figures appear to be computed with the 83 unallocated segments included in the denominator, which contradicts the stated total of 991. Please correct this presentation.
- [Section 4, Answer to RQ1] The abstract states that 'explainers often finished the explanation despite their perception that explainees still had gaps in knowledge,' while the results say 'a few EXs completed their explanations' despite perceived gaps. These two characterizations differ in magnitude; please align the wording.
- [Throughout] There are several typographical and formatting issues, including 'T able 1' (Table 1), 'T echnical Model' and 'The T echnical Model' in section headings, and 'segue' (should be 'segue' or 'transition' for clarity). These do not affect the content but should be cleaned up.
- [Section 5, Methodological Considerations and Limitations] The paper states that 'the rather small sample size was sufficient for our findings' without providing an explicit saturation argument or a justification in terms of the qualitative method. A brief statement on how saturation was determined, or a clearer acknowledgment that the findings are exploratory, would strengthen the limitations section.
Circularity Check
No significant circularity: the developmental claims are empirical coding results; the sole author self-citation is motivational, not load-bearing.
full rationale
The paper's central claims (answers to RQ1 and RQ2) are descriptive summaries of coded interview segments reported in Tables 3-6, not outputs of a model fitted to those same data. The dual nature theory supplies the coding categories (Architecture and Relevance), but whether explainers' assumptions shift in the reported way is an empirical finding based on 991 coded segments with intercoder reliability k=.75. The paper contains self-citations to the authors' prior study [49], but that citation is used as background ('In a recent study, the structure of everyday explanations... [49]') and later as a point of comparison ('This process is already known from an observation study'); it is not the source of the new interview-based results. The acknowledged limitations (video recall is retrospective; scene counts differ across phases: 15 start, 24 middle, 11 end; and the 'knowledge needs' prompt was used to elicit interests) are validity concerns about whether the interviews captured real-time mental representations, not circular reductions of the results to the method's assumptions. No equation, fitted parameter, or definition makes a predicted quantity equal to an input by construction. Therefore no circular step is present; score 1 reflects only the minor, non-load-bearing self-citation.
Assumptions & free parameters
assumptions (3)
- domain assumption Technological artifacts have a dual nature comprising Architecture and Relevance, and both perspectives are needed for understanding.
- domain assumption Video recall-interviews provide valid access to explainers' mental representations during the explanation.
- domain assumption The coding categories (knowledge/interests by Architecture/Relevance) are sufficiently exhaustive and reliable for frequency comparisons.
Cite this review
Pith. "Pith review of Explainers' Mental Representations of Explainees' Needs in Everyday Explanations." pith.science (2026). https://pith.science/paper/XX4RKUW3
@misc{pith2026241108514,
author = {Pith},
title = {Pith review of: Explainers' Mental Representations of Explainees' Needs in Everyday Explanations},
year = {2026},
howpublished = {\url{https://pith.science/paper/XX4RKUW3}},
note = {Machine review of arXiv:2411.08514}
}
read the original abstract
In explanations, explainers have mental representations of explainees' developing knowledge and shifting interests regarding the explanandum. These mental representations are dynamic in nature and develop over time, thereby enabling explainers to react to explainees' needs by adapting and customizing the explanation. XAI should be able to react to explainees' needs in a similar manner. Therefore, a component that incorporates aspects of explainers' mental representations of explainees is required. In this study, we took first steps by investigating explainers' mental representations in everyday explanations of technological artifacts. According to the dual nature theory, technological artifacts require explanations with two distinct perspectives, namely observable and measurable features addressing "Architecture" or interpretable aspects addressing "Relevance". We conducted extended semi structured pre-, post- and video recall-interviews with explainers (N=9) in the context of an explanation. The transcribed interviews were analyzed utilizing qualitative content analysis. The explainers' answers regarding the explainees' knowledge and interests with regard to the technological artifact emphasized the vagueness of early assumptions of explainers toward strong beliefs in the course of explanations. The assumed knowledge of explainees in the beginning is centered around Architecture and develops toward knowledge with regard to both Architecture and Relevance. In contrast, explainers assumed higher interests in Relevance in the beginning to interests regarding both Architecture and Relevance in the further course of explanations. Further, explainers often finished the explanation despite their perception that explainees still had gaps in knowledge. These findings are transferred into practical implications relevant for user models for adaptive explainable systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Addison Wesley Longman, Inc., New York (2001)
Anderson, L.W., Krathwohl, D.R.: A taxonomy for learning, teaching, and assessing: a revision of Bloom’s taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc., New York (2001)
work page 2001
-
[2]
Brennan, S.E., Hanna, J.E.: Partner-Specific Adaptation in Dialog. Top. Cogn. Sci. 1, 274–291 (2009)
work page 2009
-
[3]
Buhl, H.M.: Partner orientation and speaker’s knowledge as conflicting parameters in language production. J. Psycholinguist. Res. 30, 549–567 (2001)
work page 2001
- [4]
-
[5]
Calderhead, J.: Stimulated recall: A method for research on teaching. Br. J. Educ. Psychol. 51, 211–217 (1981)
work page 1981
-
[6]
Chi, M., Siler, S., Jeong, H.: Can tutors monitor students’ understanding accurately? COGNITION INSTRUCT. 22, (2004)
work page 2004
- [7]
-
[8]
Clark, H.H., Krych, M.A.: Speaking while monitoring addressees for understanding. J. M. L. 50, 62–81 (2004)
work page 2004
Show all 53 references
-
[9]
de Jong, T., Ferguson-Hessler, M.G.M.: Types and qualities of knowledge. Educ. Psychol. 31, 105–113 (1996)
1996
-
[10]
Dillenbourg, P., Lemaignan, S., Sangin, M., Nova, N., Molinari, G.: The Symme- try of Partner Modelling. Intern. J. Comput.-Support. Collab. Learn. 11, 227–253 (2016)
2016
-
[11]
In: Stephanidis, Co., Kurosu, M., Degen, H., Reinerman- Jones, L
Ehsan, U., Riedl, M.O.: Human-Centered Explainable AI: Towards a reflective sociotechnical approach. In: Stephanidis, Co., Kurosu, M., Degen, H., Reinerman- Jones, L. (eds.) HCI International 2020—Late Breaking Papers: Multimodality and Intelligence. pp. 449–466. Springer Inte...
2020
-
[12]
K¨ unstl
Fisher, J.B., Lohmer, V., Kern, F., Barthlen, W., Gaus, S., Rohlfing, K.J.: Explor- ing monological and dialogical phases in naturally occurring explanations. K¨ unstl. Intell. 36, 317–326 (2022)
2022
-
[13]
Second Edition
Fleiss, J.L.: Statistical Methods for Rates and Proportions. Second Edition. Wiley, John and Sons, Incorporated, New York, N.Y. (1981)
1981
-
[14]
von, Steinke, I.: A Companion to Qualitative Research
Flick, U., Kardoff, E. von, Steinke, I.: A Companion to Qualitative Research. SAGE (2004)
2004
-
[15]
In: Benson, C., Lunt, J
Fox-Turnball, W.: Autophotography: A means of stimulated recall for investigating technology education. In: Benson, C., Lunt, J. (eds.) International Handbook of Primary Technology Education, Rotterdam (2011)
2011
-
[16]
Furlough, C.S., Gillan, D.J.: Mental Models: Structural Differences and the Role of Experience. J. Cogn. Eng. Decis. Mak. 12, 269–287 (2018)
2018
-
[17]
Yale University Press (1981)
Garfinkel, Alan: Forms of explanation: Rethinking the questions in social theory. Yale University Press (1981)
1981
-
[18]
International Encyclopedia of the Social & Behavioral Sciences
Gentner, D.: Mental Models, Psychology of. International Encyclopedia of the Social & Behavioral Sciences. pp. 9683–9687. Elsevier (2001)
2001
-
[19]
Greca, I.M., Moreira, M.A.: Mental models, conceptual models, and modelling. Int. J. Sci. Educ. 22, 1–11 (2000)
2000
-
[20]
Policy Insights Behav
Harackiewicz, J.M., Smith, J.L., Priniski, S.J.: Interest matters: The importance of promoting interest in education. Policy Insights Behav. Brain Sci. 3, 220–227 (2016). 24 M. E. Schaffer et al
2016
-
[21]
https://arxiv.org/abs/1812.04608,(2018)
Hoffman, R., Mueller, S., Klein, G., Litman, J.: Metrics for explainable AI: Chal- lenges and prospects. https://arxiv.org/abs/1812.04608,(2018)
2018 arXiv
-
[22]
E&S.art46 (2011)
Jones, N.A., Ross, H., Lynam, T., Perez, P., Leitch, A.: Mental models: An inter- disciplinary synthesis of theory and methods. E&S.art46 (2011)
2011
-
[23]
KNOWL-BASED SYST
Kaplan, S., Uusitalo, H., Lensu, L.: A Unified and Practical User-Centric Frame- work for Explainable Artificial Intelligence. KNOWL-BASED SYST. 283, (2024)
2024
-
[24]
Keil, F.C.: Explanation and Understanding. Annu. Rev. Psychol. 57, 227–254 (2006)
2006
-
[25]
Krapp, A.: Interest, motivation and learning: An educational-psychological per- spective. Eur. J. Psychol. Educ. 14, 23–40 (1999)
1999
-
[26]
Kroes, P.: Technological explanations: The Relation between structure and function of technological objects. Soc. Philos. Technol. Q. Electr. J. 3, 124–134 (1998)
1998
-
[27]
Krull, D.S., Anderson, C.A.: The process of explanation. Curr. Dir. Psychol. 6, 1–5 (1997)
1997
-
[28]
In: 2013 IEEE Symposium on Visual Languages and Human Centric Computing
Kulesza, T., Stumpf, S., Burnett, M., Yang, S., Kwan, I., Wong, W.-K.: Too Much, too Little, or Just Right? Ways Explanations Impact End Users’ Mental Models. In: 2013 IEEE Symposium on Visual Languages and Human Centric Computing. pp. 3–10. IEEE, San Jose, CA, USA (2013)
2013
-
[29]
In: Holyoak, K.J., Morrison, R.G
Lombrozo, T.: Explanation and Abductive Inference. In: Holyoak, K.J., Morrison, R.G. (eds.) The Oxford Handbook of Thinking and Reasoning. pp. 260–276. Oxford University Press (2012)
2012
-
[30]
Lyle, J.: Stimulated Recall: A report on its use in naturalistic research. Br. Educ. Res. J. 29, 861–878 (2003)
2003
-
[31]
Lawrence Erlbaum Associates (2005)
Mackey, Alison, Gass, Susan M.: Second language research: Methodology and de- sign. Lawrence Erlbaum Associates (2005)
2005
-
[32]
Mayer, R.E., Mathias, A., Wetzell, K.: Fostering understanding of multimedia mes- sages through pre-training: Evidence for a two-stage theory of mental model con- struction. J. Exp. Psychol. Appl. 8, 147–154 (2002)
2002
-
[33]
Miller, T.: Explanation in Artificial Intelligence: Insights from the Social Sciences. J. Artif. Intell. 267, 1–38 (2019)
2019
-
[34]
In: Harris, D
Neerincx, M.A., van Der Waa, J., Kaptein, F., van Diggelen, J., J.: Using perceptual and cognitive explanations for enhanced human-agent team performance. In: Harris, D. (ed.) Engineering Psychology and Cognitive Ergonomics. pp. 204–214. Springer International Publishing, Cham (2018)
2018
-
[35]
User Model User-Adap Inter
N¨ uckles, M., Winter, A., Wittwer, J., Herbert, M., H¨ ubner, S.: How do experts adapt their explanations to a layperson’s knowledge in asynchronous communica- tion? An experimental study. User Model User-Adap Inter. 16, 87–127 (2006)
2006
-
[36]
In: D’hondt, S., ¨Ostman, J.-O., Verschueren, J
O’Connell, D.C., Kowal, S.: Transcription systems for spoken discourse. In: D’hondt, S., ¨Ostman, J.-O., Verschueren, J. (eds.) The Pragmatics of Interaction. pp. 240–254. John Benjamins Publishing Company, Amsterdam (2009)
2009
-
[37]
Springer Fachmedien Wiesbaden, Wiesbaden (2019)
R¨ adiker, S., Kuckartz, U.: Analyse qualitativer Daten mit MAXQDA: Text, Audio und Video. Springer Fachmedien Wiesbaden, Wiesbaden (2019)
2019
-
[38]
IUI Workshops
Ribera, M., Lapedriza, A.: Can we do better explanations? A proposal of user- centered explainable AI. IUI Workshops. (2019)
2019
-
[39]
IEEE Trans
Rohlfing, K.J., et al.: Explanation as a social practice: Toward a conceptual frame- work for the social design of AI systems. IEEE Trans. Cogn. Dev. Syst. 13, 717–728 (2021)
2021
-
[40]
A.: Mental models: Issues in construction, congruency, and cognition
Savage-Knepshield, P. A.: Mental models: Issues in construction, congruency, and cognition. The State University of New Jersey, Rutgers, New Brunswick (2001) Explainers’ Mental Representations of Explainees’ Needs 25
2001
-
[41]
Schoonderwoerd, T.A.J., Jorritsma, W., Neerincx, M.A., van den Bosch, K.: Human-centered XAI: developing design patterns for explanations of clinical de- cision support systems. Int. J. Hum.-Comput. Stud. 154, Article 102684 (2021)
2021
-
[42]
SAGE Publications Ltd, London (2012)
Schreier, M.: Qualitative content analysis in practice. SAGE Publications Ltd, London (2012)
2012
-
[43]
In: Mittermeir, R.T., Sys lo, M.M
Schulte, C.: Duality reconstruction—teaching digital artifacts from a socio- technical perspective. In: Mittermeir, R.T., Sys lo, M.M. (eds.) Informatics Edu- cation—Supporting Computational Thinking. pp. 110–121. Springer Berlin Heidel- berg, Berlin, Heidelberg (2008)
2008
-
[44]
In: Proceedings of the 18th Koli Calling International Conference on Computing Education Research
Schulte, C., Budde, L.: A framework for computing education: Hybrid interaction system: The need for a bigger picture in computing education. In: Proceedings of the 18th Koli Calling International Conference on Computing Education Research. pp. 1–10. ACM, Koli Finland (2018)
2018
-
[45]
Silvia, P.J.: Expressed and Measured Vocational Interests: Distinctions and Defi- nitions. J. Vocat. Behav. 59, 382–393 (2001)
2001
-
[46]
K¨ unstl
Sokol, K., Flach, P.: One explanation does not fit all: The promise of interactive explanations for machine learning transparency. K¨ unstl. Intell.34, 235–250 (2020)
2020
-
[47]
Soloway, E.: Learning to program = learning to construct mechanisms and expla- nations. Commun. ACM.29, 850–858 (1986)
1986
-
[48]
Springer New York, New York, NY (2011
Sweller, J., Ayres, P., Kalyuga, S.: Cognitive load theory. Springer New York, New York, NY (2011
2011
-
[49]
In: Longo, L
Terfloth, L., Schaffer, M., Buhl, H.M., Schulte, C.: Adding why to what? Analyses of an everyday explanation. In: Longo, L. (eds.) Explainable Artificial Intelligence. pp. 256–279. Springer Nature Switzerland, Cham (2023)
2023
-
[50]
Vermaas, P.E., Houkes, W.: Technical functions: A drawbridge between the inten- tional and structural natures of technical artefacts. Stud. Hist. Philos. Sci. Part A. 37, 5–18 (2006)
2006
-
[51]
Wittwer, J., Renkl, A.: How effective are instructional explanations in example- based learning? A meta-analytic review. Educ. Psychol. Rev. 22, 393–409 (2010)
2010
-
[52]
Zemla, J.C., Sloman, S., Bechlivanidis, C., Lagnado, D.A.: Evaluating everyday explanations. Psychon. Bull. Rev. 24, 1488–1500 (2017)
2017
-
[53]
Zhang, Y.: The development of users’ mental models of medlineplus in information searching. LIBR. INFORM. SCI. RES. 35, 159–170 (2013)
2013
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.