{"id":"00762a13-528c-49e1-bc2c-0f543481c3db","arxiv_id":"2411.08514","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Explainers' assumptions about a listener's knowledge start with technical details (Architecture) and later broaden to include purpose and strategy (Relevance), while interest assumptions follow the opposite path.","lead":"This paper studies how people explaining a board game guess what their listener already knows and wants to know, and how those guesses shift over the course of the explanation. It suggests adaptive AI explanation systems could build similar 'partner models' to tailor their output to each user.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Video recall interviews reconstruct, not measure, real-time mental representations; the developmental pattern may largely be post-hoc storytelling.","rationale":"The paper's strongest claim is that explainers build dynamic partner models of explainee knowledge and interests that shift across explanation phases. That claim is load-bearing for the XAI implications: if the observed temporal pattern is not a real-time cognitive process but a retrospective reconstruction, then the proposed user-model component is not supported by this study. The reader identified this same weakest assumption, and I agree. The concern is not merely a philosophical worry about introspection; the paper itself reports that explainers realized knowledge gaps only during video recall, not during the explanation, which is direct evidence that recall content is not identical to online representation. The unequal number of video recall scenes per phase (15, 24, 11) further confounds the frequency tables: counts are raw segment counts, not normalized by number of scenes or by amount of talk, so the apparent growth of Relevance coding in the middle and end phases could reflect the larger number of middle scenes and the content of the selected moments. There is also an internal inconsistency in Table 2: the column percentages for Architecture and Relevance are 56.98% and 35.25%, but the corresponding raw totals are 612/991 = 61.76% and 379/991 = 38.24%, which further reduces confidence in the quantitative summaries. None of this is fatal to the qualitative, exploratory contribution: the coding scheme is theory-grounded, the procedure is described transparently, the authors acknowledge the introspective character of the method, and the developmental pattern is plausible. But the central claim is conditional on the validity of retrospective recall as a measure of online mental representations, and that condition has not been established. The right verdict is therefore to keep the conditional acceptance pending independent validation, not to accept the dynamic partner model as empirically established or to reject the paper outright.","tokens_in":18509,"tokens_out":3362,"duration_ms":33536,"concrete_test":"Run a validation study with a new sample of explainers in the same Quarto paradigm: collect concurrent think-aloud or immediate timeline ratings of the explainee's knowledge gaps and interests at regular intervals during the explanation, then administer the same video recall procedure used in this paper. Compare the Architecture/Relevance proportions and temporal trajectories between online and delayed reports. If delayed reports systematically shift toward a more orderly Architecture-to-Relevance narrative than online reports, the reported developmental pattern is an artifact of reconstruction. As a complementary check on existing data, recompute Tables 3 and 5 as per-scene rates and as phase-wise proportions to test whether the claimed trends survive the unequal scene counts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central developmental claim—that explainers' assumptions about explainees' knowledge shift from Architecture toward Architecture plus Relevance, and interests shift from Relevance toward both—rests entirely on what explainers say in retrospective video recall interviews (Section 3, Procedure and Material; Tables 3 and 5). Those interviews prompt explainers to watch selected scenes and report what knowledge needs the explainee had 'in that particular moment' and how understanding developed. This measures retrospective reconstructions, not real-time mental states. The paper itself provides direct evidence of the gap: in Section 4 (Answer to RQ1), the authors report that many explainers 'weren't aware of this in the explanation itself but realized the knowledge gaps, particularly in the video recall-interview.' That admission means the recalled content was not consciously available during the explanation, undermining the claim that these are representations used online to adapt the explanation. Scene selection also differs by phase (15 start, 24 middle, 11 end video recalls), so raw code frequencies in Tables 3 and 5 are not directly comparable across phases; the Architecture-to-Relevance shift could be an artifact of which moments were selected and of a culturally standard narrative about how explanations proceed. The data cannot distinguish an updated partner model used in real time from a coherent post-hoc account constructed after the fact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a qualitative study of explainers' mental representations of explainees' knowledge and interests during everyday explanations of a technological artifact (the board game Quarto). Nine explainers were interviewed before, during (via video recall), and after their explanations, and the transcripts were analyzed with deductive qualitative content analysis using the dual nature theory (Architecture vs. Relevance) as the coding lens. The central claims are that explainers' assumptions about explainees' knowledge begin centered on Architecture and develop toward both Architecture and Relevance, while assumptions about interests begin centered on Relevance and develop toward both; and that explainers often ended explanations despite perceiving residual knowledge gaps. The authors translate these findings into practical implications for user models in adaptive explainable systems.","tokens_in":18827,"tokens_out":5742,"duration_ms":53173,"significance":"If the developmental claims are valid, the paper provides a useful empirical starting point for XAI user modeling by identifying content dimensions (Architecture and Relevance) and temporal dynamics that an adaptive explainer should track. The study is carefully documented: the coding manual is described, inter-coder reliability is reported (Cohen's kappa = .75), and the analysis is transparent with detailed tables. The paper also explicitly limits its scope to first steps and acknowledges several methodological limitations. However, the central developmental claims currently rest on raw frequency counts that are not directly comparable across phases, and the video recall method measures retrospective reconstructions, not necessarily real-time mental states. The value of the contribution therefore depends on whether these evidentiary gaps can be resolved in revision.","major_comments":[{"comment":"The central developmental claims are based on raw code frequencies across phases that are not directly comparable. The number of video recall scenes differs by phase (15 start, 24 middle, 11 end), and the pre-interview probes knowledge and interests about board games in general, while the video recall and post-interviews probe knowledge and interests about Quarto specifically. The shift from 'Architecture-centered' to 'both Architecture and Relevance' may therefore reflect the change in measurement instrument rather than a true change in explainers' assumptions. The authors should either restrict comparisons to like-for-like categories (e.g., only Quarto-specific codes across VR-S, VR-M, VR-E, and Post) or report proportions per scene or per participant, and explicitly discuss the confound between phase and question focus.","section":"Section 4, Tables 3 and 5"},{"comment":"The paper states that 'Many EXs weren't aware of this in the explanation itself but realized the knowledge gaps, particularly in the video recall-interview.' This admission directly undermines the claim that the video recall data reflect mental representations that were available to the explainer in real time to guide and adapt the explanation. The paper's framing as 'thereby enabling explainers to react to explainees' needs' is therefore not supported by the evidence. The authors should either reframe the findings as retrospective reconstructions of explainers' assumptions or provide convergent behavioral evidence (e.g., observable adaptation in the explanation itself) to support the real-time interpretation.","section":"Section 4, Answer to RQ1"},{"comment":"The interview questions differ systematically across phases: the pre-interview asks 'Which aspects of board games does the EE enjoy?' while the video recall asks 'What knowledge needs regarding the game did the EE have in that particular moment?' The paper notes that the phrase 'knowledge needs' was deliberately chosen to elicit richer answers about interests. This procedure confounds the phase being studied with the wording and focus of the elicitation question. The observed shift in interest frequencies (from Relevance in pre-interviews to Architecture in video recall-start) may be an artifact of the question change rather than a genuine developmental change in explainers' assumptions. The authors should analyze the data separately by elicitation question type or otherwise control for this confound.","section":"Section 3, Procedure and Section 4, Answer to RQ2"},{"comment":"For board-game knowledge, the proportion of Relevance codes is nearly identical in pre- and post-interviews (35/81 = 43.2% vs. 31/73 = 42.5%), while the overall shift toward 'both Architecture and Relevance' is driven entirely by Quarto-specific codes that were not elicited in the pre-interview. This indicates that the developmental finding may be a measurement artifact: the pre-interview simply did not ask about Quarto-specific knowledge. The authors should analyze board-game knowledge and Quarto knowledge separately and temper the abstract's claim that 'the assumed knowledge of explainees in the beginning is centered around Architecture and develops toward knowledge with regard to both Architecture and Relevance.'","section":"Section 4, Table 3"}],"minor_comments":[{"comment":"The percentages in the 'Total' row of Table 2 are inconsistent with the column totals: the table reports 612 (56.98%) and 379 (35.25%) for a total of 991 (100%), but 612/991 = 61.76% and 379/991 = 38.24%. The 56.98% and 35.25% figures appear to be computed with the 83 unallocated segments included in the denominator, which contradicts the stated total of 991. Please correct this presentation.","section":"Table 2"},{"comment":"The abstract states that 'explainers often finished the explanation despite their perception that explainees still had gaps in knowledge,' while the results say 'a few EXs completed their explanations' despite perceived gaps. These two characterizations differ in magnitude; please align the wording.","section":"Section 4, Answer to RQ1"},{"comment":"There are several typographical and formatting issues, including 'T able 1' (Table 1), 'T echnical Model' and 'The T echnical Model' in section headings, and 'segue' (should be 'segue' or 'transition' for clarity). These do not affect the content but should be cleaned up.","section":"Throughout"},{"comment":"The paper states that 'the rather small sample size was sufficient for our findings' without providing an explicit saturation argument or a justification in terms of the qualitative method. A brief statement on how saturation was determined, or a clearer acknowledgment that the findings are exploratory, would strengthen the limitations section.","section":"Section 5, Methodological Considerations and Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an XAI venue and the qualitative analysis is systematic, but the primary empirical claim rests on frequency comparisons that are confounded by measurement differences across phases and by the retrospective nature of video recall. The authors' own admission that many explainers were not aware of knowledge gaps during the explanation itself is a particular concern. The issues are addressable within the manuscript's scope (e.g., re-analyzing the data with comparable categories, reporting proportions, and softening the real-time interpretation), which is why I recommend major revision rather than rejection. I would also encourage the editor to consider whether the incremental novelty relative to the authors' prior work [49] is sufficient for the target venue; the discussion of the literature could more explicitly state what is new beyond that earlier study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a careful qualitative study that delivers exactly what it promises: a first empirical look at how explainers' assumptions about explainees' knowledge and interests evolve across an explanation, coded using the dual nature theory (Architecture vs. Relevance). The new bit is the 'technical model within partner model' framing applied to XAI user models, and the specific finding that knowledge assumptions start Architecture-heavy while interest assumptions start Relevance-heavy and both become more balanced over time. That's a genuinely useful starting point for adaptive XAI design.\n\nWhat it does well: the method section is unusually transparent. Coding manual with examples, intercoder kappa .75, clear tables, and a discussion that names real limitations, including the admission that some explainers weren't aware of knowledge gaps until the video recall interview. The practical implications (Figure 1) are concrete and traceable to the data.\n\nSoft spots are real but not fatal. The central developmental claims rest on video recall interviews, which capture retrospection, not real-time mental states. The paper's own admission that explainers often realized gaps only during recall is direct evidence that the recalled content wasn't the online representation driving adaptation. Scene selection also differs by phase (15 start, 24 middle, 11 end), so the frequency columns in Tables 3 and 5 aren't comparable across phases; the shift could be an artifact of which moments were selected. That doesn't make the findings worthless, but it means they should be framed as 'what explainers say they thought' rather than 'what they actually monitored online.' Minor: Table 2 contains percentage errors (the column totals are mislabeled: 612/991 ≈ 61.8%, not 56.98%). Also, N=9 with no inferential statistics means effect sizes are unknown, but that's normal for qualitative work and the authors say so.\n\nMy take: the paper is honest, well-scoped, and useful to the user-centered XAI community. The core developmental narrative is plausible but should be treated as hypothesis-generating. With a straightforward revision that re-frames the video recall data as retrospective accounts and removes the problematic cross-phase frequency comparisons (or at least hedges them), it's a solid contribution.\n\nRecommendation: send it to peer review. A serious referee would catch the Table 2 issue and push on the method caveat, but the paper deserves the round.","headline":"A careful qualitative study of explainers' partner models, with a real method caveat that tempers but doesn't kill the developmental story.","tokens_in":19224,"tokens_out":3633,"would_cite":true,"duration_ms":31389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human explainers continuously update a partner model of the listener's knowledge and interests, and explainable AI should do the same.","keywords":["mental representations","partner model","dual nature theory","technological artifacts","everyday explanations","explainable AI","user models","qualitative content analysis"],"falsifier":"An experiment that probes explainers' assumptions in real time—for example, asking them immediately after each explanation move, or tracking where they look and what they say—would settle the question; if live probes show no Architecture-to-both trajectory in knowledge or no Relevance-to-both trajectory in interests, the recalled developmental pattern would be shown to be a post-hoc reconstruction rather than a real-time mental representation.","tokens_in":18316,"feed_emoji":"🧠","tokens_out":10120,"duration_ms":89148,"temperature":0.7,"pith_summary":"This paper tries to establish that explainers in everyday explanations build a mental representation of the listener's knowledge and interests that develops over the course of the explanation, and that this development has a regular shape. Early assumptions about the listener's knowledge focus on the artifact's Architecture (how it is built and how it works), while early assumptions about interests center on Relevance (what it is for and why it matters). As the explanation moves into a dialogical phase, both dimensions broaden: assumed knowledge comes to include Relevance and assumed interests come to include Architecture, and vague guesses harden into well-defined beliefs. The authors' larger point is that explainable AI needs a component that tracks these same two dimensions so that machine explanations can be adapted to a user's developing needs.","feed_headline":"Explainers build a shifting model of what you know and want","feed_subtitle":"How explainers' assumptions about listeners' knowledge and interests change during an explanation.","key_machinery":"The load-bearing object is the 'partner model': the explainer's mental representation of the explainee, here decomposed into a 'technical model' (what the listener knows about the artifact) and a set of assumptions about interests. The paper maps both onto the dual nature theory of technological artifacts, which distinguishes Architecture (observable, measurable features such as pieces, rules, and structure) from Relevance (interpretable aspects such as purpose, meaning, and strategy). The argument is carried by the coded distribution of interview segments across these categories in different explanation phases—before, at the start, middle, and end of the explanation, and after it—showing how the partner model is refined and updated.","core_discovery":"On the paper's own terms, the central discovery is that explainers' partner models—their assumptions about what the explainee knows and cares about—are dynamic and dual-sided. Coding of pre-, post-, and video recall-interviews from nine explainers who explained the board game Quarto shows that assumed knowledge begins Architecture-centered and later incorporates Relevance, whereas assumed interests begin Relevance-centered and later incorporate Architecture. The patterns move from vague early assumptions to well-defined beliefs, and explainers frequently end the explanation even while believing the explainee still has gaps, treating some missing knowledge as unimportant or as requiring hands-on experience rather than more talk. These findings are offered as the empirical basis for a user-model component in adaptive explainable systems.","pith_inferences":["If the two-dimensional partner model generalizes beyond the board-game case, a compact state representation—knowledge and interest values on Architecture and Relevance—could be a practical starting point for user models in XAI, though the paper itself stops at qualitative indicators.","A natural test the paper does not run is whether explainers whose assumptions follow the observed trajectory produce explanations that explainees find more satisfying; linking the model to explanation quality would strengthen its practical relevance.","For digital artifacts with many features, the partner model may need to track which subset of features a user cares about, not just the global balance of Architecture and Relevance, because not every feature will be relevant to every user.","If the retrospective-recall worry is real, future work could validate the developmental pattern with live behavioral markers such as the timing of questions or gaze direction, which would also provide XAI with observable signals to monitor."],"forward_implications":["Adaptive explainable systems should maintain a partner model that tracks both the user's knowledge and the user's interests on the Architecture and Relevance dimensions, rather than a single measure of understanding.","An explanation system could follow the decision rules the paper derives: start when knowledge is low and interest is high, continue until the demanded knowledge level is reached, change perspective when interests shift, re-explain when knowledge stagnates, and stop when knowledge is high and interest is low.","Ending an explanation despite known knowledge gaps is a normal part of human explanation behavior, so XAI systems need not aim for complete coverage of every aspect.","The observed sequence implies that explanations of technological artifacts should typically establish Architecture first, because assumptions about Relevance knowledge develop later on that foundation."],"supporting_citations":[{"why":"Provides evidence that tutors monitor students' understanding, supporting the monitoring mechanism at the center of the partner model.","marker":"[6]"},{"why":"Shows that speakers monitor addressees for understanding, supporting the claim that explainers update their partner model during dialogue.","marker":"[8]"},{"why":"Defines the partner model as the mental representation of an interlocutor, the central construct the study investigates.","marker":"[10]"},{"why":"Defines the monological and dialogical phases of explanations used to select video recall scenes at the start, middle, and end.","marker":"[12]"},{"why":"Supplies the dual nature theory of technological artifacts—Architecture versus Relevance—used to code all knowledge and interest segments.","marker":"[26]"},{"why":"Supplies the qualitative content analysis method used to code the interview transcripts.","marker":"[37]"},{"why":"Frames explanations as co-constructive social interaction, motivating the focus on how explainers adapt to explainees' needs.","marker":"[39]"},{"why":"Provides the distinction between assumed and expressed interests used in the interest coding scheme.","marker":"[45]"},{"why":"Earlier study of everyday explanations showing an Architecture-to-Relevance shift in explanation content, which this paper extends to explainers' mental representations.","marker":"[49]"}],"fun_headline_variants":["How explainers' assumptions about you change mid-explanation","Explainers revise their model of your knowledge and interests","Partner models in everyday explanations are dynamic and dual-sided","What explainers think you know and want evolves as they talk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on the assumption that video recall-interviews give valid, undistorted access to what explainers actually thought during the explanation.","fun_headline_variants_meta":{"raw":{"variants":["How explainers' assumptions about you change mid-explanation","Explainers revise their model of your knowledge and interests","Partner models in everyday explanations are dynamic and dual-sided","What explainers think you know and want evolves as they talk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00015,"raw_usage":{"total_tokens":1207,"prompt_tokens":965,"completion_tokens":242,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":176}},"tokens_in":581,"tokens_out":242,"duration_ms":3282,"temperature":1.0,"reasoning_tokens":176,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T21:30:37.884751+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An experiment that probes explainers' assumptions in real time—for example, asking them immediately after each explanation move, or tracking where they look and what they say—would settle the question; if live probes show no Architecture-to-both trajectory in knowledge or no Relevance-to-both trajectory in interests, the recalled developmental pattern would be shown to be a post-hoc reconstruction rather than a real-time mental representation.","supporting_citations":[{"cited_title":"22, (2004)","cited_arxiv_id":null,"evidence_quote":"Provides evidence that tutors monitor students' understanding, supporting the monitoring mechanism at the center of the partner model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows that speakers monitor addressees for understanding, supporting the claim that explainers update their partner model during dialogue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the partner model as the mental representation of an interlocutor, the central construct the study investigates."},{"cited_title":"K¨ unstl","cited_arxiv_id":null,"evidence_quote":"Defines the monological and dialogical phases of explanations used to select video recall scenes at the start, middle, and end."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dual nature theory of technological artifacts—Architecture versus Relevance—used to code all knowledge and interest segments."},{"cited_title":"Springer Fachmedien Wiesbaden, Wiesbaden (2019)","cited_arxiv_id":null,"evidence_quote":"Supplies the qualitative content analysis method used to code the interview transcripts."},{"cited_title":"IEEE Trans","cited_arxiv_id":null,"evidence_quote":"Frames explanations as co-constructive social interaction, motivating the focus on how explainers adapt to explainees' needs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the distinction between assumed and expressed interests used in the interest coding scheme."},{"cited_title":"In: Longo, L","cited_arxiv_id":null,"evidence_quote":"Earlier study of everyday explanations showing an Architecture-to-Relevance shift in explanation content, which this paper extends to explainers' mental representations."}],"review_version":1}