{"id":"09beae3e-ecd1-4677-8d0f-6f7e081e50b0","arxiv_id":"2506.16997","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 66 users produces a preliminary catalog of behavior, system event, emotion, and physical reaction indicators that could signal a need for explanation.","lead":"This paper surveys 66 people about moments they needed an explanation while using software, and compiles a catalog of 17 behavior-based, 8 event-based, and 14 emotion or physical-reaction indicators that may signal such needs. The catalog is offered as a starting point for detecting explanation needs at runtime or from usage telemetry, though none of the indicators are validated yet.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-need baseline is missing: the runtime-trigger claim presupposes indicator specificity that the study design never tests, so false-positive rates are unknown.","rationale":"The reader's CONDITIONAL verdict and the identified weakest assumption (retrospective self-report and recall bias) are reasonable, but the most load-bearing weakness is more specific and more fundamental: the study has no comparison condition without explanation needs, so it cannot establish that any indicator is specific to explanation need. This is a logical gap in the inference from 'users report doing X during a need' to 'X can be used as a runtime trigger,' independent of whether participants remember accurately. The paper itself acknowledges this in Section VI.C.d, stating that no hands-on validation was performed and that conclusions about actual applicability cannot be drawn. The abstract and contribution statements nonetheless assert runtime use, creating an overclaim relative to the evidence. However, the paper is explicitly framed as a first step toward a catalog and clearly describes its limitations, so the appropriate disposition remains conditional rather than reject or unverified: the catalog is a plausible starting framework, but the runtime-trigger claim should be reframed as a research hypothesis or validated with observational data. My concern partially agrees with the reader's weakest assumption: both point to validity threats in the link between self-report and runtime detection, but the missing baseline is a design limitation distinct from recall accuracy and, in my view, the stricter test of the central claim. No personal criticism is intended; the analysis is strictly about the inferential structure of the study.","tokens_in":15007,"tokens_out":2723,"duration_ms":33224,"concrete_test":"Run a small observational study that instruments a realistic web application (e.g., a checkout or form workflow) to log the candidate behavior indicators (back-and-forth navigation, canceled actions, click spamming, inactivity, hovering) with timestamps. Let participants complete tasks while a need-annotation mechanism (e.g., an in-situ 'I need an explanation' button, or immediate cued-recall prompts after task segments) records ground-truth explanation-need intervals. Compute, for each indicator, the conditional probability given need versus given no-need, and derive precision and recall at plausible thresholds. If the likelihood ratios for the most frequent indicators (e.g., back-and-forth navigation) are near 1, or if precision at usable recall falls below an acceptable level, the runtime-trigger claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that the catalog can \"trigger explanations at appropriate moments during the runtime.\" For a trigger to be appropriate, an indicator must not only occur when an explanation need arises (recall), but also be absent, or substantially rarer, when no need exists (precision or specificity). The study design samples only episodes of need: Section III.B asks participants to recall up to five explanation needs per system (Q2.1-Q4.1) and then to describe their behavior and feelings during those needs (Q2.2-Q2.3). There is no baseline condition collecting behavior from moments without explanation needs, and no usage logs from non-need intervals. Consequently, the catalog cannot distinguish indicators of explanation need from common behaviors of exploration, indecision, frustration, or normal use. For instance, \"back-and-forth navigation\" (the most frequent behavior, 44 reports) plausibly occurs often in online shopping or browsing without any explanation need; \"hovering,\" \"inactivity,\" and \"click spamming\" are similarly non-specific. The paper's own Conclusion Validity section (VI.C.d) concedes: \"we cannot draw any conclusions about the actual applicability of the indicators,\" and Section VI.B.1 states that physical reactions \"only provide rough directions, but not yet ready-to-use methods.\" These concessions directly undercut the runtime-trigger assertion. Even granting perfect recall, the missing baseline makes the leap from self-reported correlates to runtime triggers unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an online survey study (N=66) in which participants recalled up to five explanation needs for each of three recently used software systems and described their behaviors, emotions, and physical reactions. The responses were coded into four taxonomies: 17 behavior-based, 8 event-based, 6 physical-reaction, and 8 emotional-state indicators. The authors also tabulate associations between these indicators and the need types of their earlier taxonomy (Droste et al.), and propose uses in requirements engineering, post-deployment telemetry, and runtime explanation triggering.","tokens_in":15378,"tokens_out":4115,"duration_ms":40042,"significance":"If validated, the catalog would be a practical aid for eliciting explanation requirements and for detecting explanation needs from usage data. The authors report interrater agreement (B&P kappa 0.59-0.91), share coding guidelines and pseudonymized data, and openly acknowledge several validity threats. The strength of the paper is its descriptive taxonomy: a systematic first step with transparent coding. However, the paper's central contribution as stated in the abstract—that the indicators can trigger explanations at runtime—is not supported by the study design, which relies entirely on retrospective self-reports and lacks a no-need baseline. The descriptive results are therefore valuable as hypotheses, not as validated detectors.","major_comments":[{"comment":"The runtime-trigger claim in the abstract is contradicted by the conclusion-validity section. Section VI.C.d states 'we cannot draw any conclusions about the actual applicability of the indicators.' Since the abstract claims the indicators 'can be used to trigger explanations at appropriate moments during the runtime,' this is a load-bearing overclaim. Either the abstract and the 'Real-time Explanation Triggers' use case (Section VI.B.3) must be reworded to describe a future research goal, or a validation study must be added.","section":"Abstract and Section VI.C.d"},{"comment":"The study collects only episodes in which a need was present (Section III.B, Q2.1-Q4.1); it never samples moments without explanation needs. Therefore the catalog cannot support claims about specificity or false-positive rates. For example, 'back-and-forth navigation' (44 reports) and 'click spamming' (15) are plausible in normal usage without any explanation need. Without a control condition, RQ1 'Which runtime indicators exist that signal a need for explanation' cannot be answered beyond identifying candidate correlates. Please add a baseline condition or explicitly restrict the claim to hypotheses.","section":"Section III.B and Section IV.A"},{"comment":"The mapping between indicators and need types in Section V (Tables II and III) is an internal association: the same participants reported both the need and the behavior in the same questionnaire, using prompts derived from the authors' own taxonomy [4]. This does not externally validate that the indicator causes or reliably accompanies the need; it may reflect post-hoc rationalization. The paper should present this as exploratory and state that external validation is required.","section":"Section V, Tables II and III"},{"comment":"Although framed as 'captured at runtime,' all indicators were elicited as retrospective free-text memories (Sections III.B, III.C, VI.C.b). No usage logs, clickstreams, sensor data, or system telemetry were recorded. Consequently, even the behavioral indicators are not shown to be detectable by software or sensors; physical reactions are explicitly described as 'not yet ready-to-use methods' (Section VI.B.1). Please make the gap between self-report and runtime measurability explicit in the central claims.","section":"Sections III.C, III.E, and VI.C.b"}],"minor_comments":[{"comment":"Typo: 'pseundonomized' should be 'pseudonymized'.","section":"Section III.G"},{"comment":"Typo: 'beeing' should be 'being'.","section":"Section IV.D"},{"comment":"Typo: 'posessecurity' should be 'poses security'.","section":"Section IV.B"},{"comment":"The figure labels 'Concern' with 14, but the text says 'The emotion of concern ... was indicated 30 times.' The number 30 appears in the figure under 'Confusion,' so the text and figure appear to disagree about which category has which count.","section":"Figure 6 and Section IV.D"},{"comment":"Grammar: 'These indicators may not yet ready for direct use' should be 'may not yet be ready for direct use.'","section":"Section IV.C"}],"recommendation":"major_revision","confidential_remarks":"The overclaim in the abstract is the main barrier to acceptance. The descriptive work is solid for a first-step catalog, but the authors should either reframe the central contribution as a set of candidate indicators or add a validation study with a baseline condition. I would also flag that the mapping between indicators and need types is purely internal; an external test with independent data would strengthen the paper substantially."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this paper delivers a useful first-pass catalog of self-reported indicators for explanation needs, and it is more careful in the body than the abstract claims. The 17 behavior-based and 8 event-based indicators, plus the physical/emotion list, are a new inventory in the requirements engineering / explainability space. The coding is transparent, interrater agreement is reported (Brennan-Prediger kappa from 0.59 to 0.91), and the data plus coding guidelines are shared. That is more than many qualitative SE studies do. The citation pattern leans on the authors' own prior taxonomy, but that is legitimate here since they build directly on it; the wider explainability and requirements literature is adequately covered. The authors also make a smart observation: privacy/security and domain-knowledge needs are barely visible in user behavior, which is worth knowing for anyone building detectors.\n\nThe weak spot is the gap between the abstract and the evidence. The study only samples moments where participants recall an explanation need; there is no baseline of moments without needs. So the catalog cannot support claims about specificity. Behaviors like back-and-forth navigation, inactivity, or click spamming plausibly occur in ordinary use without any explanation need, and the study gives us no way to estimate that. The paper's own Conclusion Validity section concedes that \"we cannot draw any conclusions about the actual applicability of the indicators,\" which directly undercuts the abstract's statement that these indicators can be used to trigger explanations at runtime. That framing should be softened. The indicator-to-need mapping is also internal: both needs and indicators come from the same self-reports, coded with the authors' own taxonomy, so it is an association, not independent validation. The retrospective recall issue, acknowledged in Section VI.C.b, compounds this. None of this kills the descriptive value of the catalog; it is a reasonable first step, but the strong runtime-trigger claim goes beyond the data.\n\nWho is this for? Requirements engineers and explainability researchers who want a list of candidate signals to test in observational or telemetry studies. It is not a finished detector. I would send it to peer review, but the authors should reframe the abstract and either add a no-need comparison or explicitly present runtime triggering as future work. The descriptive content deserves a serious venue; the overreach should be fixed before publication.","headline":"A useful exploratory catalog of self-reported indicators, but the abstract overclaims runtime-trigger readiness since the study lacks any no-need baseline.","tokens_in":15767,"tokens_out":4445,"would_cite":true,"duration_ms":44851,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Software can recognize when a user needs an explanation by watching for 17 behavior patterns, 8 system events, and 14 emotional or physical reactions.","keywords":["explainability","explanation needs","user-based indicators","runtime detection","user behavior","requirements elicitation","telemetry","self-reported study"],"falsifier":"Record users interacting with a software system, log the proposed behavior-based indicators (back-and-forth navigation, canceled actions, click spamming, inactivity) together with system-event indicators (errors, loading times, design deviations), and interrupt users at random moments to ask whether they currently need an explanation; if the logged signals are no more frequent during self-reported need than at other times, the catalog fails.","tokens_in":14810,"feed_emoji":"💡","tokens_out":7088,"duration_ms":64049,"temperature":0.7,"pith_summary":"The paper aims to give software a way to know when a user is confused and needs an explanation, without asking the user directly. From an online study in which 66 participants recalled moments of needing explanations in software they had recently used, the authors built a catalog of 17 behavior-based indicators, 8 system-event indicators, and 14 emotion or physical-reaction indicators. They also connect these indicators to types of explanation needs, showing for instance that back-and-forth navigation is associated with interaction, system behavior, and user-interface needs. If the indicators hold up in practice, systems could use them in prototypes, telemetry from deployed apps, and runtime triggers to deliver explanations at the moment they are wanted.","feed_headline":"17 behaviors and 8 system events flag when users need explanation","feed_subtitle":"Software could watch for these patterns and explain itself exactly when users get stuck, instead of asking.","key_machinery":"The central object is a four-part taxonomy of runtime-detectable indicators derived from coded self-reports. It carries the argument by giving engineers a modular checklist: behavior-based indicators can be logged from interaction data, event-based indicators from system internals, and emotion and physical indicators would require additional sensors. The paper's contribution is not a fully validated detector but the catalog itself, offered as a foundation for choosing and implementing indicators in specific systems.","core_discovery":"The central claim is that the need for an explanation leaves observable traces that can be captured at runtime, and that these traces can be organized into a reusable catalog. The study identifies 17 behavior-based indicators (such as back-and-forth navigation, canceled actions, repetitive actions, click spamming, and inactivity), 8 system-event indicators (such as system errors, loading times, and design deviations), and 14 indicators covering emotions and physical reactions, with facial expressions the most commonly reported physical sign and annoyance the most common emotion. The authors further report that workflow interruptions, especially back-and-forth navigation and canceled actions, are the most frequently described behavioral signals, and that privacy, security, and domain-knowledge needs are harder to detect from behavior alone.","pith_inferences":["A practical detector could likely be built from common product-analytics events alone, since session paths, clicks on non-interactive elements, and session duration already approximate several behavior-based indicators.","If these indicators are validated, the same signals could feed adaptive user interfaces that not only explain but also adjust navigation or undo support, because several indicators point to usability problems rather than explanation needs specifically.","A testable extension would combine behavior-based indicators with the reported emotions through facial-expression or biometric sensing; the paper lists the reactions, but the catalog gives a target list for training such detectors.","The frequency distribution suggests a practical starting point: workflow-interruption indicators appeared most often and are among the cheapest to log, so early adopters should begin with them."],"forward_implications":["Requirements engineers can instrument high-fidelity prototypes to record the behavior indicators and identify where test users need explanations.","Deployed applications can collect telemetry matching the behavior-based indicators to locate explanation needs after release.","At runtime, source-code instrumentation can evaluate indicators and trigger explanations at appropriate moments.","The indicator-to-need-type mappings let engineers select targeted indicators, for example back-and-forth navigation for interaction or interface needs.","The catalog gives explainability research more objective, observation-based measures of whether a need for explanation has been met."],"supporting_citations":[{"why":"Supplies the taxonomy of explanation needs used to categorize participants' reported needs.","marker":"[4]"},{"why":"Documents the 'why-not mentality' that makes direct questioning unreliable and motivates runtime indicators.","marker":"[8]"},{"why":"Defines hypothetical bias, the other elicitation pitfall the indicator catalog is designed to avoid.","marker":"[16]"},{"why":"Presents the prior explanation-on-demand technique whose impracticality motivates cheaper indicator-based detection.","marker":"[5]"},{"why":"Reports earlier biometric runtime detection whose inconclusive results set the stage for broadening to behavioral indicators.","marker":"[26]"},{"why":"Provides the kappa statistic used to measure interrater agreement in coding.","marker":"[36]"},{"why":"Defines the agreement thresholds used to interpret the reported interrater kappa values.","marker":"[38]"},{"why":"Supplies the validity-threat structure used to report limitations.","marker":"[42]"}],"fun_headline_variants":["17 behaviors and 8 system events reveal when explanations are needed","Runtime signals: 17 user behaviors and 8 events that prompt explanations","When to explain: 17 behaviors, 8 events, 14 emotions","Detect explanation needs from 17 behaviors and 8 system events"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The catalog rests on the assumption that participants' retrospective self-reports accurately capture what they actually did and felt at the moment of needing an explanation, since the indicators are built entirely from those recollections.","fun_headline_variants_meta":{"raw":{"variants":["17 behaviors and 8 system events reveal when explanations are needed","Runtime signals: 17 user behaviors and 8 events that prompt explanations","When to explain: 17 behaviors, 8 events, 14 emotions","Detect explanation needs from 17 behaviors and 8 system events"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001889,"raw_usage":{"total_tokens":7370,"prompt_tokens":869,"completion_tokens":6501,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":6424}},"tokens_in":485,"tokens_out":6501,"duration_ms":49348,"temperature":1.0,"reasoning_tokens":6424,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:13:48.239233+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record users interacting with a software system, log the proposed behavior-based indicators (back-and-forth navigation, canceled actions, click spamming, inactivity) together with system-event indicators (errors, loading times, design deviations), and interrupt users at random moments to ask whether they currently need an explanation; if the logged signals are no more frequent during self-reported need than at other times, the catalog fails.","supporting_citations":[{"cited_title":"Explanations in Everyday Software Systems: Towards a Taxonomy for Explainability Needs","cited_arxiv_id":"2404.16644","evidence_quote":"Supplies the taxonomy of explanation needs used to categorize participants' reported needs."},{"cited_title":"Designing end-user personas for explainability requirements using mixed methods research,","cited_arxiv_id":null,"evidence_quote":"Documents the 'why-not mentality' that makes direct questioning unreliable and motivates runtime indicators."},{"cited_title":"Chapter 81 experimental evidence on the existence of hypothetical bias in value elicitation methods,","cited_arxiv_id":null,"evidence_quote":"Defines hypothetical bias, the other elicitation pitfall the indicator catalog is designed to avoid."},{"cited_title":"Explanations on demand - a technique for eliciting the actual need for explanations,","cited_arxiv_id":null,"evidence_quote":"Presents the prior explanation-on-demand technique whose impracticality motivates cheaper indicator-based detection."},{"cited_title":"On the pulse of requirements elicitation: Physiological triggers and explainability needs,","cited_arxiv_id":null,"evidence_quote":"Reports earlier biometric runtime detection whose inconclusive results set the stage for broadening to behavioral indicators."},{"cited_title":"Coefficient kappa: Some uses, misuses, and alternatives,","cited_arxiv_id":null,"evidence_quote":"Provides the kappa statistic used to measure interrater agreement in coding."},{"cited_title":"Wohlin, P","cited_arxiv_id":null,"evidence_quote":"Supplies the validity-threat structure used to report limitations."}],"review_version":1}