{"id":"9f03793f-e86b-49ae-a59c-6bc61211fff1","arxiv_id":"2505.23684","paper_version":4,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Interviews were the most time-efficient elicitation method for explainability needs, surveys yielded the largest volume with high redundancy, and introducing an explanation-need taxonomy only after free elicitation produced more and more diverse needs.","lead":"This paper ran a field comparison of three ways to collect what users want explained in software: focus groups, one-on-one interviews, and online surveys, using a German company's HR system as the test case. It found interviews give the most useful needs per hour, surveys produce the most ideas but with heavy repetition, and that letting users speak freely before showing them a category list yields richer results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table III's efficiency metric cannot be reproduced from the paper's own durations: the 'interviews are most efficient' headline rests on values that do not recompute.","rationale":"The reader's verdict CONDITIONAL is reasonable. I focused on the efficiency metric because it is the most concrete and checkable flaw in the paper's central claim: the exact numbers that support 'interviews were the most efficient' cannot be derived from the paper's own stated durations and participant counts. This is not a matter of statistical interpretation or coder judgment; it is an internal arithmetic inconsistency. If the table is wrong, the RQ1 answer is unsupported even if the underlying qualitative trend is plausible. The coding-reliability threat identified by the reader is also real and is acknowledged by the authors in Section V.C, but it is a standard limitation that could be addressed by an inter-rater reliability study. The metric problem is more immediate: any reviewer can verify it from the published text alone. I therefore agree that the paper should be CONDITIONAL pending a corrected, reproducible quantitative analysis; the dataset is a valuable deliverable and the descriptive observations about low overlap and redundancy are worth keeping. The concern does not change the verdict, but it narrows the condition: before the efficiency ranking can be accepted, the authors must either provide the exact formula and recomputed table or re-analyze the data with a clearly defined metric.","tokens_in":16500,"tokens_out":13851,"duration_ms":118125,"concrete_test":"Download the Zenodo dataset (doi:10.5281/zenodo.15678937); independently recompute Table III's 'Average total time', 'Personnel effort', 'Distinct needs per participant per average time', and 'Distinct needs per personnel effort' from the raw response timestamps and the Section III.D.1 formula. Also check whether the average interview durations in Section III.C.2 (11:07, 11:53, 15:53 minutes) match the Table III total times (23:28, 33:59, 38:13 minutes). If the recomputed values differ from the published ones, the RQ1 efficiency ranking must be re-derived; if they match, the discrepancy is a reporting error in the text that should still be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline efficiency result (RQ1) rests entirely on the two efficiency columns of Table III, yet those numbers cannot be reproduced from the paper's own definitions and published durations. Section III.D.1 defines personal effort as total study time multiplied by number of participants, and Section III.C.2 reports the average durations of the interview groups as 11:07, 11:53, and 15:53 minutes. Table III instead lists average total times of 23:28, 33:59, and 38:13 minutes for those same groups, an unexplained factor-of-two discrepancy. Recomputing the key metric from the table's own rows fails for most cells: for interviews without taxonomy, 96 distinct needs among 9 participants with an average total time of 23.47 minutes gives 96/(9*23.47)=0.45, not the reported 0.66; for focus groups without taxonomy, 19/(6*27.6)=0.11, not 0.15; for surveys without taxonomy, 327/(188*11.4)=0.15, not 0.25. The per-personnel-effort column also deviates from the stated formula in several rows. Since the abstract and Section V.A claim interviews are 'most efficient' based on these exact values, the central comparative claim is not currently supported by the paper's own data. The shipped Zenodo dataset could resolve this, but as written the efficiency ranking is an unexplained artifact of the numbers in the table.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a comparative case study of three requirements elicitation methods—focus groups, interviews, and online surveys—for collecting explainability requirements from users of a personnel management system at a German IT consulting company. The study uses an existing explanation-need taxonomy (Droste et al. [4]) and compares three conditions: no taxonomy, direct taxonomy introduction, and delayed taxonomy introduction. The central claims are that interviews are the most efficient method (highest distinct needs per participant per time), surveys collect the highest absolute number of distinct needs, and delayed taxonomy introduction increases the number and diversity of elicited needs. The paper recommends a hybrid survey-plus-interview strategy. The analysis is based on hand-coded counts of distinct needs, with efficiency metrics defined in Section III.D.1 and reported in Table III. The paper explicitly acknowledges that no statistical tests were conducted and that inter-rater reliability was not assessed.","tokens_in":16661,"tokens_out":3559,"duration_ms":33418,"significance":"If substantiated, the findings would offer practically useful, concrete guidance for requirements engineers choosing among elicitation methods for explainability requirements, and the comparison of direct versus delayed taxonomy introduction is a valuable design contribution. The paper has notable strengths: it combines three methods in one study, distinguishes taxonomy timing conditions, openly publishes its dataset (Zenodo [48]), and provides a transparent threats-to-validity discussion that includes the missing statistical tests, single-coder coding, and single-company context. However, the central efficiency claim currently rests on numbers in Table III that cannot be reproduced from the paper's own definitions and reported durations. Because the headline result (RQ1) depends on these irreproducible values, the contribution is not yet fully supported by the manuscript as written.","major_comments":[{"comment":"The efficiency metrics in Table III do not recompute from the stated definitions and reported durations. Section III.D.1 defines personal effort as total study time multiplied by the number of participants, and Table III reports average total times. For interviews without taxonomy: 96 distinct needs, 9 participants, average total time 23:28 (23.47 min) gives 96/(9×23.47) = 0.45, not the reported 0.66, and 9×23.47 min = 3.52 h, not the reported 7:34 h. Similar mismatches appear for focus groups without taxonomy (19/(6×27.6)=0.11 vs. 0.15), surveys without taxonomy (327/(188×11.4)=0.15 vs. 0.25), and several personnel-effort rows. Since the abstract and Section V.A use these exact values to conclude that interviews were the most efficient, the central comparative claim is not currently supported by the paper's own data. The authors should report how these values were computed, correct or justify each cell, and ideally recompute RQ1 with a clear, reproducible formula; the Zenodo dataset could resolve the discrepancy if the raw durations and formulas are supplied.","section":"Table III / Section III.D.1"},{"comment":"The interview durations reported in the method section are inconsistent with the 'average total time' in Table III. Section III.C.2 reports average interview durations of 11:07, 11:53, and 15:53 minutes for the three interview groups, but Table III lists average total times of 23:28, 33:59, and 38:13 minutes for the corresponding groups. The factor-of-two gap is unexplained. The paper should clarify whether 'total study time' in Section III.D.1 includes additional pre- or post-interview work (e.g., preparation, analysis, or moderator overhead), and if so, define it precisely so that the efficiency calculation is reproducible and the claims in Section V.A are grounded.","section":"Section III.C.2 vs. Table III"},{"comment":"The paper concedes that 'no statistical tests were conducted to assess the significance of differences observed between elicitation methods,' yet the abstract, Section V.A, and Section VI state comparative conclusions such as 'interviews were the most efficient' and 'delayed taxonomy usage led to the highest number of distinct needs per participant.' With only two focus groups (n=12 total) and 18 interviews, these differences may be well within sampling variation. The authors should either add appropriate inferential or non-parametric tests (e.g., bootstrap confidence intervals for per-participant rates) or explicitly reword the claims as descriptive observations from a single case study, not as confirmed rankings.","section":"Section V.C.3"},{"comment":"The dependent variable—distinct explanation needs—rests on the manual coding of one requirements engineer, with a second engineer consulted only in cases of uncertainty, and the paper acknowledges that inter-rater reliability 'was not explicitly measured.' Because all four research questions and the efficiency/effectiveness rankings hinge on these distinct-need counts, coder subjectivity is a load-bearing threat. The authors should provide at least a formal inter-rater reliability assessment on a sample of responses (and ideally report per-method agreement), or otherwise provide a sensitivity analysis showing the conclusions are robust to plausible re-coding. Without this, the method-level comparisons may partly reflect coding judgment rather than elicitation-method differences.","section":"Section III.D.2 / Section V.C.1"}],"minor_comments":[{"comment":"The column 'Distinct needs per participant per average time' appears to be scaled incorrectly in several rows: for example, focus groups without taxonomy (3.17 distinct per participant / 27.6 min) yields 0.115, not 0.15; surveys without taxonomy (1.74 / 11.4 min) yields 0.153, not 0.25. Please verify all cells and state the unit (e.g., per minute).","section":"Table III"},{"comment":"The text in Section IV states that for the 'without taxonomy' condition, 24 needs were shared between surveys and interviews and one need across all three methods, but the Venn diagram in Figure 5a appears to show different overlap values (2, 2, 191) and the plotted total does not match the reported counts (327+96+19 plus overlaps). Please correct the figure or the text, and make the overlap arithmetic consistent.","section":"Figure 5 / Section IV"},{"comment":"The phrase 'the two focus groups differed in composition' is followed by a description of the two groups, which is helpful; however, the later statement that 'the durations of all three elicitation methods were comparable' seems inconsistent with the large differences in Table III and the reported durations. Consider rewording this threat for clarity.","section":"Section V.C.2"},{"comment":"The paper refers to 'the most efficient' and 'the most effective' in the abstract and Section V.A without formally defining the precise ordering rule for each metric. For example, interviews are 'most efficient' by the per-participant-per-time metric, but the text also notes that surveys have the highest absolute number of distinct needs. A short definition or table caption explaining which metric defines 'efficiency' and 'effectiveness' would prevent confusion.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's strongest contribution is the empirical design and the public dataset; the central efficiency claim, however, is currently undercut by internally inconsistent numbers that a careful reader cannot reproduce. I recommend major revision rather than rejection because the authors can plausibly correct the computation, clarify the time basis, add or explicitly forgo significance testing, and strengthen the coding-reliability evidence. The paper's reliance on the authors' own taxonomy and their prior coding guidelines is a self-referential element, but it is not by itself a reason to reject; the actionable issue is the absence of reproducibility of the reported metrics."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the two things you should know. First, this is one of the first papers to actually run focus groups, interviews, and a survey head-to-head for explainability requirements, and it ships the dataset on Zenodo. That is a real contribution. Second, the central efficiency result is currently not supported by the paper's own Table III. The stress-test note is right: the efficiency values do not recompute from the reported durations. For example, interviews without taxonomy are reported at 23:28 average total time in Table III, while Section III.C.2 says the same group averaged 11:07. That factor-of-two discrepancy alone makes 'distinct needs per participant per average time' of 0.66 unexplained. Recomputing from the table's own rows gives 0.45 for that cell. Similar mismatches appear in most rows. The per-personnel-effort column also deviates from the stated formula.\n\nWhat the paper does well: it is honest and well-structured. The threats section admits no statistical tests, single-coder classification with no inter-rater reliability, and the confound that 'without taxonomy' and 'delayed taxonomy' come from the same participants. The Venn analysis and the category-level breakdowns are useful descriptions. The delayed-taxonomy finding is plausible but confounded with extra time on task.\n\nWhere the soft spots are: the efficiency ranking (RQ1) is the load-bearing claim in the abstract, and it rests on those non-recomputing numbers. The effectiveness ranking (RQ2) depends on counts of distinct needs from one coder using the authors' own taxonomy, with no reliability check. So the quantitative comparisons are best treated as hypotheses, not confirmed results. The paper's own conclusion overstates them.\n\nWho is this for? Requirements engineering researchers who want a worked example and dataset for comparing elicitation methods, and anyone planning a similar study. It deserves a serious referee because the dataset and template are valuable and the flaws are fixable with a revised metric definition, a re-analysis of the shipped data, and a small reliability study.\n\nMy recommendation: send it to peer review with a request for major revision fixing Table III and re-running the analysis; don't desk reject.","headline":"Useful empirical template and dataset for explainability elicitation, but the headline efficiency ranking doesn't recompute from the paper's own numbers.","tokens_in":17318,"tokens_out":2242,"would_cite":true,"duration_ms":19636,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A case study with 188 survey respondents, 18 interviewees, and 12 focus-group participants finds interviews are the most efficient method for eliciting explainability requirements, surveys maximize volume but repeat themselves, and…","keywords":["explainability requirements","requirements elicitation","interviews","focus groups","online surveys","taxonomy","non-functional requirements","case study"],"falsifier":"Recode the raw responses from the 188 surveys, 18 interviews, and two focus groups with two independent coders using the same taxonomy; if the recomputed distinct-need counts no longer rank interviews above surveys on per-participant-per-hour efficiency, the central claim fails.","tokens_in":16188,"feed_emoji":"🗣️","tokens_out":7781,"duration_ms":68390,"temperature":0.7,"pith_summary":"The paper asks which of three common elicitation methods—focus groups, interviews, and online surveys—best captures users' explainability requirements, and whether the timing of a taxonomy matters. Using a web-based personnel management system at a large German IT consulting company, it reports that interviews were the most efficient method, producing the most distinct needs per participant per unit of time. Surveys collected the most needs overall but with roughly 20% to 23% redundancy, and introducing a taxonomy only after an initial open elicitation phase produced more and more diverse needs than front-loading it. The practical upshot is that a hybrid survey-plus-interview approach, with taxonomy introduced after free elicitation, balances breadth and depth better than any single method.","feed_headline":"Interviews win on efficiency for explainability needs","feed_subtitle":"In a 218-person case study, surveys gather the most needs but repeat; delaying a taxonomy boosts diversity.","key_machinery":"The carrying machinery is a comparison of elicitation methods measured by distinct explanation needs per participant per unit time and per personnel effort, where personnel effort multiplies session duration by participant count. All responses are coded into an extended version of a five-category taxonomy of explanation needs—interaction, system behavior, privacy and security, domain knowledge, and user interface, plus software-specific additions such as feature missing and business needs. The taxonomy works both as an elicitation checklist and as the coding instrument that turns raw statements into countable needs. The second mechanism is the two-condition design: direct taxonomy usage from the outset versus delayed taxonomy usage after an open phase, which isolates the effect of when structure is introduced.","core_discovery":"On the authors' own terms, the paper establishes that interviews, not surveys or focus groups, are the most efficient way to elicit explainability requirements, because they yield the highest number of distinct explanation needs per participant per time spent and per personnel effort. It also establishes that surveys are the most effective in absolute volume, but their redundancy—20.05% without taxonomy and 22.72% with it—undercuts per-participant diversity. Finally, it establishes that introducing the explanation-need taxonomy only after an initial open elicitation phase yields more and more diverse needs than front-loading it, with delayed interviews reaching 14.78 distinct needs per participant versus 11.67 under direct usage. The paper concludes that no single method is universally best: interviews maximize efficiency, surveys maximize coverage, and a two-phase hybrid approach is recommended.","pith_inferences":["Beyond what the paper tests, the same delayed-taxonomy mechanism may generalize to other non-functional requirements, because the mechanism is about when structure is imposed rather than about explainability itself.","The paper's own redundancy figures imply a cost model it does not build: if duplicate survey needs cost as much to process as distinct ones, the volume advantage of surveys shrinks once processing effort is priced in.","An untested extension would randomly assign participants to direct versus delayed taxonomy conditions instead of measuring both in the same session, which would separate the taxonomy's effect from practice, fatigue, and order effects."],"forward_implications":["A requirements engineer with a limited budget should choose interviews over surveys or focus groups when the goal is the number of distinct explainability needs collected per hour.","A team needing broad coverage should run a survey, accepting that 20% to 23% of the collected needs will duplicate earlier ones.","Elicitation should start with an open phase and introduce a taxonomy afterward; front-loading the taxonomy reduces the number and diversity of needs, especially in interviews.","Because each method captures largely different need categories, relying on any single method leaves categories uncovered; a hybrid survey-plus-interview design is the paper's recommended path."],"supporting_citations":[{"why":"Supplies the explanation-need taxonomy used as the elicitation checklist and the basis for coding all responses into categories.","marker":"[4]"},{"why":"Provides the extension of the taxonomy with software-specific categories that lets each response be split into multiple distinct needs.","marker":"[9]"},{"why":"Identifies interviews, focus groups, workshops, surveys, and personas as effective methods for eliciting explainability requirements, framing the method selection.","marker":"[11]"},{"why":"Systematic literature review supporting the claim that combining multiple elicitation methods improves coverage.","marker":"[32]"},{"why":"Source for the argument that method selection and combination affect the quality and accuracy of collected requirements.","marker":"[35]"},{"why":"Empirical guideline that experts prefer conversational methods for deep domain knowledge and questionnaires when analysts guide users, used to justify the method comparison.","marker":"[37]"},{"why":"Provides the goal-definition template and the threats-to-validity structure used to frame the research design and limitations.","marker":"[46]"}],"fun_headline_variants":["Interviews beat surveys for efficient explainability elicitation","Delay taxonomy to boost diversity in requirement elicitation","Hybrid surveys plus interviews best for explainability needs","Study: interviews most efficient for explainability requirements"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings rest on counts of distinct explanation needs produced by one coder's manual application of an extended taxonomy, with no measured inter-rater reliability; a different coder could produce different counts and a different ranking.","fun_headline_variants_meta":{"raw":{"variants":["Interviews beat surveys for efficient explainability elicitation","Delay taxonomy to boost diversity in requirement elicitation","Hybrid surveys plus interviews best for explainability needs","Study: interviews most efficient for explainability requirements"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000142,"raw_usage":{"total_tokens":1154,"prompt_tokens":917,"completion_tokens":237,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":176}},"tokens_in":533,"tokens_out":237,"duration_ms":2536,"temperature":1.0,"reasoning_tokens":176,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:40:42.601722+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recode the raw responses from the 188 surveys, 18 interviews, and two focus groups with two independent coders using the same taxonomy; if the recomputed distinct-need counts no longer rank interviews above surveys on per-participant-per-hour efficiency, the central claim fails.","supporting_citations":[{"cited_title":"Requirements elicitation tech- niques: a systematic literature review based on the maturity of the techniques,","cited_arxiv_id":null,"evidence_quote":"Systematic literature review supporting the claim that combining multiple elicitation methods improves coverage."},{"cited_title":"A model for evaluating requirements elicitation techniques in software development projects","cited_arxiv_id":null,"evidence_quote":"Source for the argument that method selection and combination affect the quality and accuracy of collected requirements."},{"cited_title":"A practical guide to requirements elicitation techniques selection-an empirical study,","cited_arxiv_id":null,"evidence_quote":"Empirical guideline that experts prefer conversational methods for deep domain knowledge and questionnaires when analysts guide users, used to justify the method comparison."}],"review_version":1}