Pith. sign in

REVIEW 4 major objections 4 minor 48 references

How to Elicit Explainability Requirements? A Comparison of Interviews, Focus Groups, and Surveys

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A case study with 188 survey respondents, 18 interviewees, and 12 focus-group participants finds interviews are the most efficient method for eliciting explainability requirements, surveys maximize volume but repeat themselves, and…

desk verdict Useful empirical template and dataset for explainability elicitation, but the headline efficiency ranking doesn't recompute from the paper's own numbers. read the letter →

arxiv 2505.23684 v4 pith:Y5FMSUGP submitted 2025-05-29 cs.SE

classification cs.SE
keywords explainabilityrequirementselicitationinterviewsfocusgroupsonlinesurveystaxonomynon-functionalcasestudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks which of three common elicitation methods—focus groups, interviews, and online surveys—best captures users' explainability requirements, and whether the timing of a taxonomy matters. Using a web-based personnel management system at a large German IT consulting company, it reports that interviews were the most efficient method, producing the most distinct needs per participant per unit of time. Surveys collected the most needs overall but with roughly 20% to 23% redundancy, and introducing a taxonomy only after an initial open elicitation phase produced more and more diverse needs than front-loading it. The practical upshot is that a hybrid survey-plus-interview approach, with taxonomy introduced after free elicitation, balances breadth and depth better than any single method.

What carries the argument

The carrying machinery is a comparison of elicitation methods measured by distinct explanation needs per participant per unit time and per personnel effort, where personnel effort multiplies session duration by participant count. All responses are coded into an extended version of a five-category taxonomy of explanation needs—interaction, system behavior, privacy and security, domain knowledge, and user interface, plus software-specific additions such as feature missing and business needs. The taxonomy works both as an elicitation checklist and as the coding instrument that turns raw statements into countable needs. The second mechanism is the two-condition design: direct taxonomy usage from the outset versus delayed taxonomy usage after an open phase, which isolates the effect of when structure is introduced.

What would settle it

Recode the raw responses from the 188 surveys, 18 interviews, and two focus groups with two independent coders using the same taxonomy; if the recomputed distinct-need counts no longer rank interviews above surveys on per-participant-per-hour efficiency, the central claim fails.

Watch

Extended reading notes

Core claim

On the authors' own terms, the paper establishes that interviews, not surveys or focus groups, are the most efficient way to elicit explainability requirements, because they yield the highest number of distinct explanation needs per participant per time spent and per personnel effort. It also establishes that surveys are the most effective in absolute volume, but their redundancy—20.05% without taxonomy and 22.72% with it—undercuts per-participant diversity. Finally, it establishes that introducing the explanation-need taxonomy only after an initial open elicitation phase yields more and more diverse needs than front-loading it, with delayed interviews reaching 14.78 distinct needs per participant versus 11.67 under direct usage. The paper concludes that no single method is universally best: interviews maximize efficiency, surveys maximize coverage, and a two-phase hybrid approach is recommended.

Load-bearing premise

The rankings rest on counts of distinct explanation needs produced by one coder's manual application of an extended taxonomy, with no measured inter-rater reliability; a different coder could produce different counts and a different ranking.

Editorial extensions

If this is right

  • A requirements engineer with a limited budget should choose interviews over surveys or focus groups when the goal is the number of distinct explainability needs collected per hour.
  • A team needing broad coverage should run a survey, accepting that 20% to 23% of the collected needs will duplicate earlier ones.
  • Elicitation should start with an open phase and introduce a taxonomy afterward; front-loading the taxonomy reduces the number and diversity of needs, especially in interviews.
  • Because each method captures largely different need categories, relying on any single method leaves categories uncovered; a hybrid survey-plus-interview design is the paper's recommended path.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond what the paper tests, the same delayed-taxonomy mechanism may generalize to other non-functional requirements, because the mechanism is about when structure is imposed rather than about explainability itself.
  • The paper's own redundancy figures imply a cost model it does not build: if duplicate survey needs cost as much to process as distinct ones, the volume advantage of surveys shrinks once processing effort is priced in.
  • An untested extension would randomly assign participants to direct versus delayed taxonomy conditions instead of measuring both in the same session, which would separate the taxonomy's effect from practice, fatigue, and order effects.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper reports a comparative case study of three requirements elicitation methods—focus groups, interviews, and online surveys—for collecting explainability requirements from users of a personnel management system at a German IT consulting company. The study uses an existing explanation-need taxonomy (Droste et al. [4]) and compares three conditions: no taxonomy, direct taxonomy introduction, and delayed taxonomy introduction. The central claims are that interviews are the most efficient method (highest distinct needs per participant per time), surveys collect the highest absolute number of distinct needs, and delayed taxonomy introduction increases the number and diversity of elicited needs. The paper recommends a hybrid survey-plus-interview strategy. The analysis is based on hand-coded counts of distinct needs, with efficiency metrics defined in Section III.D.1 and reported in Table III. The paper explicitly acknowledges that no statistical tests were conducted and that inter-rater reliability was not assessed.

Significance. If substantiated, the findings would offer practically useful, concrete guidance for requirements engineers choosing among elicitation methods for explainability requirements, and the comparison of direct versus delayed taxonomy introduction is a valuable design contribution. The paper has notable strengths: it combines three methods in one study, distinguishes taxonomy timing conditions, openly publishes its dataset (Zenodo [48]), and provides a transparent threats-to-validity discussion that includes the missing statistical tests, single-coder coding, and single-company context. However, the central efficiency claim currently rests on numbers in Table III that cannot be reproduced from the paper's own definitions and reported durations. Because the headline result (RQ1) depends on these irreproducible values, the contribution is not yet fully supported by the manuscript as written.

major comments (4)
  1. [Table III / Section III.D.1] The efficiency metrics in Table III do not recompute from the stated definitions and reported durations. Section III.D.1 defines personal effort as total study time multiplied by the number of participants, and Table III reports average total times. For interviews without taxonomy: 96 distinct needs, 9 participants, average total time 23:28 (23.47 min) gives 96/(9×23.47) = 0.45, not the reported 0.66, and 9×23.47 min = 3.52 h, not the reported 7:34 h. Similar mismatches appear for focus groups without taxonomy (19/(6×27.6)=0.11 vs. 0.15), surveys without taxonomy (327/(188×11.4)=0.15 vs. 0.25), and several personnel-effort rows. Since the abstract and Section V.A use these exact values to conclude that interviews were the most efficient, the central comparative claim is not currently supported by the paper's own data. The authors should report how these values were computed, correct or justify each cell, and ideally recompute RQ1 with a clear, reproducible formula; the Zenodo dataset could resolve the discrepancy if the raw durations and formulas are supplied.
  2. [Section III.C.2 vs. Table III] The interview durations reported in the method section are inconsistent with the 'average total time' in Table III. Section III.C.2 reports average interview durations of 11:07, 11:53, and 15:53 minutes for the three interview groups, but Table III lists average total times of 23:28, 33:59, and 38:13 minutes for the corresponding groups. The factor-of-two gap is unexplained. The paper should clarify whether 'total study time' in Section III.D.1 includes additional pre- or post-interview work (e.g., preparation, analysis, or moderator overhead), and if so, define it precisely so that the efficiency calculation is reproducible and the claims in Section V.A are grounded.
  3. [Section V.C.3] The paper concedes that 'no statistical tests were conducted to assess the significance of differences observed between elicitation methods,' yet the abstract, Section V.A, and Section VI state comparative conclusions such as 'interviews were the most efficient' and 'delayed taxonomy usage led to the highest number of distinct needs per participant.' With only two focus groups (n=12 total) and 18 interviews, these differences may be well within sampling variation. The authors should either add appropriate inferential or non-parametric tests (e.g., bootstrap confidence intervals for per-participant rates) or explicitly reword the claims as descriptive observations from a single case study, not as confirmed rankings.
  4. [Section III.D.2 / Section V.C.1] The dependent variable—distinct explanation needs—rests on the manual coding of one requirements engineer, with a second engineer consulted only in cases of uncertainty, and the paper acknowledges that inter-rater reliability 'was not explicitly measured.' Because all four research questions and the efficiency/effectiveness rankings hinge on these distinct-need counts, coder subjectivity is a load-bearing threat. The authors should provide at least a formal inter-rater reliability assessment on a sample of responses (and ideally report per-method agreement), or otherwise provide a sensitivity analysis showing the conclusions are robust to plausible re-coding. Without this, the method-level comparisons may partly reflect coding judgment rather than elicitation-method differences.
minor comments (4)
  1. [Table III] The column 'Distinct needs per participant per average time' appears to be scaled incorrectly in several rows: for example, focus groups without taxonomy (3.17 distinct per participant / 27.6 min) yields 0.115, not 0.15; surveys without taxonomy (1.74 / 11.4 min) yields 0.153, not 0.25. Please verify all cells and state the unit (e.g., per minute).
  2. [Figure 5 / Section IV] The text in Section IV states that for the 'without taxonomy' condition, 24 needs were shared between surveys and interviews and one need across all three methods, but the Venn diagram in Figure 5a appears to show different overlap values (2, 2, 191) and the plotted total does not match the reported counts (327+96+19 plus overlaps). Please correct the figure or the text, and make the overlap arithmetic consistent.
  3. [Section V.C.2] The phrase 'the two focus groups differed in composition' is followed by a description of the two groups, which is helpful; however, the later statement that 'the durations of all three elicitation methods were comparable' seems inconsistent with the large differences in Table III and the reported durations. Consider rewording this threat for clarity.
  4. [Throughout] The paper refers to 'the most efficient' and 'the most effective' in the abstract and Section V.A without formally defining the precise ordering rule for each metric. For example, interviews are 'most efficient' by the per-participant-per-time metric, but the text also notes that surveys have the highest absolute number of distinct needs. A short definition or table caption explaining which metric defines 'efficiency' and 'effectiveness' would prevent confusion.

Circularity Check

1 steps flagged · score 6.0 of 10

RQ4's delayed-taxonomy benefit reduces to the study design: the delayed condition is defined as the without-taxonomy phase plus an additional taxonomy-guided phase, so the comparison is subset-superset by construction.

  1. self definitional [Section III.C 'Methodology to Compare the Different Elicitation Methods'; Section V.C 'Threats to Validity']
    "A potential threat is that “no taxonomy usage” and “delayed taxonomy usage” were not examined in separate studies, making the “no taxonomy usage” data a subset of the “delayed taxonomy usage” data. ... In the second, needs were first collected openly without the taxonomy, after which the taxonomy was introduced to gather additional requirements (“delayed taxonomy usage”). This design allowed us to compare the without and delayed taxonomy conditions within the same group of participants."

    The delayed-taxonomy condition is defined in the same participant group as the open 'without' phase followed by an additional taxonomy-guided phase, so the delayed count is the without count plus any new needs by construction. RQ4's conclusion that 'no taxonomy usage produced fewer needs overall' and that delayed usage yielded 'a greater number ... of needs' is therefore entailed by the measurement design rather than established by the data. The paper explicitly acknowledges that the 'without' data are a subset of the 'delayed' data.

full rationale

The paper is primarily an empirical comparison, not a derivation, so most of its claims are not circular. The self-citations to the authors' own taxonomy are used as a measurement instrument and do not by themselves make the method-comparison claims circular. However, one central claim—that delayed taxonomy introduction produces more explanation needs—rests on a comparison where the delayed condition is, by design, a superset of the without-taxonomy condition for the same participants. The paper's own threats-to-validity section concedes this subset relationship, making the 'delayed > without' result a logical consequence of the design rather than an empirical effect. Other apparent problems, such as Table III efficiency values that do not recompute from the reported durations, are internal-consistency or correctness concerns, not circularity, and are not counted in this score. Because the delayed-taxonomy recommendation is a headline contribution and partially reduces to the design definition, the circularity score is 6.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

This is an empirical study, so the ledger records analytic choices rather than fitted physical constants. The free parameters are the survey exclusion rule, the hand-extended taxonomy coding categories, and the ambiguous time bases behind the efficiency metric. The axioms are domain assumptions: the taxonomy captures the space of explanation needs, self-reports are a valid proxy, single-coder classification is accurate enough, and the single-company case supports comparative inference. No invented entities are introduced.

free parameters (3)
  • Survey valid-response exclusion rule = 188 of 277 completed responses retained
    Excluding responses like 'I have no needs' and nonsensical text was a post-hoc, non-pre-specified rule that raises the average number of needs per survey participant.
  • Extended taxonomy coding categories = Droste et al. [4] base plus extensions from Obaidi et al. [9] and additional software-specific categories
    Hand-selected categories and single-coder assignment determine the counts of distinct needs that drive every research question in the paper.
  • Time basis for the efficiency metric = Inconsistent bases across methods in Table III
    The 'distinct needs per participant per average time' column does not recompute from the reported average total or elicitation times for most rows, so the choice of time window materially affects the RQ1 answer.
assumptions (4)
  • domain assumption The taxonomy by Droste et al. [4] faithfully represents the space of end-user explanation needs.
    Invoked in Section III.C; all needs are coded into its categories, and the taxonomy-timing results are interpreted through it.
  • domain assumption Self-reported explanation needs in a staged elicitation session are a valid proxy for real-world explanation needs.
    The paper itself cites hypothetical-bias and 'why-not mentality' concerns in Section II.B.1, then relies on self-reports as the only data source.
  • domain assumption One requirements engineer's coding with ad hoc consultation is accurate enough for cross-method comparison.
    Section III.D.2 describes the coding; Section V.C admits inter-rater reliability was not measured.
  • domain assumption The single-company German HR-software case is informative for comparing elicitation methods.
    External validity threat acknowledged in Section V.C.4; the comparison assumes the method-level differences are not artifacts of this one domain.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How to Elicit Explainability Requirements? A Comparison of Interviews, Focus Groups, and Surveys." pith.science (2026). https://pith.science/paper/Y5FMSUGP

@misc{pith2026250523684,
  author       = {Pith},
  title        = {Pith review of: How to Elicit Explainability Requirements? A Comparison of Interviews, Focus Groups, and Surveys},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Y5FMSUGP}},
  note         = {Machine review of arXiv:2505.23684}
}
read the original abstract

As software systems grow increasingly complex, explainability has become a crucial non-functional requirement for transparency, user trust, and regulatory compliance. Eliciting explainability requirements is challenging, as different methods capture varying levels of detail and structure. This study examines the efficiency and effectiveness of three commonly used elicitation methods - focus groups, interviews, and online surveys - while also assessing the role of taxonomy usage in structuring and improving the elicitation process. We conducted a case study at a large German IT consulting company, utilizing a web-based personnel management software. A total of two focus groups, 18 interviews, and an online survey with 188 participants were analyzed. The results show that interviews were the most efficient, capturing the highest number of distinct needs per participant per time spent. Surveys collected the most explanation needs overall but had high redundancy. Delayed taxonomy introduction resulted in a greater number and diversity of needs, suggesting that a two-phase approach is beneficial. Based on our findings, we recommend a hybrid approach combining surveys and interviews to balance efficiency and coverage. Future research should explore how automation can support elicitation and how taxonomies can be better integrated into different methods.

Figures

Figures reproduced from arXiv: 2505.23684 by the authors.

Figure 1
Figure 1. Overview of our research design in FLOW notation [45]. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 3
Figure 3. Number of explanation needs per participant with direct [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 2
Figure 2. Number of explanation needs per participant without [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Number of explanation needs per participant with [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

48 extracted references · 28 canonical work pages

  1. [4]

    Explanations in everyday software systems: Towards a taxonomy for explainability needs,

    J. Droste, H. Deters, M. Obaidi, and K. Schneider, “Explanations in everyday software systems: Towards a taxonomy for explainability needs,” in 2024 IEEE 32nd International Requirements Engineering Conference (RE), 2024, pp. 55–66

  2. [48]

    Dataset: How to elicit explainability requirements? a comparison of interviews, focus groups, and surveys,

    M. Obaidi, J. Droste, H. Deters, M. Herrmann, R. Ochsner, J. Kl ¨under, and K. Schneider, “Dataset: How to elicit explainability requirements? a comparison of interviews, focus groups, and surveys,” Jun. 2025. [Online]. Available: https://doi.org/10.5281/zenodo.15678937

  3. [1]

    Explainability as a non-functional requirement,

    M. A. K ¨ohl, K. Baum, M. Langer, D. Oster, T. Speith, and D. Bohlender, “Explainability as a non-functional requirement,” in 2019 IEEE 27th International Requirements Engineering Conference (RE). IEEE, 2019, pp. 363–368

  4. [2]

    Explainability as a non-functional requirement: challenges and recommendations,

    L. Chazette and K. Schneider, “Explainability as a non-functional requirement: challenges and recommendations,” REJ, vol. 25, no. 4, 2020

  5. [3]

    Exploring explainability: a definition, a model, and a knowledge catalogue,

    L. Chazette, W. Brunotte, and T. Speith, “Exploring explainability: a definition, a model, and a knowledge catalogue,” in RE. IEEE, 2021

  6. [5]

    Context, content, consent- how to design user-centered privacy explanations (s)

    W. Brunotte, J. Droste, and K. Schneider, “Context, content, consent- how to design user-centered privacy explanations (s).” in SEKE, 2023, pp. 86–89

  7. [6]

    Privacy explanations–a means to end-user trust,

    W. Brunotte, A. Specht, L. Chazette, and K. Schneider, “Privacy explanations–a means to end-user trust,” JSS, vol. 195, 2023

  8. [7]

    Explainable security,

    L. Vigano and D. Magazzeni, “Explainable security,” in 2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW). IEEE, 2020, pp. 293–300

Show all 48 references
  1. [8]

    The x factor: On the relationship between user experi- ence and explainability,

    H. Deters, J. Droste, A. Hess, V . Kl ¨os, K. Schneider, T. Speith, and A. V ogelsang, “The x factor: On the relationship between user experi- ence and explainability,” in Proceedings of the 13th Nordic Conference on Human-Computer Interaction , ser. NordiCHI ’24. ACM, 2024

  2. [9]

    Automating explanation need management in app reviews: A case study from the navigation app industry,

    M. Obaidi, N. V oß, J. Droste, H. Deters, M. Herrmann, J. Fischbach, and K. Schneider, “Automating explanation need management in app reviews: A case study from the navigation app industry,” in ICSE- SEIP’25, 2025

  3. [10]

    Peeking inside the black-box: a survey on explainable artificial intelligence (xai),

    A. Adadi and M. Berrada, “Peeking inside the black-box: a survey on explainable artificial intelligence (xai),” IEEE access, vol. 6, pp. 52 138– 52 160, 2018

  4. [11]

    How can we develop explainable systems? insights from a literature review and an interview study,

    L. Chazette, J. Kl ¨under, M. Balci, and K. Schneider, “How can we develop explainable systems? insights from a literature review and an interview study,” in Proceedings of the International Conference on Software and System Processes and International Conference on Global Sof...

  5. [12]

    What can be concluded from user feedback? - an empirical study,

    M. Anders, M. Obaidi, A. Specht, and B. Paech, “What can be concluded from user feedback? - an empirical study,” in 2023 IEEE 31st International Requirements Engineering Conference Workshops (REW) , 2023, pp. 122–128

  6. [13]

    A study on the men- tal models of users concerning existing software,

    M. Anders, M. Obaidi, B. Paech, and K. Schneider, “A study on the men- tal models of users concerning existing software,” in Requirements Engi- neering: Foundation for Software Quality, V . Gervasi and A. V ogelsang, Eds. Cham: Springer International Publishing, 2022, pp. 235–250

  7. [14]

    Explanation needs in app reviews: Taxonomy and automated detection,

    M. Unterbusch, M. Sadeghi, J. Fischbach, M. Obaidi, and A. V ogelsang, “Explanation needs in app reviews: Taxonomy and automated detection,” in REW. IEEE, 2023

  8. [15]

    Requirements elicitation: A survey of tech- niques, approaches, and tools,

    D. Zowghi and C. Coulin, “Requirements elicitation: A survey of tech- niques, approaches, and tools,” in Engineering and managing software requirements. Springer, 2005, pp. 19–46

  9. [16]

    A com- parison of security requirements engineering methods,

    B. Fabian, S. G ¨urses, M. Heisel, T. Santen, and H. Schmidt, “A com- parison of security requirements engineering methods,” Requirements engineering, vol. 15, pp. 7–40, 2010

  10. [17]

    Designing end-user personas for explainability requirements using mixed methods research,

    J. Droste, H. Deters, J. Puglisi, and J. Kl ¨under, “Designing end-user personas for explainability requirements using mixed methods research,” in REW. IEEE, 2023

  11. [18]

    Modeling and evaluating per- sonas with software explainability requirements,

    H. Ramos, M. Fonseca, and L. Ponciano, “Modeling and evaluating per- sonas with software explainability requirements,” in Human-Computer Interaction: 7th Iberoamerican Workshop, HCI-COLLAB 2021, Sao Paulo, Brazil, September 8–10, 2021, Proceedings 7 , 2021, pp. 136– 149

  12. [19]

    A stressful explanation: The dual effect of explainable artificial intelligence in personal health manage- ment,

    M. Gr ¨uning, T. Wolf, and M. Trenz, “A stressful explanation: The dual effect of explainable artificial intelligence in personal health manage- ment,” in Proceedings of the 57th Hawaii International Conference on System Sciences, 2024

  13. [20]

    A systematic review and taxonomy of expla- nations in decision support and recommender systems,

    I. Nunes and D. Jannach, “A systematic review and taxonomy of expla- nations in decision support and recommender systems,” User Modeling and User-Adapted Interaction , vol. 27, 2017

  14. [21]

    A review of taxonomies of explainable artificial intelligence (xai) methods,

    T. Speith, “A review of taxonomies of explainable artificial intelligence (xai) methods,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, ser. FAccT ’22. New York, NY , USA: Association for Computing Machinery, 2022, p. 2239–2250. [Onli...

  15. [22]

    How does users’ app knowledge influence the preferred level of detail and format of software explanations?

    M. Obaidi, J. Fischbach, M. Herrmann, H. Deters, J. Droste, J. Kl ¨under, and K. Schneider, “How does users’ app knowledge influence the preferred level of detail and format of software explanations?” in REFSQ’25, 2025

  16. [23]

    Do users’ explainability needs in software change with mood?

    M. Obaidi, J. Droste, H. Deters, M. Herrmann, J. Kl ¨under, and K. Schneider, “Do users’ explainability needs in software change with mood?” in REFSQ’25, 2025

  17. [24]

    How explainable is your system? towards a quality model for explainability,

    H. Deters, J. Droste, M. Obaidi, and K. Schneider, “How explainable is your system? towards a quality model for explainability,” in Require- ments Engineering: Foundation for Software Quality . Cham: Springer Nature Switzerland, 2024, pp. 3–19

  18. [25]

    Automatic generation of explainability requirements and software explanations from user reviews,

    M. Obaidi, J. Fischbach, J. Droste, H. Deters, M. Herrmann, J. Kl ¨under, S. Kr ¨atzig, H. Villamizar, and K. Schneider, “Automatic generation of explainability requirements and software explanations from user reviews,” in 2025 IEEE 33rd International Requirements Engineering ...

  19. [26]

    From app features to explanation needs: Analyzing correlations and predictive potential,

    M. Obaidi, K. Qengaj, J. Droste, H. Deters, M. Herrmann, E. Schmid, J. Kl ¨under, and K. Schneider, “From app features to explanation needs: Analyzing correlations and predictive potential,” in 2025 IEEE 33rd International Requirements Engineering Conference Workshops (REW) , 2025

  20. [27]

    Exploring the means to measure explainability: Metrics, heuristics and questionnaires,

    H. Deters, J. Droste, M. Obaidi, and K. Schneider, “Exploring the means to measure explainability: Metrics, heuristics and questionnaires,” Information and Software Technology , vol. 181, p. 107682, 2025

  21. [28]

    Iden- tifying explanation needs: Towards a catalog of user-based indicators,

    H. Deters, L. Reinhardt, J. Droste, M. Obaidi, and K. Schneider, “Iden- tifying explanation needs: Towards a catalog of user-based indicators,” in 2025 IEEE 33rd International Requirements Engineering Conference (RE), Valencia, Spain, Sep. 2025

  22. [29]

    Chapter 81 experimental evidence on the existence of hypothetical bias in value elicitation methods,

    G. W. Harrison and E. E. Rutstr ¨om, “Chapter 81 experimental evidence on the existence of hypothetical bias in value elicitation methods,” in Handbook of Experimental Economics Results , C. R. Plott and V . L. Smith, Eds. Amsterdam: Elsevier, 2008, vol. 1, pp. 752–767

  23. [30]

    On the pulse of requirements elicitation: Physiological triggers and explainability needs,

    H. Deters, J. Droste, and K. Schneider, “On the pulse of requirements elicitation: Physiological triggers and explainability needs,” in REFSQ Workshops. CEUR Workshop Proceedings, 2024

  24. [31]

    Requirements on explanations: A quality framework for explainability,

    L. Chazette, V . Kl ¨os, F. Herzog, and K. Schneider, “Requirements on explanations: A quality framework for explainability,” in RE, 2022

  25. [32]

    Requirements elicitation tech- niques: a systematic literature review based on the maturity of the techniques,

    C. Pacheco, I. Garc ´ıa, and M. Reyes, “Requirements elicitation tech- niques: a systematic literature review based on the maturity of the techniques,” IET Software, vol. 12, no. 4, pp. 365–378, 2018

  26. [33]

    Successful requirement elicita- tion by combining requirement engineering techniques,

    D. Mishra, A. Mishra, and A. Yazici, “Successful requirement elicita- tion by combining requirement engineering techniques,” in 2008 First International Conference on the Applications of Digital Information and Web Technologies (ICADIWT). IEEE, 2008, pp. 258–263

  27. [34]

    Non-functional re- quirements elicitation guideline for agile methods,

    M. Younas, D. Jawawi, I. Ghani, and R. Kazmi, “Non-functional re- quirements elicitation guideline for agile methods,” Journal of Telecom- munication, Electronic and Computer Engineering (JTEC) , vol. 9, no. 3-4, pp. 137–142, 2017

  28. [35]

    A model for evaluating requirements elicitation techniques in software development projects

    N. C. Alflen, E. P. Prado, and A. Grotta, “A model for evaluating requirements elicitation techniques in software development projects.” in ICEIS (2), 2020, pp. 242–249

  29. [36]

    The role of domain knowledge in requirements elicitation via interviews: an exploratory study,

    I. Hadar, P. Soffer, and K. Kenzi, “The role of domain knowledge in requirements elicitation via interviews: an exploratory study,” Require- ments Engineering, vol. 19, pp. 143–159, 2014

  30. [37]

    A practical guide to requirements elicitation techniques selection-an empirical study,

    F. Anwar and R. Razali, “A practical guide to requirements elicitation techniques selection-an empirical study,” Middle-East Journal of Scien- tific Research, vol. 11, no. 8, pp. 1059–1067, 2012

  31. [38]

    Using combined techniques for requirements elicitation: A brazilian case study

    N. C. Alflen, L. C ´assia, E. P. Prado, and A. Grotta, “Using combined techniques for requirements elicitation: A brazilian case study.” in ICEIS (2), 2021, pp. 241–248

  32. [39]

    Combining requirements engi- neering techniques-theory and case study,

    L. Jiang, A. Eberlein, and B. H. Far, “Combining requirements engi- neering techniques-theory and case study,” in 12th IEEE International Conference and Workshops on the Engineering of Computer-Based Systems (ECBS’05). IEEE, 2005, pp. 105–112

  33. [40]

    An iterative approach for global requirements elicitation: A case study analysis,

    N. Sabahat, F. Iqbal, F. Azam, and M. Y . Javed, “An iterative approach for global requirements elicitation: A case study analysis,” in 2010 International Conference on Electronics and Information Engineering , vol. 1. IEEE, 2010, pp. V1–361

  34. [41]

    A methodology for the selection of requirement elicitation techniques,

    S. Tiwari and S. S. Rathore, “A methodology for the selection of requirement elicitation techniques,” arXiv preprint arXiv:1709.08481 , 2017

  35. [42]

    R. A. Krueger, Focus groups: A practical guide for applied research . Sage publications, 2014

  36. [43]

    Customer-driven software product devel- opment software products for the social media world–a case study,

    T. Fehlmann and E. Kranich, “Customer-driven software product devel- opment software products for the social media world–a case study,” in Systems, Software and Services Process Improvement: 20th European Conference, EuroSPI 2013, Dundalk, Ireland, June 25-27, 2013. Pro- ceedi...

  37. [44]

    Cognitive profiles in understanding and prioritizing requirements: a case study,

    N. M. Carod and A. Cechich, “Cognitive profiles in understanding and prioritizing requirements: a case study,” in 2010 Fifth International Conference on Software Engineering Advances . IEEE, 2010, pp. 341– 346

  38. [45]

    Using flow to improve com- munication of requirements in globally distributed software projects,

    K. Stapel, E. Knauss, and K. Schneider, “Using flow to improve com- munication of requirements in globally distributed software projects,” in 2009 Collaboration and Intercultural Issues on Requirements: Commu- nication, Understanding and Softskills , 2009, pp. 5–14

  39. [46]

    Wohlin, P

    C. Wohlin, P. Runeson, M. H ¨ost, M. C. Ohlsson, B. Regnell, and A. Wessl´en, Experimentation in software engineering . Springer, 2012

  40. [47]

    Framing what can be explained – an operational taxonomy for explainability needs,

    J. Droste, H. Deters, M. Obaidi, J. Kl ¨under, and K. Schneider, “Framing what can be explained – an operational taxonomy for explainability needs,” Requirements Engineering , 2025. [Online]. Available: https://doi.org/10.1007/s00766-025-00440-x

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.