Pith. sign in

REVIEW 3 major objections 61 references

ISTQB certifications help careers and shared language, but practitioners and experts still dispute their practical testing value.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 21:06 UTC pith:USA5RM7T

load-bearing objection First careful multivocal map of ISTQB endorsements vs criticisms, with transparent AI-MLR logs and expert triangulation; useful for the testing profession even if the 20-source pool cannot fully support global claims. the 3 major comments →

arxiv 2603.14572 v2 pith:USA5RM7T submitted 2026-03-15 cs.SE

ISTQB Certifications Under the Lens: Their Contributions to the Software-Testing Profession; and AI-assisted Synthesis of Practitioners' Endorsements and Criticisms

classification cs.SE
keywords ISTQBsoftware testing certificationmultivocal literature reviewpractitioner perceptionsGenAI in systematic reviewssoftware testing body of knowledgeendorsements and criticisms
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper examines the global ISTQB certification scheme, which has issued more than 1.2 million credentials across 130-plus countries, by synthesizing what practitioners actually say about it. Using an AI-assisted multivocal literature review of twenty grey-literature sources (blogs, forums, surveys) under continuous human oversight, the authors extract recurring endorsements and criticisms, then ask four independent experts to rate how precise the endorsements are and how fair the criticisms are. The synthesis shows that practitioners consistently credit the certificates with career mobility, professional recognition, and a common vocabulary, while criticizing them as overly theoretical, memorization-heavy, slow to track agile and automation practice, and weak as measures of real skill. Expert ratings largely confirm the precision of the career and terminology benefits and treat many skill-related criticisms as context-dependent rather than universally true or false. The paper therefore frames ISTQB as a real but incomplete contribution to the software-testing body of knowledge: useful for signaling and shared language, contested for hands-on competence.

Core claim

ISTQB certifications deliver recognizable career and communication value yet remain contested on practical utility; triangulating practitioner voices from an AI-assisted multivocal literature review of twenty grey-literature sources with four independent experts yields an evidence-based reflection on their strengths and weaknesses in shaping the software testing body of knowledge.

What carries the argument

AI-assisted Multivocal Literature Review (MLR) under continuous human oversight: ChatGPT deep-research agent searches and thematically codes grey literature, researchers apply inclusion/exclusion criteria and quality-assurance checks against known AI error types, then a four-expert panel rates endorsement precision and criticism fairness on five-point Likert scales.

Load-bearing premise

The final set of twenty grey-literature sources plus four experts is treated as a sufficiently representative sample of global practitioner sentiment for thematic patterns and average ratings to be generalized.

What would settle it

A large-scale, stratified survey or interview study of certified and non-certified testers across multiple regions and industries that measures actual career outcomes, on-the-job skill change, and employer hiring filters would show whether the reported endorsement and criticism themes hold beyond the sampled online voices.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The paper offers a pragmatic review of the ISTQB certification portfolio (Section 3) and an AI-assisted Multivocal Literature Review of 20 grey-literature sources that synthesizes five endorsement themes (RQ1) and ten criticism themes (RQ2), with a cross-perspective synthesis (RQ3). Four independent experts then rate endorsement precision and criticism fairness on Likert scales (Tables 3–4, RQ4), attributing residual tensions to schools of thought and context. The central claim is that ISTQB supplies recognizable career and communication value while remaining contested on practical utility, and that the AI-assisted MLR plus expert triangulation yields an evidence-based reflection on its role in the software-testing body of knowledge.

Significance. ISTQB is the dominant global testing credential (1.2M+ certificates). A transparent, multi-source synthesis of practitioner endorsements and criticisms, triangulated by independent experts and accompanied by an open empirical dataset and detailed AI-oversight log (Appendix, Table 2), is useful for testers, employers, educators, and the certification body itself. The methodological documentation of human-supervised ChatGPT deep-research for grey-literature MLR is a secondary contribution that other SE evidence-synthesis studies can reuse. Strengths include explicit validity discussion (7.4), multi-author checks, and public artifacts.

major comments (3)
  1. Section 5 and Appendix (Tables 6–8): the entire thematic synthesis of RQ1–RQ3 rests on a final pool of only 20 grey-literature items (blogs, Reddit/MoT threads, a few surveys). No saturation metric, no inter-coder reliability, and no quantitative check for over-representation of English-language or forum-centric voices are reported. Section 7.4 acknowledges selection bias but still frames the result as an ‘evidence-based’ global reflection. Either enlarge/justify the pool or substantially soften the generalizability language that underpins the abstract and conclusions.
  2. Section 6 / Tables 3–4: expert precision and fairness are summarized by arithmetic means of four ordinal Likert ratings. The paper itself notes the ordinal-averaging debate (citing [47]) yet still treats the averages as the primary quantitative baseline for interpreting root causes. Report full rating distributions or medians/modes, and clarify that the numeric averages are only pragmatic summaries, not interval-scale evidence.
  3. Section 4.2 and Appendix: the claim that continuous human oversight plus the QA strategies in Table 2 render ChatGPT deep-research extractions reliable is asserted but not quantified. No inter-rater agreement between AI themes and human re-coding, nor any count of hallucinations caught and corrected, is supplied. Without such metrics the methodological contribution remains under-supported relative to the weight placed on the AI-assisted pipeline.

Circularity Check

0 steps flagged

No circularity: empirical synthesis of external grey literature and independent expert ratings; no derivation reduces to fitted parameters or self-definition.

full rationale

This is an empirical multivocal literature review (MLR) plus expert panel study, not a first-principles derivation or predictive model. RQ1–RQ2 themes are extracted from 20 external grey-literature sources (blogs, forums, surveys, Reddit/Ministry-of-Testing threads) listed in Tables 6–8 and the Appendix; RQ3 is a cross-perspective synthesis of those same external statements; RQ4 is Likert ratings by four independent experts on the precision/fairness of those statements (Tables 3–4). None of the central claims (career/communication value, contested practical utility, schools-of-thought root causes) is obtained by fitting a parameter to a subset of the data and then “predicting” a related quantity, nor by defining a construct in terms of the result it is said to produce. Author prior experience with ISTQB is disclosed (Section 2.7) and mitigated by multi-author checks, expert triangulation, and open artifacts; self-citations to the authors’ earlier education/MLR papers supply background and method, not uniqueness theorems or load-bearing uniqueness claims that force the present conclusions. The paper’s own validity section (7.4) and Appendix correctly flag selection bias and AI risks without converting those risks into circular reasoning. Score 0 is therefore the correct, proportionate finding.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

As an empirical qualitative synthesis the paper rests on standard MLR methodology, the assumption that grey-literature practitioner voices are the primary evidence base for certification perceptions, and the claim that continuous human oversight plus an explicit AI-error QA table sufficiently controls LLM hallucinations and bias. No free parameters are fitted; no new physical or mathematical entities are postulated.

axioms (4)
  • domain assumption Grey-literature sources (blogs, forums, surveys) constitute the primary and adequate evidence base for practitioner perceptions of ISTQB value and limitations.
    Stated in Section 4.2.1 and used to justify the MLR source pool; academic literature is treated as complementary only.
  • ad hoc to paper ChatGPT deep-research outputs under continuous human oversight and the tabulated QA strategies (Table 2) yield reliable thematic extractions.
    Core methodological premise of the AI-assisted MLR (Sections 4.2.5–4.2.7 and Appendix); without it the synthesis collapses.
  • domain assumption Four independent experts rating precision/fairness on a 5-point Likert scale provide external validation of the practitioner themes.
    RQ4 design (Section 4.3 and 6); panel composition and rating protocol are taken as sufficient triangulation.
  • standard math MLR guidelines of Garousi et al. (2019) correctly prescribe the search, selection, quality-assessment and synthesis steps.
    Explicitly trained into the AI and followed throughout (Section 4.2.4).

pith-pipeline@v1.1.0-grok45 · 46957 in / 2601 out tokens · 35061 ms · 2026-07-14T21:06:05.799651+00:00 · methodology

0 comments
read the original abstract

Context: The International Software Testing Qualifications Board (ISTQB) certification dominates global software testing practice, with 1.2+ million certifications issued across 130+ countries. Yet it remains contested: practitioners value it for career advancement and shared terminology, while others criticize it as overly theoretical and somewhat disconnected from real-world testing practices. Objective: This study investigates the perceived value and critique of ISTQB certifications, the most widely recognized testing qualifications worldwide. Method: We conducted an AI-assisted Multivocal Literature Review (MLR), combining academic and grey literature to synthesize practitioner endorsements (RQ1) and criticisms (RQ2). ChatGPT's deep research capability was employed under continuous human oversight, with QA strategies ensuring transparency and reliability. As another analysis, we asked a panel of four independent experts to evaluate the precision of endorsements and fairness of criticisms. Results: Practitioner endorsements emphasized career benefits, improved communication, and a shared vocabulary as the main values of ISTQB certifications. Criticisms focused on excessive theoretical content, limited relevance in agile and automation-intensive contexts, and weak support for real testing skills. Expert review confirmed that while many endorsements were precise, several criticisms reflected broader tensions in the discipline, including contrasting schools of thought in testing practice. Conclusions: ISTQB certifications provide recognizable career and communication value but remain contested in terms of practical utility. By triangulating practitioner voices with expert validation, this study delivers an evidence-based reflection on the strengths and weaknesses of ISTQB in shaping the software testing body of knowledge.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 1 canonical work pages

  1. [1]

    Current State of the Software Testing Education in North American Academia and Some Recommendati ons for the New Educators

    V. Garousi and A. Mathur, "Current State of the Software Testing Education in North American Academia and Some Recommendati ons for the New Educators " in Proceedings of IEEE Conference on Software Engineering Education and Training, 2010, pp. 89–96

  2. [2]

    S oftware-testing education: A systematic literature mapping,

    V. Garousi, A. Rainer, P. Lauvås Jr, and A. Arcuri, "S oftware-testing education: A systematic literature mapping," Journal of Systems and Software, vol. 165, p. 110570, 2020

  3. [3]

    A pragmatic look at education and training of software test engineers: Further cooperation of academia and industry is needed,

    V. Garousi and A. B. Kele ş, "A pragmatic look at education and training of software test engineers: Further cooperation of academia and industry is needed," in IEEE International Conference on Software Testing, Verification and Validation, 2024: IEEE, pp. 354–360

  4. [4]

    Graham, E

    D. Graham, E. v. Veenendaal, I. Evans, and R. Black, Foundations of software testing: ISTQB certification. Cengage Learning, 2006. 40

  5. [5]

    ISTQB Certification: Why You Need It and How to Get It,

    R. Black, "ISTQB Certification: Why You Need It and How to Get It," Testing Experience The Magazine for Professional Testers, vol. 1, pp. 26–30, 2008

  6. [6]

    Certification and licensing for software professionals and organizations,

    L. Werth, "Certification and licensing for software professionals and organizations," in Proceedings Conference on Software Engineering Education, 1998, pp. 151–160

  7. [7]

    Estimating the Global QA Workforce: How Many Testers Are There Worldwide?,

    R. Desyatnikov, "Estimating the Global QA Workforce: How Many Testers Are There Worldwide?," Sept. 2025. [Online]. Availabl e: www.linkedin.com/pulse/estimating-global-qa-workforce-how-many-testers-ruslan-desyatnikov- xmzie/www.statista.com/statistics/627312/worldwide-developer-population/

  8. [8]

    5-6 Million softwsre testers in the world? ,

    J. Arbon, "5-6 Million softwsre testers in the world? ," Sept. 2025. [Online]. Available: www.linkedin.com/posts/jasonarbon_softwsre-testers-activity- 7084731918356774912-E2wu/

  9. [9]

    An Open Modern Software Testing Laboratory Courseware: An Experience Report

    V. Garousi, "An Open Modern Software Testing Laboratory Courseware: An Experience Report " in Proceedings of the IEEE Conference on Software Engineering Education and Training, 2010, pp. 177–184

  10. [10]

    Cho osing the right testing tools and systems under test (SUTs) for practical exercises in testing education,

    V. Garousi, "Cho osing the right testing tools and systems under test (SUTs) for practical exercises in testing education," in 8th Workshop on Teaching Software Testing (WTST), IEEE Transactions on Education, 2009

  11. [11]

    Closing the gap be tween software engineering education and ind ustrial needs,

    V. Garousi, G. Giray, E. Tüzün, C. Catal, and M. Felderer, "Closing the gap be tween software engineering education and ind ustrial needs," IEEE Software, In press, 2019

  12. [12]

    A Bibliometric/Geographic Assessment of 40 Years of Software Engineering Research (1969-2009),

    V. Garousi and G. Ruhe, "A Bibliometric/Geographic Assessment of 40 Years of Software Engineering Research (1969-2009)," International Journal of Software Engineering and Knowledge Engineering, vol. 23, pp. 1343–1366, 2013

  13. [13]

    Citations, research topics and active countries in software engineering: A bibliometrics st udy

    V. Garousi and M. V. Mäntylä, "Citations, research topics and active countries in software engineering: A bibliometrics st udy " Elsevier Computer Science Review, vol. 19, pp. 56–77, 2016

  14. [14]

    Evolution of software testing strategies and trends: Semantic content analysis of software research corpus of the last 40 years,

    F. Gurcan, G. G. M. Dalveren, N. E. Cagiltay, D. Roman, and A. Soylu, "Evolution of software testing strategies and trends: Semantic content analysis of software research corpus of the last 40 years," IEEE Access, vol. 10, pp. 106093–106109, 2022

  15. [15]

    Data-set of Software Testing Books (1979-2019),

    V. Garousi, "Data-set of Software Testing Books (1979-2019)," Zenodo, Sept. 2025. [Online]. Available: www.doi.org/10.5281/zenodo.16811266

  16. [16]

    Cultivating So ftware Quality Engineers,

    T. Farley, "Cultivating So ftware Quality Engineers," in Pacific NW Software Quality Conference, 2014

  17. [17]

    Black-Box Testing for Practitioners: A Case of the New ISTQB Test Analyst Syllabus,

    M. Hamburg and A. Roman, "Black-Box Testing for Practitioners: A Case of the New ISTQB Test Analyst Syllabus," in IEEE Conference on Software Testing, Verification and Validation, 2025, pp. 634–645

  18. [18]

    A method to improve test process in federal enterprise architecture framework using the ISTQB framework,

    H. Mahdavifar, R. Nassiri, and A. Bagh eri, "A method to improve test process in federal enterprise architecture framework using the ISTQB framework," International Journal of Computer and Information Engineering, vol. 6, no. 10, pp. 1199–1203, 2012

  19. [19]

    From certifications to international standards in software testing: mapping from ISTQB to ISO/IEC/IEEE 29119-2,

    M.-L. Sánchez-Gordón and R. Colomo-Palacios, "From certifications to international standards in software testing: mapping from ISTQB to ISO/IEC/IEEE 29119-2," in European Conference on Software Process Improvement, 2018, pp. 43–55

  20. [20]

    ISTQB-based software testing education: Advantages and challenges,

    A. Szatmári, T. Gergely, and Á. Be szédes, "ISTQB-based software testing education: Advantages and challenges," in IEEE International Conference on Software Testing, Verification and Validation, 2023, pp. 389–396

  21. [21]

    Challenges and best practices in industry-academia collaborations in software engi neering: a systematic literature review,

    V. Garousi, K. Pete rsen, and B. Özkan, "Challenges and best practices in industry-academia collaborations in software engi neering: a systematic literature review," Information and Software Technology, vol. 79, pp. 106–127, 2016

  22. [22]

    Selecting the right topics for industry-academia collabora tions in software testing: an experience report,

    V. Garousi and K. Herkilo ğlu, "Selecting the right topics for industry-academia collabora tions in software testing: an experience report," in IEEE International Conference on Software Testing, Verification, and Validation, 2016, pp. 213–222

  23. [23]

    Industry-academia collaborations in software testing: experience and success stories from Canada and Turkey,

    V. Garousi, M. M. Eskandar, and K. Herkilo ğlu, "Industry-academia collaborations in software testing: experience and success stories from Canada and Turkey," Software Quality Journal, vol. 25, no. 4, pp. 1091–1143, 2017

  24. [24]

    What industry wants from academia in software testing? Hearing practitioners’ opinions,

    V. Garousi, M. Felderer, M. Kuhrmann, and K. Herkilo ğlu, "What industry wants from academia in software testing? Hearing practitioners’ opinions," in International Conference on Evaluation and Assessment in Software Engineering, Karlskrona, Sweden, 2017, pp. 65–69

  25. [25]

    I ndustry-academia collaborations in software engin eering: An empirical analysis of challenges, patterns and anti-patterns in research projects,

    V. Garousi, M. Felderer, J. M. Fernandes, D. Pfahl, and M. V. Mantyla, "I ndustry-academia collaborations in software engin eering: An empirical analysis of challenges, patterns and anti-patterns in research projects," in Proceedings of International Conference on Evaluation and Assessment in Software Engineering, Karlskrona, Sweden, 2017, pp. 224–229

  26. [26]

    Successful engagement of practitioners and so ftware engineering researchers: Evidence from 26 international industry-academia collaborative projects,

    V. Garousi, D. C. Shepherd, and K. Herkilo ğlu, "Successful engagement of practitioners and so ftware engineering researchers: Evidence from 26 international industry-academia collaborative projects," IEEE Software, In press, 2020

  27. [27]

    Characterizing industry-academia collaborations in software engineering: evidence from 101 projects,

    V. Garousi et al. , "Characterizing industry-academia collaborations in software engineering: evidence from 101 projects," Empirical Software Engineering, In press, 2019

  28. [28]

    Bloom’s taxonomy,

    M. Forehand, "Bloom’s taxonomy," Emerging perspectives on learning, teaching, and technology, vol. 41, no. 4, pp. 47–56, 2010

  29. [29]

    Benefitting from the grey literature in software engineering resea rch,

    V. Garousi, M. Felderer, M. V. Mäntylä, and A. Rainer, "Benefitting from the grey literature in software engineering resea rch," in Contemporary Empirical Methods in Software Engineering: Springer, 2020

  30. [30]

    Grey liter ature versus academic literature in softwa re engineering: A call for epistemological analysis,

    V. Garousi and A. Rainer, "Grey liter ature versus academic literature in softwa re engineering: A call for epistemological analysis," IEEE Software, vol. 38, no. 5, pp. 65–72, 2021. 41

  31. [31]

    Guidelines for including grey literature and conducting multivocal literature reviews in software engineering,

    V. Garousi, M. Felderer, and M. V. Mäntylä, "Guidelines for including grey literature and conducting multivocal literature reviews in software engineering," Information and Software Technology, vol. 106, pp. 101–121, 2019

  32. [32]

    Gui delines for Performing Systematic Litera ture Reviews in Software engineering,

    B. Kitchenham and S. Charters, "Gui delines for Performing Systematic Litera ture Reviews in Software engineering," Technical report, School of Computer Science, Keele University, EBSE-2007-01, 2007

  33. [33]

    Guidelines for conducting systematic mapping studies in software engineering : An update,

    K. Petersen, S. Vakkalanka, and L. Ku zniarz, "Guidelines for conducting systematic mapping studies in software engineering : An update," Information and Software Technology, vol. 64, pp. 1–18, 2015, doi: http://dx.doi.org/10.1016/j.infsof.2015.03.007

  34. [34]

    Experience-based guidelines for effective and efficient data extraction in systematic reviews in software engineering,

    V. Garousi and M. F elderer, "Experience-based guidelines for effective and efficient data extraction in systematic reviews in software engineering," in International Conference on Evaluation and Assessment in Software Engineering, Karlskrona, Sweden, 2017, pp. 170–179

  35. [35]

    How to optimize the systematic review process using AI tools,

    N. Fabiano et al., "How to optimize the systematic review process using AI tools," JCPP advances, vol. 4, no. 2, p. e12234, 2024

  36. [36]

    Leveraging artificial intelligence to enhance systematic reviews in health research: advanced tools and challenges,

    L. Ge et al., "Leveraging artificial intelligence to enhance systematic reviews in health research: advanced tools and challenges," Systematic reviews, vol. 13, no. 1, p. 269, 2024

  37. [37]

    The use of generative AI for scientific literature searches for systematic reviews: ChatGPT and Microsoft Bing AI performanc e evaluation,

    Y. N. Gwon et al., "The use of generative AI for scientific literature searches for systematic reviews: ChatGPT and Microsoft Bing AI performanc e evaluation," JMIR Medical Informatics, vol. 12, p. e51187, 2024

  38. [38]

    Using ChatGPT and other forms of generative AI in systematic reviews: Challenges and opportunities,

    M. M. Hossain, "Using ChatGPT and other forms of generative AI in systematic reviews: Challenges and opportunities," Journal of Medical Imaging and Radiation Sciences, vol. 55, no. 1, pp. 11–12, 2024

  39. [39]

    A I-assisted vs human-only evidence review,

    M. Egan, L. Leak-Smith, A. Hanna-Amodio, and M. Sirera, "A I-assisted vs human-only evidence review," 2025. [Online]. Avail able: www.gov.uk/government/publications/ai-assisted-vs-human-only-evidence-review/ai-assisted-vs-human-only-evidence-review-results-from-a- comparative-study

  40. [40]

    S. Lee, "The Era of Deep Research: How AI Conducts Research and Humans Validate It-A New Research ParadigmSubtitle: From A I as an Assistant to AI as a Researcher: The Shift in Academic Knowledge ProductionAuthors: Leehyo Jae, GPT-4o, Grok 3 (xAI) Date: March 2025Keyw ords: AI Research Automation, Deep Research, AI-Human Collaboration, Future of AcademiaS...

  41. [41]

    Deep research and analysis of ChatGPT based on multiple testing experiments,

    S. Wei, Y. Luo, S. Chen, T. Huang, and Y. Xiang, "Deep research and analysis of ChatGPT based on multiple testing experiments," in 2023 International Conference on Cyber-Enabled Distributed Computing and Knowledge Discovery (CyberC), 2023: IEEE, pp. 123–131

  42. [42]

    Gültekin et al

    O. Gültekin et al. , "Evaluating DeepResearch and DeepThink in anterior cruciate ligament surgery patient educat ion: ChatGPT-4o excels in comprehensiveness, DeepSeek R1 leads in clarity and readability of orthopaedic information," Knee Surgery, Sports Traumatology, Arthroscopy, 2025

  43. [43]

    X AI—Explainable artificial intelligence,

    D. Gunning, M. Stefik, J. Choi, T. Miller, S. Stumpf, and G.-Z. Yang, "X AI—Explainable artificial intelligence," Science robotics, vol. 4, no. 37, p. eaay7120, 2019

  44. [44]

    Prompt engineering with ChatGPT: a guide for academic writers,

    L. Giray, "Prompt engineering with ChatGPT: a guide for academic writers," Annals of biomedical engineering, vol. 51, no. 12, pp. 2629–2633, 2023

  45. [45]

    A prompt pattern catalog to enhance prompt engineering with chatgpt,

    J. White et al., "A prompt pattern catalog to enhance prompt engineering with chatgpt," arXiv preprint arXiv:2302.11382, 2023

  46. [46]

    Prompt engineering a prompt engineer,

    Q. Ye, M. Axmed, R. Pryzant, and F. Kh ani, "Prompt engineering a prompt engineer," arXiv preprint arXiv:2311.05661, 2023

  47. [47]

    How to analyze Likert and other rating scale data,

    S. E. Harpe, "How to analyze Likert and other rating scale data," Currents in pharmacy teaching and learning, vol. 7, no. 6, pp. 836–850, 2015

  48. [48]

    Software Testing Schools of ThoughtResearch log,

    M. Stevens, "Software Testing Schools of ThoughtResearch log," Sept. 2025. [Online]. Available: https://pmo.its.uconn.edu/2017/12/19/schools-of- thought/

  49. [49]

    Incorporating Real-World Industrial Testing Projects in Software Testing Courses: Opportunities, Challenges, and Lessons Learned,

    V. Garousi, "Incorporating Real-World Industrial Testing Projects in Software Testing Courses: Opportunities, Challenges, and Lessons Learned," in Proceedings of the IEEE Conference on Software Engineering Education and Training (CSEE&T), 2011, pp. 396–400

  50. [50]

    Bug hunt: Making early software testing lessons engaging and affordable,

    S. Elbaum, S. Person, J. Dokulil, and M. Jorde, "Bug hunt: Making early software testing lessons engaging and affordable," in International Conference on Software Engineering, 2007: IEEE, pp. 688–697

  51. [51]

    Guidelines for mana ging threats to validity of secondary stu dies in software engineering,

    A. Ampatzoglou, S. Bibi, P. Avgeriou , and A. Chatzigeorgiou, "Guidelines for mana ging threats to validity of secondary stu dies in software engineering," Contemporary empirical methods in software engineering, pp. 415–441, 2020

  52. [52]

    Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT,

    A. Adel and N. Alan i, "Can generative AI reliably synthesise literature? exploring hallucination issues in ChatGPT," AI & SOCIETY, pp. 1–14, 2025

  53. [53]

    Understanding and Av oiding Hallucinated References: An AI Writing Experiment,

    R. Cole, L. Maher, and R. Rice, "Understanding and Av oiding Hallucinated References: An AI Writing Experiment," https://wac.colostate.edu/repository/collections/continuing-experiments/august-2025/ai-literacy/understanding-avoiding-hallucinated-references/, 2025

  54. [54]

    Generative AI in writing research pa pers: A new type of algorithmic bias and uncertainty in scholarl y work,

    R. Jain and A. Jain, "Generative AI in writing research pa pers: A new type of algorithmic bias and uncertainty in scholarl y work," in Intelligent Systems Conference, 2024: Springer, pp. 656–669

  55. [55]

    Researc h log,

    University College London, "Researc h log," Sept. 2025. [Online]. Available: www.ucl.ac.uk/brain-sciences/research-log

  56. [56]

    Making supervision relation ships accountable: graduate student logs,

    A. Yeatman, "Making supervision relation ships accountable: graduate student logs," Australian Universities' Review, The, vol. 38, no. 2, pp. 9–11, 1995

  57. [57]

    Ph. D. study and the use of a reflective diary: A dialogue with self,

    J. Glaze, "Ph. D. study and the use of a reflective diary: A dialogue with self," 2002

  58. [58]

    The natural history of a doctoral research study: The role of a research diary and reflexivity,

    S. Li, "The natural history of a doctoral research study: The role of a research diary and reflexivity," in Emotions and reflexivity in health & social care field research: Springer, 2017, pp. 13–37. 42

  59. [59]

    Synthesizing evidence in software engineering research,

    D. S. Cruzes and T. Dybå , "Synthesizing evidence in software engineering research," in Proceedings of the ACM-IEEE International Symposium on Empirical Software Engineering and Measurement, 2010, pp. 1–10. 43 Appendix-Execution details (log) of the AI-assisted MLR with Human Oversight Having outlined the design of the AI-assisted MLR in Section 4, we now...

  60. [60]

    Study quality assessment

  61. [61]

    qualitative coding

    Data synthesis (and you should use "qualitative coding" for that) -Consider all the relevant Academic Literature and Grey Literature, about all shared Perspectives on Value, benefits, and Limitations of the ISTQB Certifications -Include both English and non-English sources, as they may be interesting insights in a source which is written in a non-English ...