A systematic review of 50 studies identifies 69 LLM-assisted tasks in empirical software engineering, concentrated in data processing and analysis with gaps in human-centered integration and reproducibility reporting.
Cruzes and Tore Dyba
5 Pith papers cite this work, alongside 869 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.SE 5representative citing papers
Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
Systematic review of 97 studies on breaking changes in five software ecosystems, producing a four-dimensional taxonomy, reason/impact categories, 43 detection approaches, and 66 mitigation strategies.
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
Generative AI suitability in qualitative research depends primarily on the approach (small-q positivist/post-positivist or Big Q non-positivist) along with skills, ethics, and personal preferences.
citing papers explorer
-
LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda
A systematic review of 50 studies identifies 69 LLM-assisted tasks in empirical software engineering, concentrated in data processing and analysis with gaps in human-centered integration and reproducibility reporting.
-
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Gemini 3 Flash achieved the highest accuracy on PSM I-style questions among three tested LLMs, with low intra-model variability and systematic error patterns by question format and topic.
-
Breaking Changes in Software Ecosystems: A Systematic Literature Review
Systematic review of 97 studies on breaking changes in five software ecosystems, producing a four-dimensional taxonomy, reason/impact categories, 43 detection approaches, and 66 mitigation strategies.
-
Industry Practitioners Perspectives on AI Model Quality: Perceptions, Challenges, and Solutions
Industry AI practitioners view model quality through nine attributes with context-dependent priorities, where data imbalance is a key challenge addressed by strategies like active learning, as confirmed by interviews and a follow-up survey.
-
To Vibe Research or Not to Vibe Research? Generative AI in Qualitative Research
Generative AI suitability in qualitative research depends primarily on the approach (small-q positivist/post-positivist or Big Q non-positivist) along with skills, ethics, and personal preferences.