REVIEW 1 major objections 4 minor 38 references
A customized AI interviewer can run short self-administered interviews that completers rate positively, without claiming they match human interviews.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 02:01 UTC pith:3WYVWNOT
load-bearing objection Useful first data point on AI interviewers in ESE, honestly scoped, with limitations that are real but non-fatal. the 1 major comments →
AI-Conducted Interviews in Empirical Software Engineering: An Experience Report
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that a self-administered, voice-based AI interview workflow, built on a customized conversational agent following a predefined protocol, can be executed in real empirical software engineering studies and is acceptable to the participants who finish it. Evidence comes from 66 valid submissions: 92.4% contained the expected structured synthesis, 90.9% rated the overall experience positively, 90.9% felt comfortable, 95.5% found questions clear, and 89.4% would participate again. At the same time, five artifacts did not contain the expected synthesis, two conflicted with the reported protocol, and the most cited limitations were generic questions, limited sensitivity to answ
What carries the argument
The engine is a MyGPT, a customized conversational agent configured through a shared link, that follows a common interview protocol: it asks one question at a time, requests concrete examples, adapts to the participant's preferred natural language, avoids collecting identifying information, and closes by generating a structured synthesis based on the conversation. That synthesis, voluntarily pasted into a submission form, is the unit of analysis, together with a post-interview questionnaire. A post-hoc structural audit checks artifact format, language, length, and protocol consistency. Because the agent's output is a summary rather than a verbatim transcript, the paper consistently treats th
Load-bearing premise
The acceptability and feasibility conclusions rest on trusting the self-reported questionnaire answers of the 66 participants who finished, with no data from those who started but did not submit, no inter-rater replication of the artifact audit, and no independent validation of what those ratings mean.
What would settle it
A deployment that logs invitations, starts, and abandonments and finds that most invitees never finish the interview would undercut the operational-viability claim; likewise, a controlled comparison in which the same participants rate the AI interviewer significantly worse on a validated rapport or richness instrument would undercut the acceptability claim.
If this is right
- Researchers running short, focused, low-risk studies can collect completed interviews asynchronously, without scheduling a live session, and receive an immediately available structured artifact.
- Every submitted artifact requires individual inspection: in this deployment, five of 66 submissions lacked the expected synthesis and two conflicted with the reported protocol, making verification a required step rather than an option.
- High participant-rated clarity and pace can coexist with perceived lack of depth, so acceptance metrics alone do not imply probing quality or data richness.
- The workflow is positioned as complementary to human interviews, not a replacement; controlled studies comparing AI-only, human-only, and hybrid setups are the stated next step for assessing equivalence.
Where Pith is reading between the lines
- Inference: because the sample is completers-only, the positive acceptability figures may overstate how the workflow would fare among all invited participants; a log-based replication tracking invitations, starts, and abandonments would test this directly.
- Inference: the near-universal Portuguese-language artifact set means the claimed multilingual capability remains largely untested; a deliberate cross-language deployment with recorded interaction language would be the natural extension.
- Inference: the paper flags platform dependence as a threat, and a testable extension would re-run the same prompts on different conversational-agent backends to separate workflow design from vendor behavior.
- Inference: the 22.7% who cited privacy concerns point to a concrete design constraint: future protocols for sensitive or confidential topics need explicit platform-privacy notices and account-settings guidance before the interview starts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an experience of deploying a customized MyGPT as a self-administered, voice-based interviewer in two ESE studies: one on refactoring practices and one on the use of generative AI in Scrum-related activities. The workflow uses shared links, participant-selected natural language, a fixed prompt with follow-up instructions, and a final structured synthesis that participants voluntarily submit. The paper analyzes 66 completed submissions through a post-hoc structural audit of artifact format, language, length, and topic consistency, plus descriptive statistics from a post-interview questionnaire and open-ended comments. The central conclusion is explicitly narrowed to the 66 completers: the workflow is operationally viable and generally acceptable to these participants. The paper repeatedly disclaims any estimate of completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. The main RQ1 evidence is a single-coder audit classifying 92.4% (61/66) of artifacts as expected structured syntheses; RQ2 reports generally positive ratings (e.g., 90.9% overall positive, 95.5% question clarity, 89.4% would participate again) with a mixed comparative item and a substantive list of limitations; RQ3 draws lessons about participant instructions, voice-mode risks, privacy, artifact validation, and human oversight. The artifact audit is partly a prompt-adherence check, but the authors frame it as structural/topical consistency rather than as
Significance. If the results are taken at face value, this is a useful and unusually disciplined experience report. Its main value is not the novelty of AI interviewing per se, but the degree of scoping and transparency: exact denominators, explicit disclaimers about what was not measured, a replication package, and a clear separation between structural completeness of artifacts and qualitative fidelity. The inclusion of a mixed comparative item and multiple negative-option survey items provides some internal evidence against uniform social-desirability inflation. The single-coder audit and absence of non-completer data are acknowledged, and the conclusions do not overreach. The paper is a solid methodological contribution for a venue that publishes experience reports, and it gives concrete, reproducible lessons for researchers who might adopt similar workflows.
major comments (1)
- [Sections 2.9 and 3.1 (RQ1, artifact audit)] The headline operational metric is that 92.4% (61/66) of submitted artifacts followed the expected structured-synthesis format. This classification was made by one researcher, with ambiguous and protocol-inconsistent cases discussed by the team but no independent second coder. Because this number is central to RQ1, the single-coder nature of the audit is a load-bearing limitation. The manuscript acknowledges it in Section 2.9 and wisely limits what the number means (structural/topical consistency, not fidelity), so I do not view the limitation as fatal. However, the results section should either report a reliability check on a subset of artifacts or at least include the single-coder caveat directly in Section 3.1 when the 92.4% figure is first stated.
minor comments (4)
- [Section 5 and reference [33]] The text attributes the study cited as [33] to 'Cuevas et al.' (also mentioned in the Introduction), but the reference list entry [33] is by Villalba, Scurrell, Brown, Entenmann, and Daepp. Please harmonize the in-text attribution with the bibliography.
- [Figure 4] Panel (e) heading contains a typo ('Y ears of experience'); panels (c) and (e) use bare values such as '3' and '1.5' without percent signs. Make the formatting consistent with the other panels. Also, the row 'No weaknesses only 19.7%' in panel (b) is ambiguous for a multi-select item; rephrase as 'No weaknesses: 19.7%'.
- [Reference [9]] The DOI/URL for the replication package is broken across a line ('10.\n5281'); format it as a single working hyperlink.
- [Throughout] Non-standard symbols such as the lightbulb before the 'Finding' paragraphs may not render reliably in all venues; consider replacing them with standard 'Finding' or subsection formatting.
Circularity Check
No significant circularity: the central claim is bounded to completers and rests on independent questionnaire data; the prompt-defined format audit is a conformance check, not a derivation.
full rationale
The paper is an experience report whose central claim—operational viability and acceptability of the MyGPT workflow among participants who completed and submitted—is explicitly bounded to completers and is supported by independent questionnaire responses and submitted artifacts. No fitted parameter is presented as a prediction, and no claimed result is equivalent by construction to its inputs. The format-adherence audit (Section 2.9, Section 3.1) measures submitted artifacts against the synthesis structure that the MyGPT was prompted to produce; while this makes the audit a prompt-conformance check rather than an independent validation of interview quality, the paper explicitly frames it as a structural/topical consistency check and does not use it to claim data fidelity, richness, or equivalence to human interviews. The observed 92.4% adherence is an empirical outcome that could have been lower, not a quantity derived from the prompt. The paper repeatedly disclaims completion rates, researcher time savings, summary fidelity, and equivalence to human-conducted interviews (Abstract, Sections 2.1, 2.8, 2.9, 4), and it acknowledges limitations such as single-coder audit judgment, self-reported perceptions, and absence of full transcripts. Self-citations in the reference list (e.g., the replication package [9] and related-work items [20], [23], [29]) are not load-bearing for the central conclusion: they provide context or supporting evidence but are not used to force the paper's claims. Protocol–artifact mismatches are treated as data-quality caveats, not hidden assumptions. Accordingly, no circular step meeting the required evidentiary standard is present.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Self-reported questionnaire responses are a valid proxy for participant experience.
- ad hoc to paper The structural audit's classification of 'expected structured synthesis' is a meaningful measure of operational success.
- domain assumption Completers-only data are sufficient for the scoped claims.
- domain assumption LLM platform behavior (ChatGPT/MyGPT) is treated as a stable enough background for the reported workflow.
read the original abstract
Semi-structured interviews are widely used in empirical software engineering (ESE), but they are resource-intensive and difficult to coordinate across schedules, locations, and natural languages. This experience report examines a customized MyGPT used to conduct short, self-administered interviews in two ESE studies: one on refactoring practices and another on generative AI in Scrum-related activities. Participants accessed the interviewer through shared links, used voice interaction, selected a preferred natural language, and completed the interview without a researcher present. The AI followed a predefined protocol and generated a structured synthesis that participants voluntarily submitted; these artifacts were not treated as verbatim transcripts. We analyzed 66 submissions and questionnaire responses, and audited artifact format, language, length, and protocol consistency. Of the submitted artifacts, 92.4% followed the expected synthesis format, 65 were predominantly in Portuguese and one in English, and two conflicted with the reported protocol. Participants generally rated the experience positively: 90.9% reported a positive overall experience and comfort, 95.5% considered the questions clear, 97.0% rated the pace positively, and 89.4% would participate again. Reported limitations included generic questions, limited sensitivity to answers, insufficient depth, privacy concerns, and missed human interaction. The findings support the operational viability and acceptability of this workflow among analyzed respondents, but do not establish completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. AI interviewers should therefore be treated as a complementary option for short, focused, low-risk studies, with protocol design, privacy guidance, artifact validation, and human oversight.
Figures
Reference graph
Works this paper leans on
-
[1]
Mendon c a, and Julio C \' e sar Sampaio P
Camila Almeida, Isaque Copque, Alvaro Oliveira, Murilo Guerreiro Arouca, Adriano Barbosa, S \' a vio Freire, Manoel G. Mendon c a, and Julio C \' e sar Sampaio P. Leite. From elicitation interviews to software requirements: Evaluating LLM performance in requirement generation. In Workshop on Requirements Engineering , 2025
2025
-
[2]
Basili, Gianluigi Caldiera, and H
Victor R. Basili, Gianluigi Caldiera, and H. Dieter Rombach. The Goal Question Metric Approach . Wiley, 1994
1994
-
[3]
Brown et al
Tom B. Brown et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems , 2020
2020
-
[4]
Conducting qualitative interviews with AI
Felix Chopra and Ingar Haaland. Conducting qualitative interviews with AI . CESifo Working Paper 10666, CESifo, 2023
2023
-
[5]
Towards LLM -augmented multiagent systems for agile software engineering
Konrad Cinkusz and Jaroslaw A Chudziak. Towards LLM -augmented multiagent systems for agile software engineering. In Automated Software Engineering , pages 2476--2477, 2024
2024
-
[6]
Automated testing of refactoring engines
Brett Daniel, Danny Dig, Kely Garcia, and Darko Marinov. Automated testing of refactoring engines. In Foundations of Software Engineering , pages 185--194. ACM , 2007
2007
-
[7]
Refactoring: improving the design of existing code
Martin Fowler. Refactoring: improving the design of existing code . Addison-Wesley, 1999
1999
-
[8]
Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering , 31(1):13, 2024
Marco Aur \' e lio Gerosa, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering , 31(1):13, 2024
2024
-
[9]
AI-Conducted Interviews in Empirical Software Engineering: An Experience Report
Rohit Gheyi, Danyllo Albuquerque, Márcio Ribeiro, and Mirko Perkusich. AI-Conducted Interviews in Empirical Software Engineering: An Experience Report . https://doi.org/10.5281/zenodo.20277901, 2026
-
[10]
One thousand and one stories: a large-scale survey of software refactoring
Yaroslav Golubev, Zarina Kurbatova, Eman Abdullah AlOmar, Timofey Bryksin, and Mohamed Wiem Mkaouer. One thousand and one stories: a large-scale survey of software refactoring. In Foundations of Software Engineering , pages 1303--1313. ACM , 2021
2021
-
[11]
Experiences from conducting semi-structured interviews in empirical software engineering research
Siw Elisabeth Hove and Bente Anda. Experiences from conducting semi-structured interviews in empirical software engineering research. In International Symposium on Software Metrics , page 23, 2005
2005
-
[12]
Designing the conversational agent: Asking follow-up questions for information elicitation
Jiaxiong Hu, Jingya Guo, Ningjing Tang, Xiaojuan Ma, Yuan Yao, Changyuan Yang, and Yingqing Xu. Designing the conversational agent: Asking follow-up questions for information elicitation. Proceedings of the ACM on Human-Computer Interaction , 8(CSCW1):1--30, 2024
2024
-
[13]
Envisioning AI support during semi-structured interviews across the expertise spectrum
Zhe Liu, Jiamin Dai, Cristina Conati, and Joanna McGrenere. Envisioning AI support during semi-structured interviews across the expertise spectrum. Proceedings of the ACM on Human-Computer Interaction , 9(2):1--29, 2025
2025
-
[14]
Scalable requirements elicitation education through simulated interview practice with large language models
Nelson Lojo. Scalable requirements elicitation education through simulated interview practice with large language models. Master's thesis, University of California, Berkeley, 2025
2025
-
[15]
Real world Scrum a grounded theory of variations in practice
Zainab Masood, Rashina Hoda, and Kelly Blincoe. Real world Scrum a grounded theory of variations in practice. IEEE Transactions on Software Engineering , 48(5):1579--1591, 2022
2022
-
[16]
Equity In The Preparation Of Students For Software Engineering Coding Interviews: ChatGPT as a Mock Interviewer
Marlon Mejias, Zef Vargas, William Ted Edwards, Gloria Washington, Legand Burge, Dale-Marie Wilson, and Luce-Melissa Kouaho. Equity In The Preparation Of Students For Software Engineering Coding Interviews: ChatGPT as a Mock Interviewer . In Congress in Computer Science, Computer Engineering, and Applied Computing , pages 1016--1020, 2023
2023
-
[17]
How We Refactor, and How We Know It
Emerson Murphy-Hill, Chris Parnin, and Andrew Black. How We Refactor, and How We Know It . IEEE Transactions on Software Engineering , 38(1):5--18, 2012
2012
-
[18]
Impact of expert interviews in software engineering: Challenges and benefits
Wajeeha Nasar. Impact of expert interviews in software engineering: Challenges and benefits. In International MultiConference of Engineers and Computer Scientists , 2023
2023
-
[19]
Between policy and practice: GenAI adoption in agile software development teams
Michael Neumann, Lasse Bischof, Nic Elias Hinz, Abdullah Altun, Luca Stockmann, Dennis Schrader, Ana Carolina Ahaus, Erim Can Demirci, Benjamin Gabel, Maria Rauschenberger, Philipp Diebold, Henning Fritzemeier, and Adam Przyby ek. Between policy and practice: GenAI adoption in agile software development teams. In Agile Processes in Software Engineering an...
2026
-
[20]
Revisiting the refactoring mechanics
Jonhnanthan Oliveira, Rohit Gheyi, Melina Mongiovi, Gustavo Soares, M\' a rcio Ribeiro, and Alessandro Garcia. Revisiting the refactoring mechanics. Information and Software Technology , 110:136--138, 2019
2019
-
[21]
Refactoring: An Aid in Designing Application Frameworks and Evolving Object-Oriented Systems
William Opdyke and Ralph Johnson. Refactoring: An Aid in Designing Application Frameworks and Evolving Object-Oriented Systems . In Symposium on Object-Oriented Programming emphasizing Practical Applications , pages 274--282, 1990
1990
-
[22]
OpenAI . Creating and editing GPTs . https://help.openai.com/en/articles/8554397-creating-and-editing-gpts, 2026
arXiv 2026
-
[23]
Adoption of large language models in Scrum management: Insights from Brazilian practitioners
Mirko Perkusich, Danyllo Albuquerque, Allysson Allex Ara \'u jo, Matheus Paix \ a o, Rohit Gheyi, Marcos Kalinowski, and Angelo Perkusich. Adoption of large language models in Scrum management: Insights from Brazilian practitioners. In International Conference on Agile Software Development , pages 255--273, 2026
2026
-
[24]
An empirical investigation into the impact of refactoring on regression testing
Napol Rachatasumrit and Miryung Kim. An empirical investigation into the impact of refactoring on regression testing. In International Conference on Software Maintenance , pages 357--366, 2012
2012
-
[25]
Simulating the software engineering interview process using a decision-based serious computer game
Adrian Rusu, Robert Russell, and Remo Cocco. Simulating the software engineering interview process using a decision-based serious computer game. In International Conference on Computer Games , pages 235--239, 2011
2011
-
[26]
Agile Software Development with Scrum
Ken Schwaber and Mike Beedle. Agile Software Development with Scrum . Prentice Hall, 2001
2001
-
[27]
Carolyn B. Seaman. Qualitative methods in empirical studies of software engineering. Transactions on Software Engineering , 25(4):557--572, 1999
1999
-
[28]
Why we refactor? Confessions of GitHub contributors
Danilo Silva, Nikolaos Tsantalis, and Marco T \' u lio Valente. Why we refactor? Confessions of GitHub contributors . In Foundations of Software Engineering , pages 858--870, 2016
2016
-
[29]
Automated behavioral testing of refactoring engines
Gustavo Soares, Rohit Gheyi, and Tiago Massoni. Automated behavioral testing of refactoring engines. IEEE Transactions on Software Engineering , 39(2):147--162, 2013
2013
-
[30]
Ethical interviews in software engineering
Per Erik Strandberg. Ethical interviews in software engineering. In Empirical Software Engineering and Measurement , pages 1--11. IEEE , 2019
2019
-
[31]
Barriers to Refactoring
Ewan Tempero, Tony Gorschek, and Lefteris Angelis. Barriers to Refactoring . Communications of the ACM , 60(10):54--61, 2017
2017
-
[32]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need . Advances in Neural Information Processing Systems , pages 5998--6008, 2017
2017
-
[33]
Collecting qualitative data at scale with large language models: A case study
Alejandro Villalba, Jennifer Scurrell, Eva Brown, Jason Entenmann, and Madeleine Daepp. Collecting qualitative data at scale with large language models: A case study. Human Computer Interaction , 9(2):1--27, 2025
2025
-
[34]
Towards understanding refactoring engine bugs
Haibo Wang, Zhuolin Xu, Huaien Zhang, Nikolaos Tsantalis, and Shin Hwei Tan. Towards understanding refactoring engine bugs. Transactions on Software Engineering and Methodology , 35(5):1--55, 2026
2026
-
[35]
InterFlow : Designing unobtrusive AI to empower interviewers in semi-structured interviews
Yi Wen, Yu Zhang, Sriram Suresh, Zhicong Lu, Can Liu, and Meng Xia. InterFlow : Designing unobtrusive AI to empower interviewers in semi-structured interviews. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1--21. ACM, 2026
2026
-
[36]
Zhou, Wenxi Chen, Huahai Yang, and Changyan Chi
Ziang Xiao, Michelle X. Zhou, Wenxi Chen, Huahai Yang, and Changyan Chi. If I Hear You Correctly : Building and evaluating interview chatbots with active listening skills. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1--14. ACM, 2020
2020
-
[37]
Ziang Xiao, Michelle X. Zhou, Q. Vera Liao, Gloria Mark, Changyan Chi, Wenxi Chen, and Huahai Yang. Tell me about yourself: Using an AI -powered chatbot to conduct conversational surveys with open-ended questions. ACM Transactions on Computer-Human Interaction , 27(3):1--37, 2020
2020
-
[38]
He Zhang, Yueyan Liu, Xin Guan, Jie Cai, and John M. Carroll. Harnessing the power of AI in qualitative research: Role assignment, engagement, and user perceptions of ai-generated follow-up questions in semi-structured interviews, 2025. URL: https://arxiv.org/abs/2509.12709, https://arxiv.org/abs/2509.12709 arXiv:2509.12709
arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.