Pith. sign in

REVIEW 1 major objections 4 minor 38 references

A customized AI interviewer can run short self-administered interviews that completers rate positively, without claiming they match human interviews.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 02:01 UTC pith:3WYVWNOT

load-bearing objection Useful first data point on AI interviewers in ESE, honestly scoped, with limitations that are real but non-fatal. the 1 major comments →

arxiv 2607.14452 v1 pith:3WYVWNOT submitted 2026-07-16 cs.SE

AI-Conducted Interviews in Empirical Software Engineering: An Experience Report

classification cs.SE
keywords AI-conducted interviewslarge language modelssemi-structured interviewsempirical software engineeringqualitative researchparticipant perceptionsartifact auditself-administered interviews
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Interviews are central to empirical software engineering but costly to schedule across time zones and languages. This experience report asks whether a customized AI interviewer, accessed through a shared link and operated by voice, can conduct short semi-structured interviews without a researcher present. Across 66 completed submissions in two studies, the authors found the workflow operationally viable and generally well received: 92.4% of submitted artifacts followed the expected structured format, and roughly nine in ten respondents rated the experience, comfort, clarity, and pace positively. The authors are explicit about what they do not claim: no completion rates, no time savings, no summary fidelity, and no equivalence with human interviews. The paper's contribution is a concrete baseline for what a self-administered AI interview workflow can and cannot support in software engineering research.

Core claim

The central claim is that a self-administered, voice-based AI interview workflow, built on a customized conversational agent following a predefined protocol, can be executed in real empirical software engineering studies and is acceptable to the participants who finish it. Evidence comes from 66 valid submissions: 92.4% contained the expected structured synthesis, 90.9% rated the overall experience positively, 90.9% felt comfortable, 95.5% found questions clear, and 89.4% would participate again. At the same time, five artifacts did not contain the expected synthesis, two conflicted with the reported protocol, and the most cited limitations were generic questions, limited sensitivity to answ

What carries the argument

The engine is a MyGPT, a customized conversational agent configured through a shared link, that follows a common interview protocol: it asks one question at a time, requests concrete examples, adapts to the participant's preferred natural language, avoids collecting identifying information, and closes by generating a structured synthesis based on the conversation. That synthesis, voluntarily pasted into a submission form, is the unit of analysis, together with a post-interview questionnaire. A post-hoc structural audit checks artifact format, language, length, and protocol consistency. Because the agent's output is a summary rather than a verbatim transcript, the paper consistently treats th

Load-bearing premise

The acceptability and feasibility conclusions rest on trusting the self-reported questionnaire answers of the 66 participants who finished, with no data from those who started but did not submit, no inter-rater replication of the artifact audit, and no independent validation of what those ratings mean.

What would settle it

A deployment that logs invitations, starts, and abandonments and finds that most invitees never finish the interview would undercut the operational-viability claim; likewise, a controlled comparison in which the same participants rate the AI interviewer significantly worse on a validated rapport or richness instrument would undercut the acceptability claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Researchers running short, focused, low-risk studies can collect completed interviews asynchronously, without scheduling a live session, and receive an immediately available structured artifact.
  • Every submitted artifact requires individual inspection: in this deployment, five of 66 submissions lacked the expected synthesis and two conflicted with the reported protocol, making verification a required step rather than an option.
  • High participant-rated clarity and pace can coexist with perceived lack of depth, so acceptance metrics alone do not imply probing quality or data richness.
  • The workflow is positioned as complementary to human interviews, not a replacement; controlled studies comparing AI-only, human-only, and hybrid setups are the stated next step for assessing equivalence.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Inference: because the sample is completers-only, the positive acceptability figures may overstate how the workflow would fare among all invited participants; a log-based replication tracking invitations, starts, and abandonments would test this directly.
  • Inference: the near-universal Portuguese-language artifact set means the claimed multilingual capability remains largely untested; a deliberate cross-language deployment with recorded interaction language would be the natural extension.
  • Inference: the paper flags platform dependence as a threat, and a testable extension would re-run the same prompts on different conversational-agent backends to separate workflow design from vendor behavior.
  • Inference: the 22.7% who cited privacy concerns point to a concrete design constraint: future protocols for sensitive or confidential topics need explicit platform-privacy notices and account-settings guidance before the interview starts.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This paper reports an experience of deploying a customized MyGPT as a self-administered, voice-based interviewer in two ESE studies: one on refactoring practices and one on the use of generative AI in Scrum-related activities. The workflow uses shared links, participant-selected natural language, a fixed prompt with follow-up instructions, and a final structured synthesis that participants voluntarily submit. The paper analyzes 66 completed submissions through a post-hoc structural audit of artifact format, language, length, and topic consistency, plus descriptive statistics from a post-interview questionnaire and open-ended comments. The central conclusion is explicitly narrowed to the 66 completers: the workflow is operationally viable and generally acceptable to these participants. The paper repeatedly disclaims any estimate of completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. The main RQ1 evidence is a single-coder audit classifying 92.4% (61/66) of artifacts as expected structured syntheses; RQ2 reports generally positive ratings (e.g., 90.9% overall positive, 95.5% question clarity, 89.4% would participate again) with a mixed comparative item and a substantive list of limitations; RQ3 draws lessons about participant instructions, voice-mode risks, privacy, artifact validation, and human oversight. The artifact audit is partly a prompt-adherence check, but the authors frame it as structural/topical consistency rather than as

Significance. If the results are taken at face value, this is a useful and unusually disciplined experience report. Its main value is not the novelty of AI interviewing per se, but the degree of scoping and transparency: exact denominators, explicit disclaimers about what was not measured, a replication package, and a clear separation between structural completeness of artifacts and qualitative fidelity. The inclusion of a mixed comparative item and multiple negative-option survey items provides some internal evidence against uniform social-desirability inflation. The single-coder audit and absence of non-completer data are acknowledged, and the conclusions do not overreach. The paper is a solid methodological contribution for a venue that publishes experience reports, and it gives concrete, reproducible lessons for researchers who might adopt similar workflows.

major comments (1)
  1. [Sections 2.9 and 3.1 (RQ1, artifact audit)] The headline operational metric is that 92.4% (61/66) of submitted artifacts followed the expected structured-synthesis format. This classification was made by one researcher, with ambiguous and protocol-inconsistent cases discussed by the team but no independent second coder. Because this number is central to RQ1, the single-coder nature of the audit is a load-bearing limitation. The manuscript acknowledges it in Section 2.9 and wisely limits what the number means (structural/topical consistency, not fidelity), so I do not view the limitation as fatal. However, the results section should either report a reliability check on a subset of artifacts or at least include the single-coder caveat directly in Section 3.1 when the 92.4% figure is first stated.
minor comments (4)
  1. [Section 5 and reference [33]] The text attributes the study cited as [33] to 'Cuevas et al.' (also mentioned in the Introduction), but the reference list entry [33] is by Villalba, Scurrell, Brown, Entenmann, and Daepp. Please harmonize the in-text attribution with the bibliography.
  2. [Figure 4] Panel (e) heading contains a typo ('Y ears of experience'); panels (c) and (e) use bare values such as '3' and '1.5' without percent signs. Make the formatting consistent with the other panels. Also, the row 'No weaknesses only 19.7%' in panel (b) is ambiguous for a multi-select item; rephrase as 'No weaknesses: 19.7%'.
  3. [Reference [9]] The DOI/URL for the replication package is broken across a line ('10.\n5281'); format it as a single working hyperlink.
  4. [Throughout] Non-standard symbols such as the lightbulb before the 'Finding' paragraphs may not render reliably in all venues; consider replacing them with standard 'Finding' or subsection formatting.

Circularity Check

0 steps flagged

No significant circularity: the central claim is bounded to completers and rests on independent questionnaire data; the prompt-defined format audit is a conformance check, not a derivation.

full rationale

The paper is an experience report whose central claim—operational viability and acceptability of the MyGPT workflow among participants who completed and submitted—is explicitly bounded to completers and is supported by independent questionnaire responses and submitted artifacts. No fitted parameter is presented as a prediction, and no claimed result is equivalent by construction to its inputs. The format-adherence audit (Section 2.9, Section 3.1) measures submitted artifacts against the synthesis structure that the MyGPT was prompted to produce; while this makes the audit a prompt-conformance check rather than an independent validation of interview quality, the paper explicitly frames it as a structural/topical consistency check and does not use it to claim data fidelity, richness, or equivalence to human interviews. The observed 92.4% adherence is an empirical outcome that could have been lower, not a quantity derived from the prompt. The paper repeatedly disclaims completion rates, researcher time savings, summary fidelity, and equivalence to human-conducted interviews (Abstract, Sections 2.1, 2.8, 2.9, 4), and it acknowledges limitations such as single-coder audit judgment, self-reported perceptions, and absence of full transcripts. Self-citations in the reference list (e.g., the replication package [9] and related-work items [20], [23], [29]) are not load-bearing for the central conclusion: they provide context or supporting evidence but are not used to force the paper's claims. Protocol–artifact mismatches are treated as data-quality caveats, not hidden assumptions. Accordingly, no circular step meeting the required evidentiary standard is present.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No fitted parameters or invented entities. The central assumptions are empirical/domain-level: self-reports stand in for experience, the artifact-format audit is a valid operationalization despite being prompt-defined, completers stand in for the population of interest only within the paper's scoped claims, and the MyGPT platform behavior is treated as a reproducible enough black box.

axioms (4)
  • domain assumption Self-reported questionnaire responses are a valid proxy for participant experience.
    RQ2 acceptability claims rely on Likert-scale self-reports without external validation; novelty or social desirability could bias results. Entered in Sections 2.7 and 2.9.
  • ad hoc to paper The structural audit's classification of 'expected structured synthesis' is a meaningful measure of operational success.
    The categories are defined by the paper's own prompt structure, introducing partial circularity; the audit itself was not independently re-coded. Sections 2.6 and 2.9.
  • domain assumption Completers-only data are sufficient for the scoped claims.
    All analyses are conditional on submissions; no invitation or start logs exist, so unobserved non-completers could have very different experiences. Sections 2.3 and 3.1.
  • domain assumption LLM platform behavior (ChatGPT/MyGPT) is treated as a stable enough background for the reported workflow.
    Model version, account tier, and voice behavior were not controlled; replication relies on prompt text rather than platform behavior. Sections 2.5 and 4.

pith-pipeline@v1.3.0-alltime-deepseek · 18120 in / 10461 out tokens · 101000 ms · 2026-08-02T02:01:18.446309+00:00 · methodology

0 comments
read the original abstract

Semi-structured interviews are widely used in empirical software engineering (ESE), but they are resource-intensive and difficult to coordinate across schedules, locations, and natural languages. This experience report examines a customized MyGPT used to conduct short, self-administered interviews in two ESE studies: one on refactoring practices and another on generative AI in Scrum-related activities. Participants accessed the interviewer through shared links, used voice interaction, selected a preferred natural language, and completed the interview without a researcher present. The AI followed a predefined protocol and generated a structured synthesis that participants voluntarily submitted; these artifacts were not treated as verbatim transcripts. We analyzed 66 submissions and questionnaire responses, and audited artifact format, language, length, and protocol consistency. Of the submitted artifacts, 92.4% followed the expected synthesis format, 65 were predominantly in Portuguese and one in English, and two conflicted with the reported protocol. Participants generally rated the experience positively: 90.9% reported a positive overall experience and comfort, 95.5% considered the questions clear, 97.0% rated the pace positively, and 89.4% would participate again. Reported limitations included generic questions, limited sensitivity to answers, insufficient depth, privacy concerns, and missed human interaction. The findings support the operational viability and acceptability of this workflow among analyzed respondents, but do not establish completion rates, time savings, summary fidelity, or equivalence to human-conducted interviews. AI interviewers should therefore be treated as a complementary option for short, focused, low-risk studies, with protocol design, privacy guidance, artifact validation, and human oversight.

Figures

Figures reproduced from arXiv: 2607.14452 by Danyllo Albuquerque, M\'arcio Ribeiro, Mirko Perkusich, Rohit Gheyi.

Figure 1
Figure 1. Figure 1: Overview of the study design. Participants completed a self-administered voice-based interview with an AI interviewer, generated an interview artifact from the conversation, submitted the artifact and questionnaire responses through a post-interview form, and provided data for descriptive and qualitative analysis [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

38 extracted references · 1 canonical work pages

  1. [1]

    Mendon c a, and Julio C \' e sar Sampaio P

    Camila Almeida, Isaque Copque, Alvaro Oliveira, Murilo Guerreiro Arouca, Adriano Barbosa, S \' a vio Freire, Manoel G. Mendon c a, and Julio C \' e sar Sampaio P. Leite. From elicitation interviews to software requirements: Evaluating LLM performance in requirement generation. In Workshop on Requirements Engineering , 2025

  2. [2]

    Basili, Gianluigi Caldiera, and H

    Victor R. Basili, Gianluigi Caldiera, and H. Dieter Rombach. The Goal Question Metric Approach . Wiley, 1994

  3. [3]

    Brown et al

    Tom B. Brown et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems , 2020

  4. [4]

    Conducting qualitative interviews with AI

    Felix Chopra and Ingar Haaland. Conducting qualitative interviews with AI . CESifo Working Paper 10666, CESifo, 2023

  5. [5]

    Towards LLM -augmented multiagent systems for agile software engineering

    Konrad Cinkusz and Jaroslaw A Chudziak. Towards LLM -augmented multiagent systems for agile software engineering. In Automated Software Engineering , pages 2476--2477, 2024

  6. [6]

    Automated testing of refactoring engines

    Brett Daniel, Danny Dig, Kely Garcia, and Darko Marinov. Automated testing of refactoring engines. In Foundations of Software Engineering , pages 185--194. ACM , 2007

  7. [7]

    Refactoring: improving the design of existing code

    Martin Fowler. Refactoring: improving the design of existing code . Addison-Wesley, 1999

  8. [8]

    Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering , 31(1):13, 2024

    Marco Aur \' e lio Gerosa, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering , 31(1):13, 2024

  9. [9]

    AI-Conducted Interviews in Empirical Software Engineering: An Experience Report

    Rohit Gheyi, Danyllo Albuquerque, Márcio Ribeiro, and Mirko Perkusich. AI-Conducted Interviews in Empirical Software Engineering: An Experience Report . https://doi.org/10.5281/zenodo.20277901, 2026

  10. [10]

    One thousand and one stories: a large-scale survey of software refactoring

    Yaroslav Golubev, Zarina Kurbatova, Eman Abdullah AlOmar, Timofey Bryksin, and Mohamed Wiem Mkaouer. One thousand and one stories: a large-scale survey of software refactoring. In Foundations of Software Engineering , pages 1303--1313. ACM , 2021

  11. [11]

    Experiences from conducting semi-structured interviews in empirical software engineering research

    Siw Elisabeth Hove and Bente Anda. Experiences from conducting semi-structured interviews in empirical software engineering research. In International Symposium on Software Metrics , page 23, 2005

  12. [12]

    Designing the conversational agent: Asking follow-up questions for information elicitation

    Jiaxiong Hu, Jingya Guo, Ningjing Tang, Xiaojuan Ma, Yuan Yao, Changyuan Yang, and Yingqing Xu. Designing the conversational agent: Asking follow-up questions for information elicitation. Proceedings of the ACM on Human-Computer Interaction , 8(CSCW1):1--30, 2024

  13. [13]

    Envisioning AI support during semi-structured interviews across the expertise spectrum

    Zhe Liu, Jiamin Dai, Cristina Conati, and Joanna McGrenere. Envisioning AI support during semi-structured interviews across the expertise spectrum. Proceedings of the ACM on Human-Computer Interaction , 9(2):1--29, 2025

  14. [14]

    Scalable requirements elicitation education through simulated interview practice with large language models

    Nelson Lojo. Scalable requirements elicitation education through simulated interview practice with large language models. Master's thesis, University of California, Berkeley, 2025

  15. [15]

    Real world Scrum a grounded theory of variations in practice

    Zainab Masood, Rashina Hoda, and Kelly Blincoe. Real world Scrum a grounded theory of variations in practice. IEEE Transactions on Software Engineering , 48(5):1579--1591, 2022

  16. [16]

    Equity In The Preparation Of Students For Software Engineering Coding Interviews: ChatGPT as a Mock Interviewer

    Marlon Mejias, Zef Vargas, William Ted Edwards, Gloria Washington, Legand Burge, Dale-Marie Wilson, and Luce-Melissa Kouaho. Equity In The Preparation Of Students For Software Engineering Coding Interviews: ChatGPT as a Mock Interviewer . In Congress in Computer Science, Computer Engineering, and Applied Computing , pages 1016--1020, 2023

  17. [17]

    How We Refactor, and How We Know It

    Emerson Murphy-Hill, Chris Parnin, and Andrew Black. How We Refactor, and How We Know It . IEEE Transactions on Software Engineering , 38(1):5--18, 2012

  18. [18]

    Impact of expert interviews in software engineering: Challenges and benefits

    Wajeeha Nasar. Impact of expert interviews in software engineering: Challenges and benefits. In International MultiConference of Engineers and Computer Scientists , 2023

  19. [19]

    Between policy and practice: GenAI adoption in agile software development teams

    Michael Neumann, Lasse Bischof, Nic Elias Hinz, Abdullah Altun, Luca Stockmann, Dennis Schrader, Ana Carolina Ahaus, Erim Can Demirci, Benjamin Gabel, Maria Rauschenberger, Philipp Diebold, Henning Fritzemeier, and Adam Przyby ek. Between policy and practice: GenAI adoption in agile software development teams. In Agile Processes in Software Engineering an...

  20. [20]

    Revisiting the refactoring mechanics

    Jonhnanthan Oliveira, Rohit Gheyi, Melina Mongiovi, Gustavo Soares, M\' a rcio Ribeiro, and Alessandro Garcia. Revisiting the refactoring mechanics. Information and Software Technology , 110:136--138, 2019

  21. [21]

    Refactoring: An Aid in Designing Application Frameworks and Evolving Object-Oriented Systems

    William Opdyke and Ralph Johnson. Refactoring: An Aid in Designing Application Frameworks and Evolving Object-Oriented Systems . In Symposium on Object-Oriented Programming emphasizing Practical Applications , pages 274--282, 1990

  22. [22]

    Creating and editing GPTs

    OpenAI . Creating and editing GPTs . https://help.openai.com/en/articles/8554397-creating-and-editing-gpts, 2026

  23. [23]

    Adoption of large language models in Scrum management: Insights from Brazilian practitioners

    Mirko Perkusich, Danyllo Albuquerque, Allysson Allex Ara \'u jo, Matheus Paix \ a o, Rohit Gheyi, Marcos Kalinowski, and Angelo Perkusich. Adoption of large language models in Scrum management: Insights from Brazilian practitioners. In International Conference on Agile Software Development , pages 255--273, 2026

  24. [24]

    An empirical investigation into the impact of refactoring on regression testing

    Napol Rachatasumrit and Miryung Kim. An empirical investigation into the impact of refactoring on regression testing. In International Conference on Software Maintenance , pages 357--366, 2012

  25. [25]

    Simulating the software engineering interview process using a decision-based serious computer game

    Adrian Rusu, Robert Russell, and Remo Cocco. Simulating the software engineering interview process using a decision-based serious computer game. In International Conference on Computer Games , pages 235--239, 2011

  26. [26]

    Agile Software Development with Scrum

    Ken Schwaber and Mike Beedle. Agile Software Development with Scrum . Prentice Hall, 2001

  27. [27]

    Carolyn B. Seaman. Qualitative methods in empirical studies of software engineering. Transactions on Software Engineering , 25(4):557--572, 1999

  28. [28]

    Why we refactor? Confessions of GitHub contributors

    Danilo Silva, Nikolaos Tsantalis, and Marco T \' u lio Valente. Why we refactor? Confessions of GitHub contributors . In Foundations of Software Engineering , pages 858--870, 2016

  29. [29]

    Automated behavioral testing of refactoring engines

    Gustavo Soares, Rohit Gheyi, and Tiago Massoni. Automated behavioral testing of refactoring engines. IEEE Transactions on Software Engineering , 39(2):147--162, 2013

  30. [30]

    Ethical interviews in software engineering

    Per Erik Strandberg. Ethical interviews in software engineering. In Empirical Software Engineering and Measurement , pages 1--11. IEEE , 2019

  31. [31]

    Barriers to Refactoring

    Ewan Tempero, Tony Gorschek, and Lefteris Angelis. Barriers to Refactoring . Communications of the ACM , 60(10):54--61, 2017

  32. [32]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need . Advances in Neural Information Processing Systems , pages 5998--6008, 2017

  33. [33]

    Collecting qualitative data at scale with large language models: A case study

    Alejandro Villalba, Jennifer Scurrell, Eva Brown, Jason Entenmann, and Madeleine Daepp. Collecting qualitative data at scale with large language models: A case study. Human Computer Interaction , 9(2):1--27, 2025

  34. [34]

    Towards understanding refactoring engine bugs

    Haibo Wang, Zhuolin Xu, Huaien Zhang, Nikolaos Tsantalis, and Shin Hwei Tan. Towards understanding refactoring engine bugs. Transactions on Software Engineering and Methodology , 35(5):1--55, 2026

  35. [35]

    InterFlow : Designing unobtrusive AI to empower interviewers in semi-structured interviews

    Yi Wen, Yu Zhang, Sriram Suresh, Zhicong Lu, Can Liu, and Meng Xia. InterFlow : Designing unobtrusive AI to empower interviewers in semi-structured interviews. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1--21. ACM, 2026

  36. [36]

    Zhou, Wenxi Chen, Huahai Yang, and Changyan Chi

    Ziang Xiao, Michelle X. Zhou, Wenxi Chen, Huahai Yang, and Changyan Chi. If I Hear You Correctly : Building and evaluating interview chatbots with active listening skills. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1--14. ACM, 2020

  37. [37]

    Ziang Xiao, Michelle X. Zhou, Q. Vera Liao, Gloria Mark, Changyan Chi, Wenxi Chen, and Huahai Yang. Tell me about yourself: Using an AI -powered chatbot to conduct conversational surveys with open-ended questions. ACM Transactions on Computer-Human Interaction , 27(3):1--37, 2020

  38. [38]

    He Zhang, Yueyan Liu, Xin Guan, Jie Cai, and John M. Carroll. Harnessing the power of AI in qualitative research: Role assignment, engagement, and user perceptions of ai-generated follow-up questions in semi-structured interviews, 2025. URL: https://arxiv.org/abs/2509.12709, https://arxiv.org/abs/2509.12709 arXiv:2509.12709