{"id":"a4495240-d607-4d5e-8cda-92c19f8fa320","arxiv_id":"2607.10856","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SE-agent development follows a recurring seven-stage loop where evaluation drives iteration, and challenges such as unreliable evaluation signals and comprehension debt emerge.","lead":"Through interviews with 20 practitioners and a survey of 80 more, this study maps how teams build AI coding agents. It finds a seven-stage workflow and shows that cheaper code shifts effort to evaluation, review, and maintenance rather than eliminating these tasks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Interview guide was iteratively revised during data collection; the seven-stage workflow may be an artifact of researcher steering rather than a discovered pattern.","rationale":"The reader's weakest assumption focuses on the lack of inter-rater reliability in the card-sorting analysis and on sample representativeness. I share the general concern about subjective interpretation, but the more specific and load-bearing issue is that the data collection itself was iterative: the researchers refined a provisional workflow during the study and adjusted interview probes accordingly. This threatens the central claim of a 'recurring' workflow more directly than the absence of a kappa statistic, because it introduces a possible confirmation pathway from the researchers' evolving model back into the interview data. The reader did not explicitly identify this data-collection steering; hence partial agreement. Despite this concern, the paper is transparent about its methods and limitations, frames the workflow as an iterative/agile-like pattern with a conditional component, and reports survey corroboration. The appropriate verdict remains CONDITIONAL (as the reader already assigned), so I recommend UNCHANGED. A decisive test would require access to the full interview materials and independent re-analysis; without that, the workflow should be treated as a tentative framework.","tokens_in":19989,"tokens_out":6740,"duration_ms":70628,"concrete_test":"Using the replication package [58], compare the interview guide used in the first five interviews with the version used in the last five. If the later guide contains probes that explicitly mention workflow stages (e.g., 'How does evaluation steer your iteration?', 'How do you handle model updates?') that were absent from the earlier guide, this supports the steering concern. Additionally, have two independent coders who are blind to the final workflow apply the same hybrid card-sorting method to the first five and last five interview transcripts (if available) and measure inter-rater agreement (e.g., Cohen's kappa) on whether the seven stages appear without prompting. If agreement is low, or if the early transcripts do not naturally yield the seven stages, the workflow's empirical grounding is weak.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a recurring seven-stage workflow (requirements, evaluation, data, construction, testing/deployment, human feedback, adaptive maintenance) genuinely reflects practitioner processes, with evaluation as a central steering mechanism. Section III-B2 states that after the 5th, 10th, and 15th interviews, the researchers 'discussed newly emerging findings and refined a provisional development workflow,' and that these discussions 'also informed adjustments to the questions and probes used in subsequent interviews.' This means the interview protocol was not fixed: later participants were asked questions informed by a provisional workflow built from earlier data. Consequently, the reported 'recurrence' of the workflow and the saturation claim (Section III-B2: 'After 18 interviews, thematic saturation was approaching') may be partly self-fulfilling—if later interviews probed for the provisional stages, participants' confirmations would inflate apparent recurrence. The survey validation (91% agreement with W1) does not break this circularity because the survey statements were themselves derived from the same evolving analysis. The paper acknowledges confirmation bias as a threat (Section VI) and uses member checking, but member checking only confirms that participants recognize the researchers' framing, not that the workflow would emerge independently. Thus, the strongest empirical claim—that the seven-stage loop is a discovered pattern rather than a researcher-imposed frame—rests on an insecure methodological foundation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods study of how practitioners build LLM-based software engineering (SE) agents. The authors conducted semi-structured interviews with 20 practitioners from 12 organizations (mostly large technology companies), analyzed the transcripts using thematic analysis and hybrid card sorting, and then ran an online survey of 80 practitioners to validate the findings. The central claim is that building SE agents does not eliminate traditional software engineering but reorganizes it into a recurring seven-stage loop (requirements, evaluation, data, system construction, testing/deployment, human feedback, adaptive maintenance) in which evaluation increasingly acts as a central steering mechanism. The paper also identifies five process shifts (cheaper implementation, unmasked/created effort, evaluation-driven development, shrinking role boundaries, specifications as first-class artifacts), six challenges, and 12 associated practices. The authors report that member checking confirmed the findings and that the survey showed broad agreement (91% for the workflow, 71–95% for the shifts).","tokens_in":20233,"tokens_out":6341,"duration_ms":63910,"significance":"If the empirical characterization is trustworthy, this is a valuable contribution: it is among the first interview-based studies of SE-agent construction, and it provides concrete evidence for process changes that have previously been anecdotal. The proposed constructs—evaluation-driven development as an observed practice, 'comprehension debt', the 'change nothing, change everything' effect, and 'regenerative software'—are plausible and useful for future research. The paper is also transparent about its sample and limitations, and it commits to a replication package. However, the independence of the evidence for the central workflow is weaker than the text suggests: the interview guide was revised on the basis of provisional analyses, the coding used consensus without inter-rater reliability, and the survey items were derived from the same researchers' interpretations. These issues are addressable, but they currently leave the central claim more vulnerable than the paper's confident framing warrants.","major_comments":[{"comment":"The paper states that after the 5th, 10th, and 15th interviews the researchers 'discussed newly emerging findings and refined a provisional development workflow' and that these discussions 'also informed adjustments to the questions and probes used in subsequent interviews.' Because later participants were asked about a workflow that had already been constructed from earlier data, the claim that the workflow 'recurs' across participants (Section IV-A) and that saturation was approached after 18 interviews is weaker than it appears: the recurring pattern may be partly an artifact of the probes. Please (i) report exactly what was changed in the protocol after each revision point, (ii) show whether the seven stages are present in the first five interviews before any provisional workflow was discussed, and (iii) provide a saturation analysis that separates themes mentioned spontaneously from","section":"Section III-B2"},{"comment":"The analysis section reports that the three authors 'jointly clustered the units and resolved disagreements through discussion and consensus rather than independent coding or inter-rater reliability.' This is a legitimate mode of analysis, but it makes the central workflow's validity depend on the researchers' joint judgment. The member-checking step (Section III-B4) only shows that participants recognize the researchers' framing, not that the workflow would emerge independently. Please provide the codebook and stage definitions, an audit trail of the card-sorting process, and at least a subset of transcripts independently coded by two researchers with agreement reported. If inter-rater reliability is not used, justify this choice in relation to standard qualitative practices and provide alternative trustworthiness evidence such as negative-case or deviant-case analysis.","section":"Section III-B3"},{"comment":"The survey is described as validating the interview findings, but the survey statements are reformulations of the same analysis that produced the workflow and shifts. The 91% agreement for W1 (Table II) therefore does not break the circularity: respondents are endorsing the authors' interpretation, not independently discovering the same structure. The survey is better characterized as an extension of the interview findings to a wider population than as an independent corroboration. Please reframe the language accordingly and, if possible, include at least some survey items that are direct quotes or minimally paraphrased raw statements from participants, and report open-ended survey responses separately.","section":"Section III-C1"},{"comment":"Figure 2 defines data as a 'conditional component: not appear in every team,' yet the paper repeatedly describes the workflow as a 'seven-stage' loop (e.g., the Section IV-A takeaway and Section IV-C). If one of the seven components is absent for some teams, the universality of the seven-stage structure is unclear. The paper should either restrict the claim to teams that perform data work or state more precisely which components are invariant and which are conditional.","section":"Section IV-A / Fig. 2"}],"minor_comments":[{"comment":"The text contains rendering artifacts such as '♂lightbulbImplications' and '/char◎-barx%'. These should be replaced with standard labels and percentage formatting.","section":"Section V"},{"comment":"Please clarify whether the 'Agreement' column is the percentage of respondents choosing 'Agree' or 'Strongly agree' (or the sum), and report the number of respondents per row, since some items had an 'I don't know' option.","section":"Tables II and III"},{"comment":"The criteria for selecting 'five process findings and six challenges' from the 52 findings are under-specified. State the selection rule explicitly or provide the full list of findings in the replication package.","section":"Section III-B3"},{"comment":"The confirmation-bias threat is acknowledged in the threats section, but its concrete implications for the iterative guide (Section III-B2) and consensus coding (Section III-B3) should also be discussed in the methods section, not only in the limitations.","section":"Section VI"},{"comment":"The abstract claims this is the first study of process changes in SE agent development; the text more cautiously says 'To our knowledge.' Keep the cautious phrasing in the abstract as well.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the SE community, but the iterative-protocol issue is the main risk. I would encourage the authors to use the revision to provide a transparent account of how the workflow evolved over the interviews and to weaken the validation language. The replication package should be checked for the interview guide versions and codebook. The survey's Prolific portion, while screened, may introduce a population mismatch; the Welch test result is reassuring but should be interpreted alongside the qualitative sample's big-tech skew."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth your time. This is the first interview study I know of that looks specifically at how teams build SE agents, and that alone makes it a useful reference for anyone working on AI4SE. The authors did 20 interviews across 12 orgs and an 80-person survey, and they are transparent about their method and their limitations. The headline contributions—the seven-stage loop (requirements, evaluation, data, construction, test/deploy, human feedback, adaptive maintenance), evaluation-driven development, comprehension debt, and the 'change nothing, change everything' effect—are genuinely clarifying. These are not just re-labeled old ideas; they capture things practitioners actually describe, like tests becoming the oracle just because they exist, and provider model updates invalidating harness work. The paper gives proper credit to prior work and is careful to distinguish what is new.\n\nThe soft spots are real but not fatal. The biggest is the adaptive interview protocol. After the 5th, 10th, and 15th interviews, the researchers discussed emerging findings, refined a provisional workflow, and adjusted the interview questions and probes accordingly. That means later interviews were partly probing for the stages the researchers already had in mind. The 'recurring' workflow and the saturation claim are then less independent than they look. The survey doesn't break the circularity because the survey statements were derived from the same evolving analysis. The authors acknowledge confirmation bias as a threat in Section VI, and they do member checking, but member checking only confirms participants recognize the framing; it doesn't show the workflow would emerge independently. I'd like the paper to be more explicit that the seven-stage loop is a researcher-facilitated synthesis rather than a pattern that simply 'recurred' in the data.\n\nAlso worth noting: consensus coding without inter-rater reliability is common in qualitative SE research, but it does mean another team might cluster the meaning units differently. And the sample is heavily big-tech and applied-scientist. Those limits are acknowledged, and the survey's demographic spread is wider, but the core workflow evidence rests on the 20 interviews.\n\nOn balance, this is a solid, honest study that organizes a nascent field. The process model should be treated as a tentative framework, not a discovered law. I'd send it to peer review—it deserves serious referee time and a request for a sharper treatment of the protocol-evolution issue. I'd also cite it when writing about agent development practices.\n\nBest.","headline":"First real interview study of how practitioners build SE agents; the process model is useful but the adaptive interview protocol makes the 'discovered' workflow less secure than claimed.","tokens_in":20731,"tokens_out":2341,"would_cite":true,"duration_ms":24125,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Building SE agents reorganizes the old software process into a seven-stage loop steered by evaluation.","keywords":["software engineering agents","LLM agents","development workflow","evaluation-driven development","comprehension debt","agent harness","mixed-methods study","software process"],"falsifier":"Have two independent teams, blinded to this paper's framework, segment and cluster the same 20 interview transcripts; if they do not converge on a similar stage structure, or if measured agreement between independent coders on assigning passages to the seven stages is low, then the claimed recurring workflow would be an artifact of the authors' coding rather than a property of practitioner experience.","tokens_in":19846,"feed_emoji":"🔁","tokens_out":4027,"duration_ms":45722,"temperature":0.7,"pith_summary":"The paper tries to establish, from interviews with 20 practitioners who build LLM-based software-engineering agents and a survey of 80 more, that agent development does not abolish the familiar software process but reorganizes it. Its central claim is a recurring seven-stage loop — requirements, evaluation, data, construction, testing/deployment, human feedback, adaptive maintenance — in which evaluation is defined early and repeatedly steers iteration. If true, it gives teams a concrete picture of where work actually goes: coding becomes cheaper, but specification, review, coordination, and deployment remain bottlenecks, and new kinds of work (reviewing generated code, evaluating agent behavior) take center stage. The paper also names six practical failure points, including unreliable evaluation signals and 'comprehension debt,' along with twelve practices teams use. A sympathetic reader would take this as a systematic map of an emerging engineering discipline.","feed_headline":"SE agents are built in a seven-stage loop, evaluation-first","feed_subtitle":"20 practitioner interviews and 80 survey responses show cheaper coding shifts the bottleneck to review and evaluation","key_machinery":"The central object is the seven-stage workflow itself, with evaluation as its steering mechanism: evaluation criteria are set early, reused during construction, and revisited after deployment, so that every other stage is organized around producing and interpreting an evaluation signal. The second key piece is the combination of a cheapest-first model strategy (API/prompt, then LoRA or fine-tuning, then pretraining) with a persistent harness — the wrapper that supplies context, memory, tools, skills, permissions, and orchestration — which survives model swaps and becomes the main locus of engineering effort. This machinery explains both why evaluation recurs throughout the loop and why provi","core_discovery":"The discovery is that building SE agents in practice follows an agile-like seven-stage loop rather than a linear pipeline, with evaluation as the recurring mechanism that decides whether an iteration is good enough to continue. Construction follows a cheapest-first model strategy — start with prompts on an existing model, escalate to fine-tuning, only rarely pretrain — while a shared harness (the layer managing context, memory, tools, skills, permissions, and orchestration) persists across model choices. Evaluation is defined early and reused at every stage, and specifications such as prompts, skills, and context definitions become first-class artifacts tested and versioned alongside code. T","pith_inferences":["The 'regenerative software' idea could be tested by comparing teams that preserve specifications, tests, and skills as the durable asset against teams that preserve the generated artifact itself, measuring maintainability and time-to-repair over several months.","The 'change nothing, change everything' effect implies that model choice becomes a coupling point: SE-agent teams may need model-agnostic harness abstractions and explicit upgrade/pinning policies, analogous to dependency management for the model itself.","If evaluation-steered development generalizes, the natural next step is to check whether the same seven-stage loop appears in organizations that are not big tech or in agent-building teams with less research exposure; the paper's own sample skew leaves that open.","Comprehension debt may be measurable with proxy signals such as growth in the unreviewed pull-request backlog or the time from merge to a developer's first successful explanation of a change — signals future work could turn into dashboard metrics."],"forward_implications":["If the seven-stage loop is accurate, teams should expect to spend more early effort on requirements and evaluation design, because those stages now steer all subsequent iteration.","If evaluation-driven development is real, agent-building teams need continuous evaluation infrastructure — regression-style benchmarks and production-derived signals — rather than a single final validation pass.","If specifications are first-class artifacts, then prompts, skills, context definitions, and scaffold configuration should be versioned, reviewed, and tested in the same pipeline as code.","If 'change nothing, change everything' holds, provider model updates are a routine source of behavioral drift that teams must monitor and re-evaluate against even when their own code is unchanged.","If comprehension debt accumulates, code volume alone is a misleading productivity metric; review burden and the ability to later modify generated code are the more relevant costs."],"fun_headline_variants":["SE agents: 7-step loop, evaluation-first, specs as code","Building SE agents: evaluation steers, specs become code","Evaluation-first loop: how teams build SE agents","Seven-stage SE agent loop: eval first, specs as artifacts","Coding cheap, review central: SE agent building evolves"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"That the seven-stage workflow and the five process shifts genuinely reflect how SE-agent builders work, rather than how the three researchers who coded the 20 interview transcripts chose to group the material, since the categories were settled by consensus without independent coding or inter-rater reliability checks.","fun_headline_variants_meta":{"raw":{"variants":["SE agents: 7-step loop, evaluation-first, specs as code","Building SE agents: evaluation steers, specs become code","Evaluation-first loop: how teams build SE agents","Seven-stage SE agent loop: eval first, specs as artifacts","Coding cheap, review central: SE agent building evolves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000674,"raw_usage":{"total_tokens":2903,"prompt_tokens":739,"completion_tokens":2164,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":2081}},"tokens_in":483,"tokens_out":2164,"duration_ms":13823,"temperature":1.0,"reasoning_tokens":2081,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T07:04:53.540309+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent teams, blinded to this paper's framework, segment and cluster the same 20 interview transcripts; if they do not converge on a similar stage structure, or if measured agreement between independent coders on assigning passages to the seven stages is low, then the claimed recurring workflow would be an artifact of the authors' coding rather than a property of practitioner experience.","supporting_citations":[],"review_version":2}