REVIEW 2 major objections 6 minor
When coding agents make writing code cheap, the real bottlenecks move to evaluation, review, and specifications that both humans and agents must read.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 08:44 UTC pith:4AMDZFIH
load-bearing objection Solid first interview study of how people actually build SE agents: usable process map and named challenges, with the sample-scope limit already owned by the authors. the 2 major comments →
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
As implementation becomes cheaper when building SE agents, bottlenecks shift rather than disappear: long-standing non-coding work becomes more visible, reviewing and evaluating agent output becomes new and central, evaluation steers iteration, and specifications become versioned dual-audience artifacts. The study characterizes a seven-stage workflow and six challenges (including unreliable evaluation, comprehension debt, and change-nothing-change-everything model updates) together with practices teams adopt.
What carries the argument
A seven-stage agent-building workflow (requirements, evaluation, data, system construction with cheapest-first model strategy plus harness, testing/deployment, human feedback, adaptive maintenance) plus the process shift the authors call evaluation-driven development, in which evaluation is designed early and steers iteration while agent specifications become versioned dual-audience artifacts.
Load-bearing premise
That what twenty largely big-tech, research-engineer builders and eighty survey respondents said in early 2026 is a fair picture of how SE agents are built in practice beyond large technology companies and this short window.
What would settle it
A broader multi-site study of SE-agent builders outside large tech firms, or a later replication after major model and tool shifts, that finds coding still dominates effort, evaluation is still treated as a late gate, or the six named challenges are rare rather than recurring.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an exploratory sequential mixed-methods study of how practitioners build Software Engineering (SE) agents. Through semi-structured interviews with 20 builders from 12 organizations (to thematic saturation) and a validation survey of 80 practitioners, the authors characterize a seven-stage agile-like workflow (requirements, evaluation, data, system construction with cheapest-first model strategy and harness, testing/deployment, human feedback, adaptive maintenance), five process shifts (cheaper implementation at higher abstraction; unmasked and newly created non-coding effort; evaluation-driven development; shrinking role boundaries; specifications as dual-audience versioned artifacts), and six challenges with associated practices (unreliable evaluation signals; provider-side model updates that change behavior without local code changes; safety–performance trade-offs; tacit unwritten knowledge; comprehension debt; broken productivity metrics). Member checking (8 of 15 responders) and survey agreement (71–95% on process findings; 73–85% on challenges) corroborate the interview themes. The central claim is that as implementation becomes cheaper, bottlenecks shift rather than disappear, with evaluation and dual-audience specifications becoming central engineering artifacts.
Significance. If the reported characterization holds for the studied population, the paper is a timely first interview-based account of SE-agent construction, filling a gap left by repository-mining and deployment-focused studies. The seven-stage workflow, the framing of evaluation-driven development as already emerging in practice, and the named constructs (comprehension debt; the “change nothing, change everything” effect; regenerative software as an alternative to preserving implementations) give the community a shared vocabulary for process research and tool design. Strengths include an appropriate mixed-methods design, hybrid card-sorting by three authors, member checking, a broader survey with Welch comparison of recruitment channels, explicit threats, and a replication package. The work is useful to SE process researchers, agent-platform builders, and organizations adopting coding agents, even if external validity is limited to large-tech and early-adopter settings.
major comments (2)
- Abstract and §I frame the contribution as how developers build SE agents “in practice,” while Table I and §VI show 15 of 20 interviewees from big tech, applied-scientist/research-engineer heavy roles, dual builder–user status, and a Feb–May 2026 window. The threats section is candid, but the abstract/intro claims should be scoped more tightly to early-adopter and large-technology builders (or the survey demographics should be used more explicitly to justify broader transfer). Without that alignment, the central “in practice” claim overreaches the sample the design can support.
- §V-E and Table III: practices offered for comprehension debt received markedly weaker survey support (53.9% “return maintenance to AI”; 67.1% “preserve regenerability”) than other practices (typically ≥70%). The RQ2 response correctly notes that comprehension debt “remains unresolved,” yet the abstract and section framing still present these as practices teams adopt. Distinguish established practices from exploratory or minority responses so the load-bearing challenge claim is not diluted by weakly supported remedies.
minor comments (6)
- §III-B3: the reduction from 52 findings to the five process findings and six challenges is described only as “most relevant… and most clearly extending current understanding.” A short selection criterion or mapping (even in the replication package summary) would improve transparency.
- Throughout §V, the “♂lightbulb” markers appear to be broken emoji/markup for Implications callouts; replace with a consistent, journal-safe heading or icon.
- Fig. 2: the dashed-arrow feedback paths and the “conditional component” notation are dense; a one-sentence caption note on which stages are optional (data; model adaptation) would help readers who only skim the figure.
- Table II/III agreement bars are described in the table footer but not rendered as visual bars in the text source; ensure the camera-ready version shows the stacked agreement visualization consistently.
- §II-C and related work: a brief explicit contrast with Nahar et al. (LLM integration at Microsoft) and AgentOps/LLMOps literature on what is new once systems become agentic (vs. LLM-component integration) would sharpen novelty for readers familiar with that line of work.
- Skills findings are omitted for space (§III-B3); a one-sentence pointer in the main text to the replication package’s skills section would avoid the impression that skillset results were discarded rather than deferred.
Circularity Check
No circular derivation: inductive interview findings triangulated by survey, not restatements of fitted inputs or self-definitional claims.
full rationale
This is an exploratory sequential mixed-methods empirical paper (20 semi-structured interviews → hybrid card-sorting thematic analysis → member checking → confirmatory survey of 80 builders). The load-bearing claims—a seven-stage workflow, five process shifts (implementation cheaper; effort unmasked/created; evaluation-driven development; shrinking role boundaries; specifications as dual-audience versioned artifacts), and six challenges with practices—are induced from participant accounts and then checked against a broader sample (Tables II–III report agreement rates). There are no equations, fitted parameters renamed as predictions, uniqueness theorems imported from the authors, or ansatzes smuggled via self-citation that force the result by construction. Survey items deliberately restate interview themes for validation; that is standard triangulation, not circular proof. Self-citations appear in related work and methodology (e.g., prior AI/SE studies) but do not underwrite the central characterization of SE-agent building. Threats (self-report, big-tech sample, temporal window) are external-validity limits the authors state in §VI, not circular reductions. Score 0 is therefore the correct outcome.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Semi-structured interview self-reports from builders, after member checking, adequately represent how SE agents are developed in the studied organizations.
- domain assumption Thematic saturation after ~18–20 interviews plus hybrid card-sorting consensus is sufficient to stabilize the reported workflow, shifts, and challenges.
- domain assumption Survey agreement and effectiveness ratings from 80 filtered practitioners corroborate interview themes beyond the original sample.
- domain assumption SE agents are software systems whose development process can be usefully compared to classical SE, ML-enabled, and LLMOps processes.
invented entities (3)
-
Comprehension debt
independent evidence
-
Regenerative software (as practice)
no independent evidence
-
Change nothing, change everything effect
independent evidence
read the original abstract
The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how developers build these systems in practice: existing studies mine repositories or examine deployment, but few investigate how SE agents are constructed. Through semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners, this paper is the first to study how SE processes are changing in the development of SE agents and what challenges developers face. We find that as implementation becomes cheaper, bottlenecks shift rather than disappear: long-standing work in requirements, coordination, and deployment becomes more visible, while reviewing generated code and evaluating agent behavior become new and increasingly central forms of work. We characterize a seven-stage workflow and five process shifts, including a move toward evaluation-driven development, in which evaluation is increasingly defined early and steers iteration, and the emergence of specifications as first-class artifacts that teams test and version alongside code. We further identify six challenges that teams face, together with 12 corresponding practices they use or propose to address them, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates.
Figures
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.