Pith. sign in

REVIEW 2 major objections 6 minor

When coding agents make writing code cheap, the real bottlenecks move to evaluation, review, and specifications that both humans and agents must read.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 08:44 UTC pith:4AMDZFIH

load-bearing objection Solid first interview study of how people actually build SE agents: usable process map and named challenges, with the sample-scope limit already owned by the authors. the 2 major comments →

arxiv 2607.10856 v2 pith:4AMDZFIH submitted 2026-07-12 cs.SE cs.AIcs.HC

How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study

classification cs.SE cs.AIcs.HC
keywords software engineering agentsLLM agentsevaluation-driven developmentsoftware processcomprehension debtmixed-methods studyAI4SEagent harness
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks how practitioners actually build software-engineering agents—LLM systems that can work on large codebases with limited oversight—and how that work is changing ordinary software process. From interviews with twenty builders across twelve organizations and a survey of eighty more, it finds that implementation has become much cheaper, often because the same agents are used to build the next agents. Bottlenecks do not vanish; they shift. Requirements, coordination, review, deployment, and especially evaluation become more visible, while reviewing agent-generated code becomes new central work. The authors describe a seven-stage agile-like loop in which evaluation is defined early and steers iteration, and in which prompts, skills, and scaffold behavior become versioned engineering artifacts for both people and models. They also surface six recurring challenges—unreliable evaluation signals, silent provider-side model updates that rewrite behavior without local code changes, safety traded for performance, tacit knowledge agents cannot retrieve, comprehension debt as generated code outruns understanding, and productivity metrics that inflate without measuring value—along with practices teams use in response.

Core claim

As implementation becomes cheaper when building SE agents, bottlenecks shift rather than disappear: long-standing non-coding work becomes more visible, reviewing and evaluating agent output becomes new and central, evaluation steers iteration, and specifications become versioned dual-audience artifacts. The study characterizes a seven-stage workflow and six challenges (including unreliable evaluation, comprehension debt, and change-nothing-change-everything model updates) together with practices teams adopt.

What carries the argument

A seven-stage agent-building workflow (requirements, evaluation, data, system construction with cheapest-first model strategy plus harness, testing/deployment, human feedback, adaptive maintenance) plus the process shift the authors call evaluation-driven development, in which evaluation is designed early and steers iteration while agent specifications become versioned dual-audience artifacts.

Load-bearing premise

That what twenty largely big-tech, research-engineer builders and eighty survey respondents said in early 2026 is a fair picture of how SE agents are built in practice beyond large technology companies and this short window.

What would settle it

A broader multi-site study of SE-agent builders outside large tech firms, or a later replication after major model and tool shifts, that finds coding still dominates effort, evaluation is still treated as a late gate, or the six named challenges are rare rather than recurring.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper reports an exploratory sequential mixed-methods study of how practitioners build Software Engineering (SE) agents. Through semi-structured interviews with 20 builders from 12 organizations (to thematic saturation) and a validation survey of 80 practitioners, the authors characterize a seven-stage agile-like workflow (requirements, evaluation, data, system construction with cheapest-first model strategy and harness, testing/deployment, human feedback, adaptive maintenance), five process shifts (cheaper implementation at higher abstraction; unmasked and newly created non-coding effort; evaluation-driven development; shrinking role boundaries; specifications as dual-audience versioned artifacts), and six challenges with associated practices (unreliable evaluation signals; provider-side model updates that change behavior without local code changes; safety–performance trade-offs; tacit unwritten knowledge; comprehension debt; broken productivity metrics). Member checking (8 of 15 responders) and survey agreement (71–95% on process findings; 73–85% on challenges) corroborate the interview themes. The central claim is that as implementation becomes cheaper, bottlenecks shift rather than disappear, with evaluation and dual-audience specifications becoming central engineering artifacts.

Significance. If the reported characterization holds for the studied population, the paper is a timely first interview-based account of SE-agent construction, filling a gap left by repository-mining and deployment-focused studies. The seven-stage workflow, the framing of evaluation-driven development as already emerging in practice, and the named constructs (comprehension debt; the “change nothing, change everything” effect; regenerative software as an alternative to preserving implementations) give the community a shared vocabulary for process research and tool design. Strengths include an appropriate mixed-methods design, hybrid card-sorting by three authors, member checking, a broader survey with Welch comparison of recruitment channels, explicit threats, and a replication package. The work is useful to SE process researchers, agent-platform builders, and organizations adopting coding agents, even if external validity is limited to large-tech and early-adopter settings.

major comments (2)
  1. Abstract and §I frame the contribution as how developers build SE agents “in practice,” while Table I and §VI show 15 of 20 interviewees from big tech, applied-scientist/research-engineer heavy roles, dual builder–user status, and a Feb–May 2026 window. The threats section is candid, but the abstract/intro claims should be scoped more tightly to early-adopter and large-technology builders (or the survey demographics should be used more explicitly to justify broader transfer). Without that alignment, the central “in practice” claim overreaches the sample the design can support.
  2. §V-E and Table III: practices offered for comprehension debt received markedly weaker survey support (53.9% “return maintenance to AI”; 67.1% “preserve regenerability”) than other practices (typically ≥70%). The RQ2 response correctly notes that comprehension debt “remains unresolved,” yet the abstract and section framing still present these as practices teams adopt. Distinguish established practices from exploratory or minority responses so the load-bearing challenge claim is not diluted by weakly supported remedies.
minor comments (6)
  1. §III-B3: the reduction from 52 findings to the five process findings and six challenges is described only as “most relevant… and most clearly extending current understanding.” A short selection criterion or mapping (even in the replication package summary) would improve transparency.
  2. Throughout §V, the “♂lightbulb” markers appear to be broken emoji/markup for Implications callouts; replace with a consistent, journal-safe heading or icon.
  3. Fig. 2: the dashed-arrow feedback paths and the “conditional component” notation are dense; a one-sentence caption note on which stages are optional (data; model adaptation) would help readers who only skim the figure.
  4. Table II/III agreement bars are described in the table footer but not rendered as visual bars in the text source; ensure the camera-ready version shows the stacked agreement visualization consistently.
  5. §II-C and related work: a brief explicit contrast with Nahar et al. (LLM integration at Microsoft) and AgentOps/LLMOps literature on what is new once systems become agentic (vs. LLM-component integration) would sharpen novelty for readers familiar with that line of work.
  6. Skills findings are omitted for space (§III-B3); a one-sentence pointer in the main text to the replication package’s skills section would avoid the impression that skillset results were discarded rather than deferred.

Circularity Check

0 steps flagged

No circular derivation: inductive interview findings triangulated by survey, not restatements of fitted inputs or self-definitional claims.

full rationale

This is an exploratory sequential mixed-methods empirical paper (20 semi-structured interviews → hybrid card-sorting thematic analysis → member checking → confirmatory survey of 80 builders). The load-bearing claims—a seven-stage workflow, five process shifts (implementation cheaper; effort unmasked/created; evaluation-driven development; shrinking role boundaries; specifications as dual-audience versioned artifacts), and six challenges with practices—are induced from participant accounts and then checked against a broader sample (Tables II–III report agreement rates). There are no equations, fitted parameters renamed as predictions, uniqueness theorems imported from the authors, or ansatzes smuggled via self-citation that force the result by construction. Survey items deliberately restate interview themes for validation; that is standard triangulation, not circular proof. Self-citations appear in related work and methodology (e.g., prior AI/SE studies) but do not underwrite the central characterization of SE-agent building. Threats (self-report, big-tech sample, temporal window) are external-validity limits the authors state in §VI, not circular reductions. Score 0 is therefore the correct outcome.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 3 invented entities

This is an empirical process study, not a formal derivation. Load-bearing premises are methodological and sampling assumptions rather than free physical parameters. The central claims rest on treating practitioner self-report plus survey agreement as evidence of real process change; on thematic saturation and consensus coding as adequate for theme validity; and on conceptual labels (comprehension debt, regenerative software, evaluation-driven development as practiced) as faithful compressions of the data. No numerical constants are fitted to produce the main claims.

axioms (4)
  • domain assumption Semi-structured interview self-reports from builders, after member checking, adequately represent how SE agents are developed in the studied organizations.
    Core epistemic premise of the interview study (Sections III–V); authors themselves flag self-report and builder bias in Threats.
  • domain assumption Thematic saturation after ~18–20 interviews plus hybrid card-sorting consensus is sufficient to stabilize the reported workflow, shifts, and challenges.
    Invoked in interview analysis and stopping rule (Section III-B2/B3); standard qualitative SE practice but not independently verified by second coding team with IRR.
  • domain assumption Survey agreement and effectiveness ratings from 80 filtered practitioners corroborate interview themes beyond the original sample.
    Used throughout Tables II–III and RQ responses; depends on item wording matching interview constructs and on valid hands-on experience screening.
  • domain assumption SE agents are software systems whose development process can be usefully compared to classical SE, ML-enabled, and LLMOps processes.
    Framing assumption in Introduction and Related Work that licenses process-shift claims.
invented entities (3)
  • Comprehension debt independent evidence
    purpose: Name the pattern where agent-generated code is accepted and shipped faster than developers can understand, review, and safely maintain it.
    Introduced from interview themes and Fig. 3; related to prior 'cognitive debt' citations but operationalized here for coding-agent output volume and review backlog.
  • Regenerative software (as practice) no independent evidence
    purpose: Describe an emerging alternative of preserving specs, tests, and infrastructure to regenerate implementations rather than maintaining a particular code artifact.
    Attributed mainly to one participant vision (P8) and offered as a response to comprehension debt; survey effectiveness for the related practice is only moderate (~67%).
  • Change nothing, change everything effect independent evidence
    purpose: Label provider-side model updates that alter agent behavior and invalidate harnesses without local code/prompt changes.
    Complement to Sculley's CACE principle, grounded in multiple interview quotes; survey agreement high (~80%).

pith-pipeline@v1.1.0-grok45 · 25270 in / 3181 out tokens · 41174 ms · 2026-07-14T08:44:04.393126+00:00 · methodology

0 comments
read the original abstract

The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how developers build these systems in practice: existing studies mine repositories or examine deployment, but few investigate how SE agents are constructed. Through semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners, this paper is the first to study how SE processes are changing in the development of SE agents and what challenges developers face. We find that as implementation becomes cheaper, bottlenecks shift rather than disappear: long-standing work in requirements, coordination, and deployment becomes more visible, while reviewing generated code and evaluating agent behavior become new and increasingly central forms of work. We characterize a seven-stage workflow and five process shifts, including a move toward evaluation-driven development, in which evaluation is increasingly defined early and steers iteration, and the emergence of specifications as first-class artifacts that teams test and version alongside code. We further identify six challenges that teams face, together with 12 corresponding practices they use or propose to address them, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates.

Figures

Figures reproduced from arXiv: 2607.10856 by Chao Peng, David Lo, David Williams, Federica Sarro, Jieke Shi, Yunbo Lyu, Zhensu Sun, Zhou Yang.

Figure 1
Figure 1. Figure 1: Overview of our research method. We adopt an exploratory sequential mixed-methods design. The interview study proceeds from scoping and interview [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Workflow for building SE agents. The process forms an agile-like loop from requirements and early evaluation to construction, deployment, feedback, [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The formation and responses to comprehension debt. Coding-agent [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.