{"id":"fb136ac7-c20c-4ba0-bc66-5e60b05c376a","arxiv_id":"2507.07548","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Developers do not feed raw requirements into LLMs; they manually decompose requirements into programming tasks and enrich prompts with design, architecture, and code context before generating code.","lead":"Interviews with 18 developers from 14 companies show that requirements alone are too vague to feed into AI code generators; developers first break them into small programming tasks and add design and architecture details. The study maps the human work between requirement and AI-generated code, which matters for anyone betting that AI will automate software engineering end to end.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The necessity claim that raw requirements cannot be used directly rests on retrospective self-reports; no logs or observed prompts verify it, and the paper's own participant table includes issue-based IDE users who are not among those reported as decomposing requirements.","rationale":"The reader's weakest assumption identifies exactly the same vulnerability: the study relies on retrospective interview accounts rather than direct observation of behavior. Our stress test sharpens that concern by pointing to the paper's own participant table: issue-based IDE users are not among the 12 participants explicitly listed as confirming decomposition, yet the process model treats all requirements artifacts uniformly. This is not an internal contradiction that invalidates the qualitative theory, but it does underdetermine the strong 'must first be manually decomposed' phrasing in Takeaway 1 and the abstract. The paper is otherwise careful: it provides a replication package, documents its coding procedure, acknowledges self-report and translation limitations, and frames the result as a grounded theory. Within qualitative research conventions, an ACCEPT verdict with moderate confidence is appropriate, and the concern we raise is best handled by future observational validation rather than by rejecting or conditionally reframing the present study. We therefore leave the reader's verdict unchanged.","tokens_in":14303,"tokens_out":5864,"duration_ms":76923,"concrete_test":"Run a two-week observational study with at least 20 developers in which IDE/chat telemetry records every prompt, whether the prompt contains raw requirement text or a decomposed programming task, and whether the generated code was integrated with no more than minor edits. If a non-negligible fraction of successful integrations originate from raw-requirement prompts, Takeaway 1 must be weakened from a necessity claim to a dominant-practice claim; if none do, the concern is settled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Takeaway 1 (Section IV-A) asserts that practitioners do not use traditional requirements artifacts as input and must first manually decompose them into programming tasks. The evidence for this necessity claim is retrospective interview accounts (Section III-B), not observed prompting behavior. Only 12 of 18 participants are listed as confirming the decomposition finding (P01, P04–P11, P13, P15, P17), while participants whose requirements are unstructured Issues used with IDE-integration (P02, P03, P18 in Table I) are absent from that list. If these developers prompt directly from issue-tracker entries with auto-completion, the process model's single 'Requirements artifact' node overstates the universality of manual decomposition. The authors report asking participants to show prompts and artifacts, but no such artifact-based evidence appears in the Findings; the central inference therefore depends entirely on what participants said they do, not on what they demonstrably did. This matters because the headline claim is normative: requirements 'must first' be decomposed. Without interaction logs, prompts, or observed code outcomes, the paper cannot distinguish a general property of requirements from a practice shaped by the participants' chosen tools, prompt styles, or their pilot-program context at their companies.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a semi-structured interview study with 18 practitioners from 14 companies to understand how software engineers incorporate requirements and design information when using large language models for code generation. Based on inductive in-vivo coding of 179 quotes, the authors propose a process model (requirements artifacts → programming tasks → prompt construction → code generation, technical exploration, or manual coding) and a content model (what context information is added to prompts). The central claim is that traditional requirements artifacts are too abstract for direct use as LLM input, and practitioners must first manually decompose them into programming tasks, which are then enriched with design decisions and architectural constraints. The paper also identifies three interaction patterns: incremental code generation, manual coding with intelligent auto-completion, and extensive code generation. The authors discuss implications for requirements engineering research and for the feasibility of fully automated requirements-to-code systems.","tokens_in":14522,"tokens_out":5146,"duration_ms":58550,"significance":"If the theory is accepted, it directly challenges the premise of many NL2Code benchmarks and automated requirements-to-code proposals, which treat raw requirements or task descriptions as sufficient prompt input. The study is one of the first to examine real-world requirements artifacts in LLM-assisted implementation, and its process/content models provide a useful vocabulary for future research. Strengths include a transparent qualitative method, a replication package with interview guide and codebook, explicit tracing from in-vivo codes to themes, multi-company and multi-domain sampling, and acknowledged threats to validity. The paper does not claim statistical generalization and is careful to position its contribution as theory building. The main weakness is that the headline normative claim rests entirely on retrospective self-reports, with inconsistencies in the reported participant support and no artifact-based evidence in the findings.","major_comments":[{"comment":"The universal normative claim is not fully supported by the reported evidence. The text states that practitioners 'unanimously stated' they first derive smaller units from requirements artifacts, but the list of confirming participants contains only 12 of 18 (P01, P04, P05, P06, P07, P08, P09, P10, P11, P13, P15, P17). The three participants who used unstructured Issues with IDE integration (P02, P03, P18 in Table I) are absent from that list, and their interaction pattern is described as 'manual coding with intelligent auto-completion,' in which suggestions are accepted while writing code as usual. These cases may not involve a separate decomposition step at all, so the process model's single 'Requirements artifact' entry node and the abstract's 'must first be manually decomposed' overstate universality. Please either supply direct evidence (prompts, artifacts, or observed sessions) for these cases, or revise the process model and wording to present decomposition as a dominant reported pattern rather than a universal necessity.","section":"Section IV-A, Takeaway 1 and abstract"},{"comment":"The central empirical support is retrospective self-report, and the stated attempt to collect artifacts is not reported in the Findings. Section III-B says participants were asked to 'show us intermediate artifacts such as user stories, prompt templates, and prompts, if possible,' but the Findings contain no such artifact-based data; no logs, prompt snapshots, or observed interaction data are presented. The claims about what practitioners 'do not use' and 'must' do are therefore based entirely on what participants said they do, not on what they demonstrably did. I am not asking for quantitative validation, but the manuscript should either report whatever artifact evidence was collected or explicitly restrict the theory's scope to stated practices and soften the modal language in Takeaway 1 and the abstract.","section":"Section III-B and Section IV-A"},{"comment":"There is an internal inconsistency in the evidence reporting. The first paragraph of this subsection says practitioners 'unanimously stated' that requirements artifacts must be decomposed, while two paragraphs later the text says 'Almost all interviewees confirmed this notion' and lists 12 of 18 participants. 'Unanimously' and 'almost all' cannot both describe the same finding. Please use a single, accurate formulation and, if necessary, explain how the non-listed participants were coded.","section":"Section IV-A, 'Deriving Programming Tasks From Requirements Artifacts'"}],"minor_comments":[{"comment":"The author name appears as 'V ogelsang' in the header and in references [27] and [44]; this is presumably a typesetting error and should be corrected to 'Vogelsang'.","section":"Throughout"},{"comment":"Please clarify whether the DeepL-translated German transcripts were reviewed against the original recordings or German text by a second researcher, as the fidelity of the quoted statements is otherwise dependent on a single machine translation.","section":"Section III-B"},{"comment":"The loop from 'Code adjustment' back to 'Refine programming task' and the optional paths around 'Use auto-completion' are visually dense; consider labeling the loop-back condition explicitly or adding a note in the caption.","section":"Figure 1"},{"comment":"The parenthesized participant IDs after each content category are useful, but a small table mapping each content category to the confirming participants would make the evidence easier to verify.","section":"Section IV-B"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the venue and makes a worthwhile contribution to requirements engineering and LLM-assisted software engineering. My main concern is calibrating the strength of the central takeaway to the qualitative evidence: the 'must first be manually decomposed' claim is presented as universal even though the reported participant support is partial. This is fixable by rewording and by either adding artifact-based evidence or explicitly limiting the scope to reported practices. I do not see a need for additional data collection at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Mark — this one is worth a read and worth a referee. The paper is a qualitative interview study (18 practitioners, 14 companies) about how developers use requirements when they generate code with LLMs. The headline finding — that raw requirements are too abstract for direct prompting, and that developers manually decompose them into 'programming tasks' enriched with design context — is clearly documented and fills a real gap. Existing NL2Code work uses contest descriptions as requirements; interaction studies stopped at the prompt. This paper shows the upstream process that comes before a prompt is ever written.\n\nWhat it does well: the method is honestly reported. In-vivo coding, 179 quotes, themes cross-checked by all authors, replication package with interview guide and codebook. The process and content models (Figures 1 and 2) are intuitive and will be useful vocabulary for future research. The related-work section correctly identifies the gap, and the authors cite their own prior work as background rather than as a crutch.\n\nThe soft spot is proportionate but real. The 'must' in Takeaway 1 and the abstract is stronger than the evidence. The data are retrospective self-reports; no logs, no observed prompts, no artifacts shown. The authors mention asking participants to show prompts and artifacts, but nothing like that appears in the findings. Also, the participant table lists P02, P03, and P18 as using unstructured Issues with IDE-integration, and none of them appear in the list of 12 who confirmed the decomposition finding. That is consistent with the 'manual coding with auto-completion' pattern the paper itself describes, but it undercuts the universal phrasing of Takeaway 1. The evidence supports 'many practitioners describe decomposing requirements before prompting'; it does not prove that requirements logically must be decomposed. That's a fixable defect, not a fatal one — the paper is explicit about its threats to validity and does not hide the self-report nature of the data.\n\nThe paper is for requirements engineering researchers, tool builders, and anyone working on Req2Code. It is a solid qualitative contribution that deserves a serious referee. The right referee request is to temper the necessity claim and to discuss the issue-based IDE users more directly. Send it to peer review.","headline":"A useful and honest interview study of how developers actually get from requirements to LLM-based code, with one 'must' claim that outruns the self-reported evidence.","tokens_in":15009,"tokens_out":2391,"would_cite":true,"duration_ms":28159,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Documented requirements are too abstract to feed directly into code-generating LLMs; developers must first decompose them into programming tasks and add design and architectural context before prompting.","keywords":["requirements engineering","LLM-assisted code generation","interview study","programming tasks","prompt engineering","requirements-to-code","design decisions","software engineering process"],"falsifier":"A log-based field study recording the actual prompts, accepted completions, and code diffs of developers over several weeks would settle the claim: if raw requirement texts are regularly pasted verbatim into prompts and the resulting code is integrated without any prior decomposition into programming tasks, the central claim is contradicted; if every code-generating prompt is preceded by a hand-written task breakdown, the theory is supported.","tokens_in":14111,"feed_emoji":"🧩","tokens_out":7848,"duration_ms":76779,"temperature":0.7,"pith_summary":"This paper tries to establish a theory, grounded in interviews with practitioners, of how developers actually use requirements and design artifacts when generating code with large language models. Based on 18 interviews across 14 companies, it argues that traditional requirements artifacts such as user stories and functional requirements are not useful as direct LLM inputs. Before prompting, developers manually decompose requirements into programming tasks and enrich those tasks with context: design decisions, architectural constraints, code context, and constraints on infrastructure, interfaces, languages, and tests. If this is right, the vision of fully automatic requirements-to-code generation is missing the core human process that production code generation depends on, and requirements engineering remains an essential activity in LLM-assisted software development.","feed_headline":"Interviews show raw requirements fail as LLM prompts","feed_subtitle":"A 14-company study finds developers must decompose requirements and add design context before generated code is usable.","key_machinery":"The central object is the programming task, the intermediate artifact that sits between a documented requirement and an LLM prompt. It carries the theory because it is both the output of the manual decomposition step and the core component of every code-generation prompt. The two models are the machinery: the process model traces a requirements artifact through decomposition into programming tasks, context construction (ad-hoc or elaborate), code generation, checking, and adjustment, with three interaction patterns; the content model specifies the five context categories—language and libraries, interface and data format, infrastructure and deployment, business logic and algorithms, and unit tests—plus the code context that practitioners add so the generated code can be integrated.","core_discovery":"The study's central claim is that LLM-assisted implementation does not bypass requirements engineering; it depends on it. Practitioners reported that pasting a user story or a catalog requirement into a chat or IDE assistant yields code that cannot be integrated into the existing code base. Instead, they derive programming tasks—smaller units that specify how a requirement is to be realized in code—and then construct prompts around those tasks, adding information about the programming language and libraries, interfaces and data formats, infrastructure and deployment constraints, unit tests, and the relevant code context. The paper organizes this into a process model (requirements artifact to programming tasks, then technical exploration, code generation, or manual coding, with prompt templates and reused chat histories to lower context-construction effort) and a content model listing the context categories that make generated code usable. The authors conclude that any scientific approach claiming to automate a requirements-centric software engineering task should explain how it represents requirements and how it handles the manual decomposition and enrichment this study observed.","pith_inferences":["The authors leave implicit that if decomposition is the bottleneck, then measuring and supporting the decomposition step—for example, with semi-automated task-splitting tools—may matter more than further improving code generation quality.","A testable extension would be a log-based field study that records actual prompts, accepted completions, and code diffs; unlike retrospective interviews, such logs could quantify how much of the effort is prompt construction versus post-generation code adjustment.","The content model suggests a structured prompt schema; one could test whether systematically supplying the five context categories reduces integration failures, such as compilation errors or failing tests, compared with unstructured prompts.","The theory predicts the programming-task intermediate step will shrink but not disappear as models gain repository grounding, because the abstraction gap is partly project-specific knowledge the model cannot infer from a requirement alone."],"forward_implications":["Benchmarks that evaluate code generation from short coding descriptions should not be treated as evidence about requirements-to-code practice; the paper draws a sharp line between the two.","Fully automated software engineering from raw requirements remains distant for complex software, because decomposition and context selection require requirements and software engineering expertise.","Prompts are currently treated as transient artifacts rather than documented, reviewed engineering artifacts, so traceability and accountability of prompt content become pressing concerns.","Tooling that supports context reuse—prompt templates, chat histories, and pre-filled agent contexts—directly addresses the effort bottleneck the practitioners described.","The optimal granularity of programming tasks is an open research question, and future studies should compare the granularity of benchmark inputs with the granularity of tasks derived from real requirements."],"supporting_citations":[{"why":"Supplies the standard definition of requirements, design, and implementation that the paper uses as its baseline process.","marker":"[1]"},{"why":"Provides the premise that generative AI adoption depends on integration with existing software development workflows.","marker":"[8]"},{"why":"Exemplifies the benchmark line that treats programming contest descriptions as requirements, motivating the paper's gap.","marker":"[10]"},{"why":"Exemplifies requirements-clarification research that also uses coding-task inputs rather than real requirements.","marker":"[11]"},{"why":"Supplies the prior interaction-pattern findings (accelerate and explore) that the paper extends into a full process model.","marker":"[32]"},{"why":"Supplies the prior hypothesis that code queries need contextual information, which the paper's content model confirms and extends.","marker":"[37]"},{"why":"Supplies prior qualitative evidence that developers constantly monitor high-level requirements during implementation.","marker":"[38]"}],"fun_headline_variants":["LLMs can't skip requirements: devs still decompose","Pasting requirements into LLMs yields unusable code","Study: Requirements need manual breakdown for LLM code","Devs must decompose requirements before prompting LLMs","LLM code generation still hinges on requirements engineering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The theory stands on what 18 practitioners said about their own workflows in interviews, rather than on logs or direct observation of their prompts, keystrokes, or code diffs; if those accounts are inaccurate, the paper describes what developers say they do, not what they actually do.","fun_headline_variants_meta":{"raw":{"variants":["LLMs can't skip requirements: devs still decompose","Pasting requirements into LLMs yields unusable code","Study: Requirements need manual breakdown for LLM code","Devs must decompose requirements before prompting LLMs","LLM code generation still hinges on requirements engineering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000616,"raw_usage":{"total_tokens":2852,"prompt_tokens":926,"completion_tokens":1926,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":1851}},"tokens_in":542,"tokens_out":1926,"duration_ms":13483,"temperature":1.0,"reasoning_tokens":1851,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:37:40.134840+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A log-based field study recording the actual prompts, accepted completions, and code diffs of developers over several weeks would settle the claim: if raw requirement texts are regularly pasted verbatim into prompts and the resulting code is integrated without any prior decomposition into programming tasks, the central claim is contradicted; if every code-generating prompt is preceded by a hand-written task breakdown, the theory is supported.","supporting_citations":[{"cited_title":"Sommerville, Software Engineering, 10th ed","cited_arxiv_id":null,"evidence_quote":"Supplies the standard definition of requirements, design, and implementation that the paper uses as its baseline process."},{"cited_title":"Navigating the complexity of generative AI adoption in software engineering,","cited_arxiv_id":null,"evidence_quote":"Provides the premise that generative AI adoption depends on integration with existing software development workflows."},{"cited_title":"Deep learning based program generation from requirements text: Are we there yet?","cited_arxiv_id":null,"evidence_quote":"Exemplifies the benchmark line that treats programming contest descriptions as requirements, motivating the paper's gap."},{"cited_title":"Clarifygpt: A framework for enhancing llm-based code generation via requirements clarification,","cited_arxiv_id":null,"evidence_quote":"Exemplifies requirements-clarification research that also uses coding-task inputs rather than real requirements."},{"cited_title":"Grounded copilot: How programmers interact with code-generating models,","cited_arxiv_id":null,"evidence_quote":"Supplies the prior interaction-pattern findings (accelerate and explore) that the paper extends into a full process model."},{"cited_title":"In-IDE code generation from natural language: Promise and challenges,","cited_arxiv_id":null,"evidence_quote":"Supplies the prior hypothesis that code queries need contextual information, which the paper's content model confirms and extends."},{"cited_title":"A qualitative study on the implementation design decisions of developers,","cited_arxiv_id":null,"evidence_quote":"Supplies prior qualitative evidence that developers constantly monitor high-level requirements during implementation."}],"review_version":1}