{"id":"526a5e14-db25-4e2d-bb61-74755b8b156e","arxiv_id":"2502.04256","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A proposal for using LLMs to classify and score system requirements under INCOSE criteria, with no results reported.","lead":"This paper outlines a plan to use large language models like GPT-4 and Claude to automatically check system requirements against INCOSE quality rules and sort them into functional and non-functional categories. It is a work-in-progress proposal with no experimental results yet, so its value lies in the design of future comparisons with human engineers.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper reports no experimental results, so its central claim is unverifiable; even granting that, the expert benchmark lacks an inter-rater reliability check, making any future AI-vs-expert comparison indeterminate.","rationale":"The reader's verdict is UNVERDICTED, and I agree: the paper contains no experimental results, data, or analysis, so there is no central claim that can be verified or falsified from the manuscript. The reader's weakest assumption—that the 107-requirement DR TOOL dataset is insufficient to support a general capability claim—is a real concern, and I partially agree with it. My stress-test identifies a different, more fundamental load-bearing condition: the expert benchmark itself must be reliable for the proposed AI-vs-expert comparison to mean anything. Section II-B specifies Cohen's kappa between AI and experts but never specifies expert-expert agreement. Because INCOSE's criteria are interpretive, a blind expert panel could disagree substantially among themselves, in which case the 'ground truth' is not a fixed standard. This does not change the verdict: the manuscript remains unverdictable because it is a work-in-progress proposal with no reported results. It does, however, sharpen the design flaw that should be fixed before the planned experiment can validate any claim. My recommendation is therefore UNCHANGED, consistent with the reader's high-confidence UNVERDICTED verdict.","tokens_in":4809,"tokens_out":1913,"duration_ms":22438,"concrete_test":"Run the planned experiment with at least three independent expert raters labeling all 107 DR TOOL requirements against the seven INCOSE criteria, and compute pairwise Cohen's kappa among experts before comparing any LLM output. If expert-expert kappa is below 0.6, or is not substantially higher than the AI-expert kappa, then the expert benchmark cannot serve as a validation standard and the paper's proposed comparison cannot support its claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that generative AI can automate requirement analysis and that this can be validated by comparing AI outputs with expert judgments—depends on the expert judgments forming a stable, reproducible ground truth. That condition is not established and, as designed, cannot be established from the described protocol. Section II-B (step 3) states that a panel of experienced systems engineers will independently classify the dataset, and step 4 compares AI results with expert classifications using Cohen's kappa. However, no expert-expert agreement measure is specified. The INCOSE criteria in Section II-A ('essential,' 'independent,' 'complete,' 'singular,' etc.) are inherently interpretive; two qualified engineers can reasonably disagree on whether a requirement is 'unambiguous' or 'complete.' If expert-expert kappa is low, then low AI-expert kappa would not indicate AI failure—it would indicate an unstable reference standard. Conversely, if expert-expert kappa is high only because most requirements are trivially easy, the comparison says little about AI's reliability on genuinely difficult requirements. Additionally, the dataset is 107 requirements from a single RFID-based medical-equipment-tracking project (Section II-C), so any quantitative result is domain-specific; generalizing to systems engineering at large would require cross-domain validation. The paper is explicitly a proposal—no results are reported—so the central claim is not yet supported by any evidence. The abstract uses present-tense statements such as 'The AI does not just classify requirements but also explains why some do not meet the standards,' which overstates what the paper actually contains. The appropriate status is UNVERDICTED, not because the idea is wrong, but because no experiment or analysis is present to evaluate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a work-in-progress proposal for using generative AI models (GPT-4, Claude Sonnet, Llama) to automate the classification and quality assessment of systems engineering requirements according to INCOSE's 'good requirement' criteria, to distinguish functional from non-functional requirements, and to generate test specifications. The proposed method involves a dataset of 107 requirements from a single RFID-based medical-equipment tracking project (DR TOOL), independent blind classification by experienced systems engineers, and comparison of AI and expert classifications using Cohen's kappa, followed by iterative refinement of the AI models. The paper reports no experimental results, no data, and no computed statistics; all evaluation steps are described as planned future work.","tokens_in":5038,"tokens_out":3973,"duration_ms":41789,"significance":"If the proposed evaluation were carried out and the results were positive, the work would address a practical gap in requirements engineering: automated, explainable classification of requirement quality against an established standard such as INCOSE's. The study's strengths are its use of a real-world requirements dataset, the explicit choice of three state-of-the-art LLMs, and a planned blind expert benchmark with a statistical agreement measure. However, the paper currently provides no evidence for its central claim, only a research design. The discussion sections (IV) and conclusion (VI) go beyond what the manuscript supports by describing the study as 'demonstrating' a practical transition that has not yet occurred. The planned benchmark also needs a critical addition—inter-rater reliability among the expert panel—without which any future AI-expert kappa is not interpretable.","major_comments":[{"comment":"The paper reports no experimental results. The central claim that generative AI can automate requirement analysis and that this can be validated by comparing AI outputs with expert judgments is not supported by any data, model output, or computed agreement statistic. All five steps in the research design are described in future/proposed terms, and Section III itself states only 'expected' contributions. For a full journal paper this is a blocking issue: the manuscript must either present completed experiments with at least one AI model and a comparison to expert classifications, or be explicitly reframed as a research proposal/vision paper with the scope limited to presenting a protocol.","section":"Section II-B (Steps 1–5) and Section III"},{"comment":"The validation protocol does not specify any measure of expert-expert agreement. Step 3 says a panel of experienced systems engineers will independently classify the dataset, and Step 4 compares AI classifications with expert classifications using Cohen's kappa. Cohen's kappa is defined for two raters, not for a panel, and more importantly, without an expert-expert agreement measure (e.g., Fleiss' kappa or pairwise Cohen's kappa) the expert benchmark is not established as a stable ground truth. INCOSE criteria such as 'unambiguous' and 'complete' are inherently interpretive; if experts disagree among themselves, a low AI-expert kappa does not indicate AI failure, and a high AI-expert kappa could be misleading. Please add an explicit expert-expert agreement analysis and a stratification of the dataset by expert difficulty.","section":"Section II-B, Steps 3 and 4"},{"comment":"The dataset consists of 107 requirements from a single RFID-based medical equipment tracking project (DR TOOL). The paper generalizes its claims to 'AI-powered engineering' without justifying the representativeness of this dataset. A positive or negative result on this small, domain-specific corpus would not support conclusions about systems engineering in general. Either the dataset needs to be expanded to multiple projects/domains, or the paper must explicitly restrict its claims to this domain and present the work as a case study rather than a general evaluation.","section":"Section II-C"},{"comment":"The protocol does not operationalize how each INCOSE criterion is assessed. For instance, is each requirement rated as pass/fail or on a Likert scale? What prompt templates and few-shot examples are used for each LLM? How is a 'hallucination' identified and recorded? Without these details, the proposed comparison is not reproducible and the future AI-expert comparison cannot be interpreted. This load-bearing detail should be specified before the experiment is run.","section":"Section II-A and II-B"}],"minor_comments":[{"comment":"The verb tense shifts from the future tense of Step 2 ('will analyze') to the past tense in Section II-C ('The models were applied'), creating ambiguity about whether the experiment has already been conducted; please standardize the tense.","section":"Section II-C, first sentence"},{"comment":"The captions for Fig. 1 and Fig. 2 are present, but the figures themselves are not included in the text; please provide the actual diagrams or remove the references.","section":"Figures 1 and 2"},{"comment":"Reference [13] lacks a year and appears to be a course report; please provide a permanent citation or remove it. Reference [15] is a preliminary review of ChatGPT rather than a technical description of GPT-4; citing the GPT-4 technical report or model card would be more appropriate.","section":"References"},{"comment":"The paper states that the study focuses on seven out of nine INCOSE parameters but does not name the two omitted parameters or explain the selection; please add a brief justification.","section":"Section II-A"},{"comment":"The phrase 'this study demonstrates how AI can effectively transition from a research tool to a practical asset' overclaims relative to the absence of results; please soften to 'proposes' or 'aims to demonstrate.'","section":"Section IV-A"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest in labeling itself as work in progress, and the proposed experimental design has some merit. However, as a journal submission it does not yet contain a research contribution; it is a research proposal. The missing expert-expert agreement measure is a design flaw that should be fixed before the study is run. The reliance on the authors' own DR TOOL dataset is not circular, but it is a narrow single-case basis for broad claims. If this is intended for a workshop or work-in-progress track, the lack of results may be acceptable; for a journal track, I would require completed experiments or a clear reframing as a protocol paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a work-in-progress proposal, not a paper with results. The one thing to know up front: nothing is evaluated or reported, so every claim about AI performance is promissory. That is not a fatal flaw given the explicit WIP framing, but it means the paper should not be judged as a completed study.\n\nWhat is actually useful: the authors propose combining three current LLMs (GPT-4, Claude, Llama) with INCOSE's seven \"good requirement\" criteria, and they plan to validate the approach on a real dataset from their prior DR TOOL project. That dataset, 107 requirements from an RFID-based medical equipment tracking system, gives the plan some concreteness. The design is straightforward: AI classifies, a panel of experts classifies blind, and agreement is measured with Cohen's kappa. The paper is clearly written, cites prior ML-based requirements classification work (Tamai and Anzai, Cheligeer et al.), and is transparent that this is an expected contribution rather than a completed one.\n\nThe soft spots are the usual ones for a proposal, plus one the authors have not yet addressed. First, there is no data and no analysis. For a WIP that is acceptable, but the abstract overstates the case by using present tense (\"The AI does not just classify requirements but also explains why...\") when no AI has run yet. Second, the protocol lacks an expert-expert agreement measure. Cohen's kappa between AI and experts is only interpretable if the experts agree with each other; without an inter-rater reliability check on the expert panel, a low AI-expert kappa could mean the AI is wrong or the reference standard is unstable. That is a fixable design gap, but it should be added before data collection. Third, the dataset is a single project in one domain, so any quantitative result will be domain-specific. That is fine for a pilot, but the paper should not imply general conclusions without cross-domain validation.\n\nThe citation pattern looks fine. Using their own prior DR TOOL dataset is transparent and not a problem. The self-citation in [11] is to the source of the dataset, which is appropriate.\n\nWho is this for? Someone running a WIP/emerging-results session or a graduate seminar on applying LLMs to requirements engineering could get value from the framing. I would not cite it in my own work yet, because there is no result to cite. If the authors bring this back with actual experimental data and the expert-expert kappa added, it could become a legitimate empirical paper. As it stands, I would not send it to a full research track, but it is appropriate for a workshop or a WIP track that explicitly solicits proposals.","headline":"A clear, honest WIP proposal with no results yet; the missing expert-expert agreement check is a real design gap, but the plan is sensible for a workshop track.","tokens_in":5595,"tokens_out":2185,"would_cite":false,"duration_ms":24124,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes that generative AI can classify system requirements against the systems-engineering 'good requirement' criteria as reliably as experienced engineers, and lays out a blind expert-comparison experiment to test that claim.","keywords":["generative AI","systems engineering","requirements engineering","good requirement criteria","requirements classification","functional and non-functional requirements","Cohen's kappa","AI evaluation"],"falsifier":"Run the described experiment: have GPT-4, Claude, and Llama classify all 107 DR TOOL requirements against the seven 'good requirement' criteria and the functional/non-functional split, have a blind panel of experienced systems engineers do the same, and compute Cohen's kappa for each model-expert pair; if all three kappas are at or near zero, the claim that AI can match a systems engineer's judgment fails.","tokens_in":4619,"feed_emoji":"🤖","tokens_out":8491,"duration_ms":77916,"temperature":0.7,"pith_summary":"This work-in-progress paper aims to establish that generative AI can analyze system requirements against the systems-engineering 'good requirement' criteria (essential, independent, unambiguous, complete, singular, feasible, verifiable) and classify each requirement as functional or non-functional at a level that agrees with experienced systems engineers. The authors argue that if this holds, an AI tool could flag poorly written requirements with explanations, generate test specifications from the classifications, and serve as a teaching aid for engineering students. To test the claim, they propose running three large language models (GPT-4, Claude, and Llama) on 107 real requirements from a medical-equipment tracking system and comparing the models' outputs with a blind panel of human experts using Cohen's kappa. The paper reports no results yet; its current contribution is the evaluation design and the argument that measuring human-AI agreement is the correct way to judge whether AI matches or only assists an engineer's judgment.","feed_headline":"Three LLMs face engineers in a blind requirements test","feed_subtitle":"A proposed experiment compares GPT-4, Claude, and Llama with expert judgment on seven quality criteria.","key_machinery":"The load-bearing mechanism is a two-rater comparison design. Three large language models (GPT-4, Claude, and Llama) and a panel of experienced systems engineers independently classify the same 107 requirements, 31 stakeholder and 76 optimized system requirements from the DR TOOL medical-equipment tracking project, against the seven 'good requirement' criteria and the functional/non-functional split; Cohen's kappa measures agreement between each model and the human benchmark, and discrepancies feed an iterative refinement loop.","core_discovery":"On its own terms, the paper's central claim is that large language models can apply the seven 'good requirement' criteria and the functional/non-functional distinction to real engineering requirements, producing classifications and justifications that agree with experienced systems engineers. The authors expect some agreement but predict that AI will show hallucinations and contextual misunderstandings that engineers catch intuitively; the proposed experiment is designed to measure exactly where agreement breaks down. The intended discovery, if the experiment goes as hoped, is that a generative-AI workflow can serve as a reliable screening and teaching tool for requirements quality, with human engineers remaining the benchmark.","pith_inferences":["Beyond the paper: the same protocol could be applied to requirements from defense, transportation, or software projects to test whether performance on the medical-equipment dataset generalizes.","A future extension could separate label agreement from explanation agreement, since a model and an expert can choose the same criterion and still disagree about the underlying defect.","If the method proves reliable, it could scale to large requirements documents where manual expert review is expensive, turning the comparison into a practical audit tool.","The educational claim could be tested by measuring whether students who train on AI-flagged requirements write better requirements than a control group."],"forward_implications":["If the comparison shows strong agreement, systems engineers could delegate the first pass of requirements-quality screening to an AI and spend their time on the flagged items.","If AI can reliably separate functional from non-functional requirements, test specifications can be generated earlier in the design process, before detailed system design is locked.","If the AI's justifications match expert reasoning, the same tool can be used in engineering education to show students why a requirement fails a quality criterion.","If agreement is weak, the refinement loop is intended to show which criteria and which requirement types need model tuning or human oversight.","Organizations could use the measured agreement to decide where AI assistance is safe and where a human expert must stay in the loop."],"supporting_citations":[{"why":"Supplies the seven 'good requirement' criteria that the AI models are asked to apply.","marker":"[8]"},{"why":"Provides the Quality Requirements Mining and Classification Process that the proposed AI approach adapts for functional/non-functional classification.","marker":"[7]"},{"why":"Provides the DR TOOL project's 107 requirements, the real-world dataset on which the evaluation is run.","marker":"[11]"},{"why":"Identifies GPT-4 as one of the three large language models tested.","marker":"[15]"},{"why":"Identifies Claude as one of the three large language models tested.","marker":"[16]"},{"why":"Identifies Llama as one of the three large language models tested.","marker":"[17]"}],"fun_headline_variants":["LLMs vs engineers: blind test on requirement quality","Can AI judge good requirements? A blind test vs experts","Three LLMs put to the test on requirement writing","Blind test: Do LLMs match engineers on requirements?","GPT-4, Claude, Llama vs engineers on requirement quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 107 requirements from one medical-equipment tracking project are treated as representative enough of system requirements in general that agreement with experts on this dataset would tell us something general about AI's classification ability.","fun_headline_variants_meta":{"raw":{"variants":["LLMs vs engineers: blind test on requirement quality","Can AI judge good requirements? A blind test vs experts","Three LLMs put to the test on requirement writing","Blind test: Do LLMs match engineers on requirements?","GPT-4, Claude, Llama vs engineers on requirement quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1148,"prompt_tokens":784,"completion_tokens":364,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":400,"completion_tokens_details":{"reasoning_tokens":282}},"tokens_in":400,"tokens_out":364,"duration_ms":4805,"temperature":1.0,"reasoning_tokens":282,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T22:57:39.818594+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the described experiment: have GPT-4, Claude, and Llama classify all 107 DR TOOL requirements against the seven 'good requirement' criteria and the functional/non-functional split, have a blind panel of experienced systems engineers do the same, and compute Cohen's kappa for each model-expert pair; if all three kappas are at or near zero, the claim that AI can match a systems engineer's judgment fails.","supporting_citations":[{"cited_title":"John Wiley and Sons, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the seven 'good requirement' criteria that the AI models are asked to apply."},{"cited_title":"Tamai and T","cited_arxiv_id":null,"evidence_quote":"Provides the Quality Requirements Mining and Classification Process that the proposed AI approach adapts for functional/non-functional classification."},{"cited_title":"Management and Detection System for Medical Surgical Equipment","cited_arxiv_id":"2211.02351","evidence_quote":"Provides the DR TOOL project's 107 requirements, the real-world dataset on which the evaluation is run."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4 as one of the three large language models tested."},{"cited_title":"2023, revised Jun","cited_arxiv_id":null,"evidence_quote":"Identifies Claude as one of the three large language models tested."}],"review_version":1}