{"id":"2f1c4060-fb35-40b3-b092-bbc56c7ede29","arxiv_id":"2506.11548","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A proposal for augmenting LLM generality in non-code software engineering tasks, with no reported experimental evidence.","lead":"This preprint is a research proposal, not a completed study. It outlines plans to train and evaluate LLMs on non-code software engineering tasks, build knowledge benchmarks, and detect hallucinations, but reports no experimental results.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's only empirical support—Section IV's claim of 'increase in performance in most cases' across 17 tasks—is unreported, unreproducible, and undefined, leaving the central feasibility thesis unsubstantiated.","rationale":"The reader correctly identifies the paper as a research proposal rather than a completed study, and the lack of datasets, code, and metrics makes it unverdictable. The reader's weakest_assumption concerns the construct validity of the proposed foundational-knowledge benchmark; this is a legitimate forward-looking concern. However, the more immediate load-bearing issue is the paper's only reported result: Section IV claims that fine-tuning improved performance on 17 non-code tasks 'in most cases,' yet supplies no quantitative detail. This is the empirical basis for the central thesis that smaller, domain-specific models can compete with large general-purpose models. Without a report of task definitions, baselines, metrics, and effect sizes, the claim is unfalsifiable and the paper's motivation lacks support. The proposed test—requiring the per-task results and a paired significance analysis—would settle whether the claim is substantive or vacuous. Because the paper is explicitly a proposal with 'expected contributions,' the correct status remains unverdictable rather than accepted or rejected; our concern strengthens the case for UNVERDICTED but does not change the reader's verdict.","tokens_in":5094,"tokens_out":5713,"duration_ms":50688,"concrete_test":"Request the author's per-task results table for the 17 fine-tuned tasks (or the forthcoming first paper). For each task, record the base-model performance, the fine-tuned performance, and the evaluation metric. Then run a paired non-parametric test (e.g., Wilcoxon signed-rank) on the 17 deltas, with a correction for multiple comparisons if needed. If the table is not provided, or if the test fails to reject the null of no improvement at α=0.05 (or the effect sizes are negligible), then the Section IV claim is unsupported and the paper's central thesis is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section I argues that the default practice of using the largest general-purpose LLMs is 'far from optimal' because 'similar or better performance could potentially be achieved with smaller LLMs' trained on domain-specific datasets. The only evidence for this feasibility is Section IV: models fine-tuned on labeled datasets for 17 non-code tasks 'show an increase in performance in most cases.' No task list, dataset description, model architecture, size, baseline, metric, effect size, or significance test is given. Consequently the phrase 'increase in performance in most cases' is vacuous: it could mean 9 of 17 tasks with negligible gains or 17 of 17 with large gains; it could be relative to an untuned checkpoint or to a general-purpose frontier model; the metric could be accuracy, F1, or something else. The promised replication kit and the forthcoming 'first paper' are not provided, so a reader cannot check any of these. Since the entire proposal's motivation depends on this empirical claim, the concern is not stylistic—without a falsifiable report, the paper does not provide scientific evidence for its central thesis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is an extended abstract or research proposal for software engineering LLM research. It argues that the common practice of using the largest general-purpose LLMs for every SE task is \"far from optimal\" and suggests that smaller LLMs with different architectures trained on domain-specific data could achieve similar or better performance on non-code tasks (Section I). The paper poses four research questions, plans four contributions (training LLMs, building foundational-knowledge benchmarks, generating nonsensical SE statements, and detecting hallucinations), and asserts in Section IV that BERT/RoBERTa/GPT-2 models fine-tuned on 17 non-code tasks showed \"an increase in performance in most cases\" without reporting any data, baselines, metrics, or experimental setup.","tokens_in":5411,"tokens_out":3580,"duration_ms":35577,"significance":"If validated, the core thesis would be practically important: it would challenge the default practice of scaling up general-purpose models and would motivate smaller, domain-specific LLMs for non-code SE tasks. The planned benchmarks on foundational SE knowledge and the hallucination-detection methodology could also become useful community resources. The paper appropriately states its intended replication kits and analysis dimensions. However, all of this significance is conditional: the only empirical evidence offered is a single unreported sentence, and the planned experiments are under-specified. As it stands, the manuscript does not yet demonstrate its central claims.","major_comments":[{"comment":"The sentence \"these models have also been fine-tuned on labeled datasets for 17 non-code tasks, showing an increase in performance in most cases\" is the sole empirical support for the paper's central feasibility claim. It gives no task list, dataset descriptions, model architecture or size, baseline definition, evaluation metric, effect size, or statistical test. This statement is unfalsifiable as written and cannot support the thesis that small domain-specific LLMs can match or exceed large general-purpose models.","section":"Section IV (Initial Results)"},{"comment":"The central argument is that similar or better performance could be achieved with smaller LLMs than with the large general-purpose models (GPT-4o, Claude 3.5, Llama 3.2) cited in Section I. However, the planned experimental design in Section III-A only compares the trained LLMs against XGBoost and FastText baselines; it does not include a comparison against those large general-purpose models. Without such a comparison, neither the reported initial results nor the proposed study can test the central claim.","section":"Sections I and III-A"},{"comment":"The benchmark for foundational SE knowledge is defined by terminology extracted from four standards bodies (ISO/IEC/IEEE 24765, ISTQB, IREB, iSAQB) plus object-oriented design scenarios with UML diagrams. The paper gives no evidence or argument that this set represents what practitioners mean by foundational SE knowledge; without a validation step such as expert agreement or comparison with an established SE curriculum, results on this benchmark cannot be claimed to measure LLM foundational knowledge in a general way.","section":"Section III-B"},{"comment":"The hallucination-detection plan relies on \"systematically modified nonsensical statements,\" but no concrete modification procedure is described, and the planned Kolmogorov-Smirnov tests and QQ-plots are reported without stating which distributions are compared, how many samples are involved, or what threshold would count as successful detection. This under-specification makes it impossible to assess whether the proposed method could answer RQ4.","section":"Sections III-C and III-D"}],"minor_comments":[{"comment":"Section III-A states that new LLMs will be pre-trained on \"over 200 GB of texts,\" while Section IV says the models have so far been trained on \"more than 23 GB of textual data\"; the relationship between these numbers should be clarified.","section":"Sections III-A and IV"},{"comment":"Reference [9] gives the incomplete arXiv identifier \"arXiv:2107.0337\"; the correct identifier appears to be 2107.03374.","section":"References"},{"comment":"The text refers to a figure whose image file is only mentioned in the arXiv metadata; no figure is described or cited in the body of the manuscript, so the pointer should either be replaced by an actual figure or removed.","section":"General"},{"comment":"The term \"foundational knowledge\" is linked to Bloom's taxonomy [16] but is not defined beyond the pointer to factual and conceptual knowledge; one sentence of explanation would improve precision.","section":"Section I"}],"recommendation":"reject","confidential_remarks":"The manuscript is more of a research proposal or workshop extended abstract than a full research paper. Its only empirical claim (17 tasks with improved performance) is unreported, and the planned methodology is too underspecified to support the advertised conclusions. The author appears to have the right research directions in mind, and the future first paper, if it reports full experimental detail, could be a strong basis for a later submission. For now, the manuscript does not meet the evidence bar of a research journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is not a completed research contribution: it is a proposal for future work, with one unreferenced sentence in Section IV claiming that fine-tuned models show 'an increase in performance in most cases' across 17 non-code tasks. No task list, datasets, model sizes, baselines, metrics, or error bars are given. Second, despite that, the proposal itself is thoughtful and points at a genuine gap: most LLM-for-SE work targets code generation, and non-code tasks like requirements, design, and foundational terminology are underexplored.\n\nWhat the paper does well is scope the problem and lay out a sensible plan. The three research questions are concrete, and the expected contributions are specific: domain-adaptive pre-training on SE text, benchmarks built from standards-body glossaries and UML design scenarios, and a hallucination-detection benchmark with planned statistical tests (QQ-plots, Kolmogorov-Smirnov). The inclusion of XGBoost and FastText baselines is a good instinct, and the promise of replication kits is exactly the right norm—though it is only a promise at this point. The citation pattern is unremarkable and appropriate, and the writing is honest about what is expected versus what exists.\n\nThe soft spot is not a minor one: the central feasibility thesis depends entirely on the unreported Section IV result. 'Increase in performance in most cases' could mean anything, and because no comparison or metric is specified, the sentence is vacuous. The abstract's 'promising' reinforces the problem. The paper also does not defend the foundational-knowledge proxy: terminology from four SDOs and a set of UML scenarios is a reasonable start, but it is not argued to represent what practitioners actually mean by foundational SE knowledge. That is a secondary concern, not the main flaw. The scale discrepancy between the planned 200+ GB of training data and the reported 23 GB is noted but not explained; that matters for interpreting the initial result, if one existed.\n\nWho is this for? Someone supervising or evaluating early-stage PhD work might find it a useful model of how to structure a research plan. But as an arXiv paper in cs.SE, it has no evidence to review. I would not cite it, and I would not bring it to reading group as a research result.\n\nMy recommendation: do not send this to peer review as-is. A serious editor should desk-reject it or redirect it to a workshop or proposal venue. The author should be told to come back when the 17-task evaluation is reported with actual metrics, baselines, task definitions, and statistical tests. The underlying plan is sound enough that a revised paper with real data would deserve a full review.","headline":"A clear, honest research proposal with zero reported evidence; Section IV's 'increase in performance in most cases' is unreported and unreproducible, so the paper currently functions as a plan, not a result.","tokens_in":5790,"tokens_out":1788,"would_cite":false,"duration_ms":19238,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Smaller domain-trained LLMs may match or beat general-purpose giants, this research proposal argues.","keywords":["large language models","software engineering","non-code tasks","foundational knowledge","benchmarking","hallucination detection","domain-specific pretraining","fine-tuning"],"falsifier":"Train the proposed smaller domain-specific models (RoBERTa, GPT-2, T5) on the SE corpus, fine-tune them on the 17 non-code tasks and the foundational-knowledge benchmarks, and compare them head-to-head with a large general-purpose model like GPT-4o under zero-shot and fine-tuned conditions. If the smaller models do not match or exceed the large model on the majority of tasks and benchmarks, the paper's central thesis is refuted. Similarly, if the proposed hallucination detector cannot distinguish nonsensical SE statements from original ones at better than chance or a simple keyword baseline, that sub-claim fails.","tokens_in":4884,"feed_emoji":"🤖","tokens_out":6624,"duration_ms":54632,"temperature":0.7,"pith_summary":"This research proposal argues that the default practice of reaching for the largest general-purpose LLM for every software-engineering task is far from optimal: for non-code tasks, similar or better performance could come from smaller models trained on domain-specific data. To test this, it designs benchmarks for foundational SE knowledge drawn from standards terminology and object-oriented design scenarios, together with a hallucination-detection benchmark built from deliberately nonsensical SE statements. The author's plan is to pre-train and fine-tune architectures such as RoBERTa, GPT-2, and T5 on over 200 GB of SE text and evaluate them on these benchmarks plus 17 non-code tasks. Initial, unreported results claim performance increases in most cases on those tasks. If the thesis holds, it would provide a cheaper, more specialized alternative to scaling up general models and a way to check SE-related hallucinations.","feed_headline":"Smaller domain-trained LLMs may match or beat general-purpose giants","feed_subtitle":"A research agenda tests specialized models on non-code software-engineering tasks and SE hallucination detection.","key_machinery":"The load-bearing machinery is a three-part experimental apparatus. First, domain-adapted LLMs: RoBERTa, GPT-2, and T5 pre-trained from scratch and fine-tuned from public checkpoints on SE text (GitHub, Stack Overflow, JIRA, ArXiv), then supervised fine-tuned on 17 non-code tasks. Second, a foundational-knowledge benchmark: terminology extracted from four standards bodies (ISO/IEC/IEEE 24765, ISTQB, IREB, iSAQB) and object-oriented design scenarios with UML diagrams, used for zero-shot discrimination tasks and generation evaluated with BLEU and qualitative error analysis. Third, a hallucination-detection benchmark: systematically corrupted SE statements that LLMs are prompted to elaborate, with detection via zero-shot classification and label-probability distributions compared using QQ-plots and Kolmogorov-Smirnov tests. These components carry the argument from the thesis (smaller domain-specific models can suffice) to measurable outcomes.","core_discovery":"The central claim is that generality and performance for non-code software-engineering tasks are best augmented not by scaling model size but by training smaller, architecturally different models on domain-specific datasets. The paper treats this as a testable hypothesis and proposes a multi-part research plan: establish benchmarks for foundational SE knowledge from ISO/IEC/IEEE 24765, ISTQB, IREB, and iSAQB terminology plus UML design scenarios; generate nonsensical SE statements to probe hallucination; and use zero-shot classification with probability comparisons (QQ-plots, Kolmogorov-Smirnov) to detect those statements. The author reports that BERT, RoBERTa, and GPT-2 models have already been pre-trained and fine-tuned on over 23 GB of textual data, with fine-tuning on 17 non-code tasks showing an increase in performance in most cases, though the evidence is not presented in this manuscript. The intended outcome is a suite of models, benchmarks, and detection methods that together support non-code SE practice.","pith_inferences":["Extension: the same benchmark-and-hallucination design could be lifted to other professional domains with formal standards vocabularies, such as medicine, law, or cybersecurity, where nonsensical-statement detection is equally valuable.","Implicit: the paper's notion of generality is per-task competence, so a collection of smaller specialized models could collectively match a single general model, suggesting a system-level rather than model-level definition of generality.","Testable: one could extend the evaluation to measure qualitative error types (fabrication versus misunderstanding) on the foundational-knowledge benchmark to see whether small domain-trained models fail differently from large general models.","The foundational-knowledge proxy is drawn from only four standards bodies, so a sensitivity analysis with additional sources would clarify how much the benchmark results depend on that particular choice."],"forward_implications":["If the thesis is correct, organizations could deploy smaller, cheaper SE assistants that match large general models on non-code tasks such as requirements analysis and design.","The proposed benchmarks would give the SE community a standardized measure of foundational knowledge that does not rely on code-generation tests.","The hallucination-detection method could be integrated into LLM-based SE tools to flag statements that are plausible but false, increasing trust.","Fine-tuning on labeled non-code datasets would become a demonstrated recipe, encouraging more labeled-data collection in SE.","The field could shift from a one-model-fits-all philosophy to a portfolio of specialized models."],"supporting_citations":[{"why":"Supplies the scaling-law premise that the paper argues is not the only route to performance.","marker":"[7]"},{"why":"Documents the coding-task imbalance in LLM-for-SE research, defining the gap this proposal targets.","marker":"[13]"},{"why":"Provides the definition of foundational knowledge (factual and conceptual) that motivates the benchmark design.","marker":"[16]"},{"why":"Provides the RoBERTa architecture used for domain-specific pre-training and fine-tuning.","marker":"[25]"},{"why":"Provides the GPT-2 architecture used for domain-specific pre-training and fine-tuning.","marker":"[26]"},{"why":"One of four standards sources from which benchmark terminology is extracted.","marker":"[27]"},{"why":"ISTQB glossary source for benchmark terminology.","marker":"[28]"},{"why":"IREB glossary source for benchmark terminology.","marker":"[29]"},{"why":"iSAQB source for benchmark terminology.","marker":"[30]"},{"why":"Motivates the hallucination-detection component by establishing hallucination as an innate limitation of LLMs.","marker":"[19]"}],"fun_headline_variants":["Smaller LLMs may rival giants on non-code SE tasks","Domain-trained LLMs show potential for non-code software engineering","Can small LLMs outdo big ones in SE? Initial results suggest yes","Specialized LLMs: a path to better non-code SE performance","New benchmarks aimed at boosting LLM generality in SE"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmarks assume that foundational software-engineering knowledge is adequately captured by terminology from four standards bodies (ISO/IEC/IEEE 24765, ISTQB, IREB, iSAQB) plus a set of UML design scenarios; if practitioners' foundational knowledge differs, the benchmark results will not generalize.","fun_headline_variants_meta":{"raw":{"variants":["Smaller LLMs may rival giants on non-code SE tasks","Domain-trained LLMs show potential for non-code software engineering","Can small LLMs outdo big ones in SE? Initial results suggest yes","Specialized LLMs: a path to better non-code SE performance","New benchmarks aimed at boosting LLM generality in SE"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000341,"raw_usage":{"total_tokens":1843,"prompt_tokens":874,"completion_tokens":969,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":882}},"tokens_in":490,"tokens_out":969,"duration_ms":9124,"temperature":1.0,"reasoning_tokens":882,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:00.492333+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the proposed smaller domain-specific models (RoBERTa, GPT-2, T5) on the SE corpus, fine-tune them on the 17 non-code tasks and the foundational-knowledge benchmarks, and compare them head-to-head with a large general-purpose model like GPT-4o under zero-shot and fine-tuned conditions. If the smaller models do not match or exceed the large model on the majority of tasks and benchmarks, the paper's central thesis is refuted. Similarly, if the proposed hallucination detector cannot distinguish nonsensical SE statements from original ones at better than chance or a simple keyword baseline, that sub-claim fails.","supporting_citations":[{"cited_title":"A revision of bloom’s taxonomy: An ove rview,","cited_arxiv_id":null,"evidence_quote":"Provides the definition of foundational knowledge (factual and conceptual) that motivates the benchmark design."},{"cited_title":"ISO/IEC/IEEE 24 765:2017(E), 2 017","cited_arxiv_id":null,"evidence_quote":"One of four standards sources from which benchmark terminology is extracted."},{"cited_title":"ISTQB Glossary,","cited_arxiv_id":null,"evidence_quote":"ISTQB glossary source for benchmark terminology."},{"cited_title":"CPRE Glossary,","cited_arxiv_id":null,"evidence_quote":"IREB glossary source for benchmark terminology."},{"cited_title":"Augmenting the Generality and Performance of Large Language Models for Software Engineering","cited_arxiv_id":"2506.11548","evidence_quote":"iSAQB source for benchmark terminology."}],"review_version":1}