{"id":"04ad1333-9c43-4cc0-b104-ceef5f9afc0d","arxiv_id":"2507.13629","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that maps LLM applications, vulnerabilities, and defenses across eight cybersecurity domains, but with significant citation and rigor problems.","lead":"This paper is a survey of how large language models are used in cybersecurity, covering applications, attacks, and defenses across eight domains. It organizes 172 recent studies into a taxonomy, but contains numerous citation errors and an unsystematic selection method.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey claims comprehensiveness, but its citation-to-reference mapping is demonstrably unreliable, making the integrated synthesis unverifiable.","rationale":"I read the paper as a broad survey whose central value is its taxonomy and synthesis of LLM applications, vulnerabilities, and defenses in cybersecurity. The strongest claim is explicitly about comprehensiveness and novel integration, so the load-bearing assumption is that the 172 cited studies are accurately identified and correctly represented. The reader identified this same assumption and provided specific citation failures; my independent reading confirms those failures and adds another (Section 4.6.3 / [38]) plus the Section 2 count inconsistency. These are not matters of disagreeing with community consensus; they are direct checks on whether the survey's evidentiary base is sound. A survey can survive a few typos, but unresolved placeholders, a citation of the Chinchilla paper for an unrelated NIDS result, and a table entry whose reference points to a different work undermine trust in every other attribution. The non-systematic methodology described in Section 2 ('timeliness, coverage, diversity, and relevance' with no search protocol) further weakens the comprehensiveness claim, though the concrete citation failures are sufficient on their own. Since the reader's CONDITIONAL verdict already captures the need for careful reference cleanup and a more rigorous methodology, my stress-test does not change the verdict. I am not raising any objection to the authors' conduct; the critique is about the verifiability of the survey's central claim.","tokens_in":39214,"tokens_out":4053,"duration_ms":46689,"concrete_test":"Reconstruct the complete citation mapping by extracting every in-text [n] marker in Sections 4 and 5 and checking it against the reference list; compute the fraction of markers whose cited author and topic match the corresponding reference entry, including the known cases [7], [38], [52], and the unresolved [?] marker. If the mismatch rate exceeds 5% of checked markers, the survey's detailed literature claims are not verifiable and the 'comprehensive enumeration' claim must be downgraded. As a second part of the same check, reconcile the Section 2 selection totals (172 vs. 114+12+31, with or without 'Others 17'); if the totals cannot be reconciled from the reported methodology, the selection process is not reproducible.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that it is 'the first survey to integrate the application landscape in cybersecurity with the attack surface and defense mechanisms' and provides 'a comprehensive enumeration' of 32 tasks across 8 domains. For this claim to hold, every textual assertion about prior work must be traceable to the correct source. That condition fails in multiple concrete places. Section 4.2.1 leaves LATTE as 'Liu et al. [?]' with an unresolved reference marker. Section 4.1.2 cites 'Houssel et al. [7]' for NetFlow-based NIDS explainability, but reference [7] is the Chinchilla paper, not the cited work. Table 2 attributes BLOCKGPT to 'Gai et al. [52]', but reference [52] is 'BC4LLM' by Luo et al. Section 4.6.3 attributes to 'Lanka et al. [38]' an LLM analysis of decoy-system data, while reference [38] is Liu et al.'s traffic detection work. These are not isolated typographical slips; they show the numeric citation markers cannot be treated as reliable links between text claims and their sources. If the mapping is unreliable, readers cannot verify whether the survey's synthesis accurately represents the literature, so the 'comprehensive' and 'first integration' claims lose their evidentiary support. The paper also reports internally inconsistent selection totals in Section 2: 15+25+64+68 = 172 studies, but '114 studies on applications, 12 on attacks, and 31 on defense techniques' sums to 157, while Figure 1 includes an unexplained 'Others 17' category. The central claim therefore rests on a partly contradicted and non-reproducible evidentiary base.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of LLM applications in cybersecurity, organized around eight domains and 32 tasks, followed by a catalog of LLM vulnerabilities (data poisoning, backdoor attacks, prompt injection, jailbreaking) and corresponding defenses. The authors claim to be the first to integrate the application landscape with the attack surface and defense mechanisms, and they derive their inventory from 172 selected publications from 2021–2024. The paper provides multiple tables and figures mapping papers to tasks and defense techniques.","tokens_in":39568,"tokens_out":8190,"duration_ms":87301,"significance":"If the bibliographic issues were repaired, the survey would have real value: the eight-domain taxonomy of Figure 2 and Table 3 is a sensible organizational scheme, Table 2 condenses many application papers, and Tables 4–5 provide a compact comparison of defense techniques. The simultaneous treatment of applications, attacks, and defenses in one document is a service to newcomers. However, in its current form the central claim of being 'comprehensive' and 'first' is not verifiable because several numeric citation markers do not point to the cited works, and at least one reference is unresolved.","major_comments":[{"comment":"The LATTE tool is introduced as 'Liu et al. [?]', with no resolved reference. Since LATTE is used in the Takeaway to demonstrate the robustness of LLM-based vulnerability detection, this unresolved placeholder prevents readers from tracing a substantive claim. Please supply the correct citation and renumber the reference list accordingly.","section":"Section 4.2.1"},{"comment":"The blockchain anomaly-detection result is attributed to 'Gai et al. [52]' in both the text and Table 2, but reference [52] is Luo et al., 'BC4LLM: A perspective of trusted artificial intelligence when blockchain meets large language models' (Neurocomputing, 2024). This is not the BLOCKGPT paper. The claim about LLM-based transaction anomaly detection is therefore not traceable; either add the correct Gai et al. reference or replace the citation marker.","section":"Section 4.5.2 and Table 2"},{"comment":"The Introduction cites 'Chinchilla [7]', but reference [7] is Houssel et al., 'Towards explainable network intrusion detection using large language models' (2024). This is a foundational-model misattribution, not a style issue. Please recheck all numeric markers attached to model names in the Introduction, since this error indicates the numbering cannot be trusted.","section":"Section 1"},{"comment":"The study-selection accounting is internally inconsistent: 15+25+64+68 = 172 studies, but the breakdown '114 studies on applications, 12 on attacks, and 31 on defense techniques' sums to 157. Figure 1 adds an unexplained 'Others 17' category, giving 174, which still does not reconcile to 172. Please provide a single consistent accounting of the corpus.","section":"Section 2 and Figure 1"},{"comment":"The row 'Liu et al. [38]' describes CharBERT-based URL and traffic detection, but reference [38] is Lanka et al., 'Intelligent threat detection – AI-driven analysis of honeypot data' (Electronics, 2024). The same number is used correctly in Section 4.6.3 for the honeypot study, so the table's source mapping is internally inconsistent. The row should presumably cite Liu et al. [57], and the authors should audit the rest of the table for similar mismatches.","section":"Table 2"}],"minor_comments":[{"comment":"The text says 'Tables 2 and 2' present the LLM-based solutions, but only one Table 2 appears in the manuscript. Please add the missing part or correct the pointer.","section":"Section 4"},{"comment":"Reference [50] has no author names ('H.-c. I. Cybersecurity'), making it impossible to verify the OllaBench citation. Please complete the entry.","section":"Reference list"},{"comment":"There are numerous spelling errors in author names and terms, e.g., 'Gangualiet al.' (should be Ganguli), 'Ribeiroetet al.', 'Kupmaret al.', 'Alone et al.', 'Contiella', 'Safty Fine Tuning' in the Table 5 header, and 'more 50%' in Section 4.2.1. A thorough proofread is needed.","section":"Multiple locations"},{"comment":"The stated time window is 2021–2024, yet several references, including [26] and [174], are dated 2025. Either adjust the selection criterion or the reported counts.","section":"Section 2"},{"comment":"The category 'Others 17' is not defined in the text. If it represents papers outside the application/attack/defense breakdown, it should be explained and included in the Section 2 accounting.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The conditional verdict is appropriate. The citation failures are concrete, checkable, and fixable, so I recommend major revision rather than rejection. At least five mismatches suggest the problem may be systemic, and I would ask the authors to verify every numeric citation in the final version before resubmission. There is no concern about novelty or scope; the survey fits the journal's readership."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plain-English take: this is a genuinely useful map of LLM-for-security work, but the citation layer is not trustworthy in its current form. A placeholder reference and several misdirected numeric markers mean you should not cite it as-is.\n\nWhat it does well: it organizes 32 security tasks across eight domains (network, software, hardware, blockchain, cloud, IoT, IR/CTI, content), writes one-paragraph takeaways per domain, and pairs each attack type (backdoor, poisoning, jailbreak, prompt injection) with a defense table. That integration is a legitimate service, though the \"first survey to integrate applications with attacks and defenses\" claim is too strong — Ferrag et al. [34] already covers apps, vulnerabilities, and defenses across all these domains. For a newcomer, this is a serviceable overview of 2021–2024 literature.\n\nThe soft spots are real but not fatal. The reference errors are the main problem. I verified the specifics: the LATTE placeholder ('Liu et al. [?]' in §4.2.1) is unresolved; BLOCKGPT is attributed to 'Gai et al. [52]' in both Table 2 and §4.5.2, while [52] is Luo et al.'s BC4LLM paper. However, two of the claimed mismatches in the stress-test note are themselves wrong: [7] is Houssel et al., not the Chinchilla paper, so the body's 'Houssel et al. [7]' is fine; the actual error is the intro's 'Chinchilla [7]'. Likewise [38] is indeed Lanka et al.'s honeypot analysis, so §4.6.3 is correct; the problem is Table 2, which mislabels Liu et al. as [38]. The takeaway is the same: the numeric citation markers cannot be trusted without cross-checking, and that undermines the 'comprehensive enumeration' claim. The selection counts are also inconsistent: the year counts sum to 172, but the application/attack/defense breakdown sums to 157, and Figure 1's 'Others 17' doesn't reconcile the difference. The methodology is non-systematic (no search protocol, no inclusion criteria), which makes the comprehensiveness claim unverifiable.\n\nWho this is for: a graduate student starting in LLM security, or a group that wants a quick domain map. Not for someone who needs reliable pointers to specific papers until the references are fixed.\n\nRecommendation: I'd accept it for review rather than desk-reject. The errors are fixable, the taxonomy is useful, and the area is important enough to warrant referee time. But I would make the authors fix every broken reference, reconcile the counts, and add a real methodology before publication. I wouldn't cite it in my own work in its current state.","headline":"Useful taxonomy, unreliable citations: worth reading, not worth citing until the references are cleaned up.","tokens_in":40032,"tokens_out":5691,"would_cite":false,"duration_ms":58810,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey integrates LLM applications, vulnerabilities, and defenses in cybersecurity into one map.","keywords":["LLMs","cybersecurity","prompt injection","jailbreaking","backdoor attacks","data poisoning","security survey","threat detection"],"falsifier":"An independent check of the reference list would settle the central claim: resolve every in-text citation marker (including the missing marker for LATTE in Section 4.2.1 and the Houssel et al. marker [7] that currently resolves to the Chinchilla paper) and verify that each cited study supports the sentence that cites it; if the markers do not check out, the survey's comprehensiveness and integration claims cannot be audited. A second check would re-derive the 32-task, eight-domain categorization from the same 172 papers and see whether the counts reproduce.","tokens_in":39044,"feed_emoji":"🛡️","tokens_out":5649,"duration_ms":58638,"temperature":0.7,"pith_summary":"This survey tries to establish that large language models have penetrated cybersecurity so broadly that only a unified map can make sense of the field. It argues for a two-part synthesis: an enumeration of LLM applications across 32 security tasks in eight domains, and a matching of four attack families against four families of defense. The point of the integration is that the same models that power threat detection, fuzzing, and malware analysis are themselves vulnerable to prompt injection, jailbreaking, data poisoning, and backdoor attacks, so secure deployment requires treating the tool and the threat as one problem.","feed_headline":"Survey maps LLM security use to 32 tasks in 8 domains","feed_subtitle":"First synthesis of applications, vulnerabilities, and defenses gives practitioners a single navigable map.","key_machinery":"The load-bearing device is a two-part classification system. The first part is the eight-domain taxonomy of security tasks (Table 3), which organizes 32 tasks into network security, software and system security, information and content security, hardware security, blockchain security, cloud security, incident response and threat intelligence, and IoT security. The second part is the attack-defense pairing (Table 4 and Figure 3), which maps four attack families — backdoor, data poisoning, prompt injection, and jailbreaking — onto defense families that include red teaming, content filtering, safety fine-tuning, and model merging. These two structures do the argumentative work: the taxonomy establishes breadth, and the pairing establishes that vulnerabilities and defenses are intrinsic to LLM deployment, not external concerns.","core_discovery":"The central claim is that LLM use in cybersecurity is wider than any prior survey captured, and that the application landscape and the vulnerability landscape must be read together. The paper categorizes LLM applications into eight domains — network, software and system, information and content, hardware, blockchain, cloud, incident response and threat intelligence, and IoT — covering 32 distinct security tasks. It then pairs four major attack types (backdoor, data poisoning, prompt injection, jailbreaking) with defense techniques such as red teaming, content filtering, safety fine-tuning, and model merging. In the authors' telling, this is the first survey to integrate these two sides, and the integration is what reveals design gaps, such as the finding that over half of LLM-generated code in one benchmark contained vulnerabilities.","pith_inferences":["The survey's separation of applications from vulnerabilities obscures a feedback loop the paper hints at: LLMs used for fuzzing and penetration testing are themselves prompt-injectable, so a compromised model could silently weaken the defenses it is supposed to operate.","The eight-domain, 32-task count is a snapshot of the literature through early 2025; a live version of this taxonomy could be re-run at intervals to track the field's growth, and the count would likely rise as autonomous agent security and AI supply-chain security become distinct categories.","The citation inconsistencies suggest that a versioned, machine-checkable reference database with author-contributed markers would make future surveys auditable; a testable extension is to rebuild the taxonomy from the same 172 papers using resolved markers and compare the resulting counts.","The 50% vulnerable-code finding, if it holds across other datasets, implies that any security application that relies on LLM-generated code should assume the code is vulnerable until a verifier proves otherwise, a stronger operational rule than the paper states."],"forward_implications":["Practitioners can use the eight-domain map to locate which security tasks already have LLM-based tooling and which domains remain thin.","The attack-defense pairing gives deployment teams a structured menu for hardening LLM-based security tools against the four main attack families.","The reported result that over 50% of LLM-generated code in the FormAI dataset contained vulnerabilities implies that any LLM-generated code entering security pipelines must be screened before use.","The survey's identification of underexplored domains — IoT, cloud, hardware, and blockchain — points to where the next generation of LLM security research is likely to concentrate.","Because the same capabilities that detect threats can also generate attacks, the survey implies that defensive and offensive LLM use cannot be governed separately."],"supporting_citations":[{"why":"Prior systematic literature review of LLMs for cyber security; the survey positions its integrated scope against this one's coverage of applications and datasets.","marker":"[31]"},{"why":"Prior review of generative AI methods in cybersecurity; supplies the attack-oriented baseline this survey contrasts with.","marker":"[32]"},{"why":"Earlier survey of 42 models, their roles in cybersecurity, and known flaws; the direct predecessor that this paper claims to extend across eight domains.","marker":"[34]"},{"why":"Prior survey of LLM potential for system-on-chip security; one of the domain-specific reviews that the eight-domain taxonomy consolidates.","marker":"[29]"},{"why":"Prior systematic review of LLMs for blockchain security; another domain-specific review that the integrated map folds in.","marker":"[30]"}],"fun_headline_variants":["LLM security survey links 32 tasks to 4 attacks, defenses","First joint map of LLM attacks and defenses across 8 domains","LLM cybersecurity survey: 32 tasks, 4 attack types, 8 domains","Survey shows LLM security gaps: half of generated code vulnerable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The synthesis is only as reliable as the accuracy of the 172 cited studies and their reference markers, and several markers in the paper are missing or mismatched, which a reader would need to resolve before trusting the survey's counts and attributions.","fun_headline_variants_meta":{"raw":{"variants":["LLM security survey links 32 tasks to 4 attacks, defenses","First joint map of LLM attacks and defenses across 8 domains","LLM cybersecurity survey: 32 tasks, 4 attack types, 8 domains","Survey shows LLM security gaps: half of generated code vulnerable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2839,"prompt_tokens":810,"completion_tokens":2029,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":1950}},"tokens_in":426,"tokens_out":2029,"duration_ms":18742,"temperature":1.0,"reasoning_tokens":1950,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:19:17.749784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent check of the reference list would settle the central claim: resolve every in-text citation marker (including the missing marker for LATTE in Section 4.2.1 and the Houssel et al. marker [7] that currently resolves to the Chinchilla paper) and verify that each cited study supports the sentence that cites it; if the markers do not check out, the survey's comprehensiveness and integration claims cannot be audited. A second check would re-derive the 32-task, eight-domain categorization from the same 172 papers and see whether the counts reproduce.","supporting_citations":[],"review_version":1}