Pith. sign in

REVIEW 2 major objections 3 minor

SOC practitioners see LLMs as useful for routine automation but not yet feasible for high-impact incident analysis, and locate the barrier in organizational readiness and human over-reliance rather than model capability alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 00:47 UTC pith:YDHLYV4P

load-bearing objection A useful, original interview study of SOC practitioners' views on LLMs, but the 25-interview sample can't carry the generalized 'practitioners rate' claim without more sampling evidence. the 2 major comments →

arxiv 2608.00672 v1 pith:YDHLYV4P submitted 2026-08-01 cs.CR cs.AIcs.HC

From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness

classification cs.CR cs.AIcs.HC
keywords Security Operations CenterSOC practitionersLarge Language ModelsLLM integrationincident analysishuman factorsover-relianceuse case taxonomy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper is an empirical study of how security operations center (SOC) practitioners with hands-on LLM experience view integrating large language models into their work. Based on 25 semi-structured interviews and interactive scenarios, it identifies 15 LLM use cases in six categories and finds a clear split: practitioners welcome LLMs for repetitive, low-ceiling tasks such as report automation, but rate high-impact tasks like incident analysis as not yet feasible. They attribute this less to model limitations than to their SOCs' lack of readiness and to human factors such as over-reliance. Despite these reservations, the same practitioners report a strong willingness to adopt LLMs, citing competitive pressure that leaves few alternatives. The paper turns these findings into concrete design and integration requirements for human-centered, operationally safe LLM-assisted security operations.

Core claim

The central discovery is a systematic map of where LLMs fit in SOC work and where they do not. Across six functional categories, the authors catalogue 15 practitioner-identified use cases, ranging from automation of routine reporting to incident analysis. The decisive pattern is a feasibility–impact gap: the tasks most valuable to automate are the ones practitioners say are still beyond current LLM competence. They report concrete limitations in technical depth, context awareness, and organization-specific knowledge, but locate the root cause not in the models themselves so much as in the readiness of their SOCs and in human factors, particularly the risk of over-reliance on model output. Th

What carries the argument

The argument is carried by the qualitative interview-and-scenario method: 25 semi-structured interviews with experienced SOC practitioners, each paired with interactive scenarios designed to surface anticipated challenges. The authors use this material to build a six-category taxonomy of 15 LLM use cases, and they use the contrast between perceived task impact and perceived feasibility as the analytical lens for distinguishing where LLMs are ready for deployment from where they are not. The notion of over-reliance serves as the key conceptual mechanism for explaining why human factors, not model capability, are seen as the binding constraint.

Load-bearing premise

The load-bearing premise is that 25 self-selected practitioners who already had LLM experience are representative of the broader SOC practitioner population, and that what they said in interviews and scenarios reflects what they would actually do in live operations—a premise the abstract does not yet support with sampling or saturation evidence.

What would settle it

A controlled deployment comparing SOC incident-analysis outcomes (time-to-detect, accuracy, false-positive rate) with and without LLM assistance, in several organizations of different maturity levels, would test the central claim. If analysts with LLM support were found to outperform unassisted analysts on high-impact incident analysis tasks without elevated over-reliance, the paper's feasibility assessment would be contradicted. Similarly, a larger-scale survey of SOC practitioners beyond the 25 interviewees that showed majority confidence in LLM incident analysis would weaken the 'not yet fe

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • LLM-based tools should be introduced first in the low-impact automation tier (e.g., report drafting, triage support), where practitioners already see value, before expanding toward incident analysis.
  • Incident analysis will remain a human-led activity for the foreseeable future, and vendors should not position LLMs as autonomous replacements in that role.
  • Improving model technical depth, context awareness, and organization-specific knowledge is necessary but not sufficient; SOC readiness—training, workflows, and data hygiene—must advance in parallel to prevent over-reliance.
  • Design requirements derived from the study should guide procurement: human-centered interfaces, clear confidence signaling, and escalation paths that keep a human in the loop.
  • The competitive pressure reported by practitioners suggests that adoption may outpace safety; organizations need explicit integration policies rather than ad-hoc experimentation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the feasibility–impact gap is real and persistent, then near-term gains from LLMs in security operations are likely to be incremental productivity improvements on routine tasks, while the highest-risk decisions stay with humans—an outcome that should temper both vendor promises and organizational cost-benefit expectations.
  • The authors' framing that limitations are located in SOC readiness rather than model capability implies a testable extension: as SOCs invest in context-rich tooling and knowledge management, the same models could become feasible for high-impact tasks without a jump in raw capability.
  • The competitive-pressure dynamic suggests a possible liability gap: if organizations adopt LLM assistance before their processes are ready, errors in incident analysis may be blamed on the model while the underlying readiness deficit goes unaddressed.
  • A natural generalization is to other high-stakes professional domains (digital forensics, threat hunting, incident response) where practitioners face similar tension between automation pressure and the cost of being wrong.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper presents a qualitative study based on 25 semi-structured interviews with SOC practitioners who had prior LLM experience, supplemented by interactive scenarios. It identifies 15 LLM use cases in six categories and reports that practitioners value LLMs for low-level automation but consider high-impact tasks like incident analysis infeasible due to limitations in technical depth, context awareness, and organization-specific knowledge. The authors interpret these limitations as rooted more in SOC readiness and human factors (e.g., over-reliance) than in the models themselves, and report strong willingness to adopt despite competitive pressure. The paper derives design and integration requirements. This report is based solely on the abstract, as the full text was not provided.

Significance. If the methodology is sound, the study is a valuable empirical contribution to the growing literature on LLMs in security operations, offering practitioner perspectives that are often missing. Its strength is the use of direct interview data and interactive scenarios. However, the significance hinges on the representativeness of the sample and the rigor of the qualitative analysis, which cannot be assessed from the abstract.

major comments (2)
  1. [Abstract] The central claim 'practitioners rate' rests on 25 interviews with practitioners who had prior LLM experience. The abstract provides no information about the sampling frame, recruitment strategy, organizational diversity, or thematic saturation. If the sample was drawn from a narrow set of early-adopter organizations, the conclusions about 'readiness' and 'human factors' would be artifacts of sample selection. Please report saturation, participant demographics, and organizational spread, or qualify the claims as specific to this sample.
  2. [Abstract] The sentence 'practitioners rate high-impact tasks such as incident analysis as not yet feasible, reporting limitations in technical depth, context awareness, and organization-specific knowledge. They locate these limitations less in the models than in the readiness of their SOCs and human factors driving over-reliance' is a strong interpretive claim. The abstract does not explain how this attribution was derived from the data. Without a coding scheme, inter-rater reliability check, or evidence that this reflects participants' own attributions rather than the researchers' framing, this conclusion is under-supported. Please specify the analysis procedure.
minor comments (3)
  1. [Abstract] The abstract does not define the 'six functional categories' or list the 15 use cases; a brief exemplar would help.
  2. [Abstract] The 'interactive scenarios' are mentioned but not described; a sentence on scenario design would clarify the method.
  3. [Abstract] The phrase 'competitive pressure' is introduced without supporting evidence in the abstract; consider citing a representative quote.

Circularity Check

0 steps flagged

No circularity: the findings are emergent from interview data, not derived from fitted inputs or self-citations.

full rationale

The paper is an abstract-only empirical interview study. Its central claims—that LLMs are valued for automating repetitive low-level tasks, that high-impact tasks are rated as not yet feasible, and that limitations are located in SOC readiness and human factors—are presented as outcomes of 25 semi-structured interviews and interactive scenarios, not as the output of a mathematical derivation or a model fitted to a subset of data. There is no equation, no fitted parameter later renamed as a prediction, and no uniqueness theorem or self-citation used to force the conclusion. The identified weakness regarding sample representativeness or saturation is a methodological threat to external validity, not a circularity: the claim does not reduce to its input by construction. Since no specific reduction between inputs and outputs can be exhibited from the abstract, the appropriate finding is no significant circularity with score 0.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

No free parameters or invented entities are relevant for a qualitative interview study. The central claims rest on three domain assumptions about sample representativeness, data validity, and coding fidelity, none of which can be verified from the abstract.

axioms (3)
  • domain assumption The 25 interviewed SOC practitioners with prior LLM experience are representative of the broader SOC practitioner population.
    The abstract generalizes from 25 interviews to SOC practices; representativeness is assumed but not demonstrated in the abstract.
  • domain assumption Practitioners' self-reports and scenario responses accurately reflect their real-world performance and beliefs.
    Interview data are treated as evidence about actual feasibility and risks; social-desirability and hypothetical-scenario bias could affect this.
  • domain assumption The researchers' thematic grouping of '15 use cases into six functional categories' is a faithful representation of the data rather than an artifact of coding.
    The abstract asserts this taxonomy without showing inter-rater reliability or coding audit; we assume the analysis is sound.

pith-pipeline@v1.3.0-daily-deepseek · 567 in / 5628 out tokens · 52557 ms · 2026-08-04T00:47:57.845798+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness." pith.science (2026). https://pith.science/paper/YDHLYV4P

@misc{pith2026260800672,
  author       = {Pith},
  title        = {Pith review of: From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YDHLYV4P}},
  note         = {Machine review of arXiv:2608.00672}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Security Operations Centers (SOCs) process large volumes of security events, requiring analysts to accurately detect and assess ongoing cyberattacks under time pressure. Recent advances in Large Language Models (LLMs) suggest potential benefits for security operations, yet their practical suitability for real-world SOC workflows remains poorly understood. To address this gap, we conducted 25 semi-structured interviews with SOC practitioners who had prior experience with LLMs, complemented by interactive scenarios to anticipate challenges and identify opportunities for the responsible integration of LLM-based tools into SOC workflows. We identified 15 LLM use cases grouped into six functional categories. While LLMs are valued for automating repetitive, low-level tasks such as report automation, practitioners rate high-impact tasks such as incident analysis as not yet feasible, reporting limitations in technical depth, context awareness, and organization-specific knowledge. They locate these limitations less in the models than in the readiness of their SOCs and human factors driving over-reliance. Despite concerns, practitioners express a strong willingness to adopt LLMs, describing competitive pressure that leaves few alternatives. This work contributes an empirical, practitioner-driven analysis of LLM use across SOC roles and organizations and derives concrete design and integration requirements for human-centered, operationally safe LLM-assisted security operations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.