REVIEW 3 major objections 6 minor 52 references
A chain-of-thought triage model trained on real SOC detections, paired with a separate calibrator that reads its reasoning, sharply raises high-confidence auto-triage recall over direct-label LLMs.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 06:39 UTC pith:3FJTOY4E
load-bearing objection Solid industrial CoT+calibrator triage paper with real SOC data and clean ablations; the headline +43 pp benign auto-close gain is overstated once you respect the validation-to-test precision drift the authors themselves report. the 3 major comments →
Cybersecurity Detection Classification with Reasoning-enabled Language Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Explicitly training a language model to produce chain-of-thought reasoning about real Windows endpoint detections, then pairing it with a separately trained calibrator that scores the full reasoning trace, yields 82.6% test accuracy and improves high-confidence benign recall by 43.0 percentage points and malicious recall by 18.3 percentage points over a conventional direct-label LLM classifier; without the trained calibrator, high-confidence recall collapses to zero.
What carries the argument
A four-stage recipe—automated prompt optimization, self-training that keeps only correct or genuinely rationalized reasoning traces, reinforcement learning with a binary format-and-correctness reward, and a second model that maps detection plus reasoning plus label to a correct/incorrect probability—together produce both the verdict and the actionable confidence.
Load-bearing premise
High-precision confidence thresholds chosen on validation, and a calibrator fit on training rollouts, will still mark the right detections as safe to auto-act on when the data comes from a later time period.
What would settle it
On a fresh, later time window of the same SOC stream, measure whether the validation-chosen high-confidence thresholds still deliver near the claimed ~98–99% precision with non-trivial recall; a large drop in precision or a collapse of high-confidence recall would falsify the automation claim.
If this is right
- High-confidence tiers of endpoint detections can be auto-closed or auto-prioritized at substantially higher coverage without raising analyst false-positive load.
- Label-token probabilities alone are not a usable confidence signal once chain-of-thought is used; a trained judge over the reasoning trace becomes mandatory for precision-targeted automation.
- Each training stage (prompt optimization, self-training, RL with verifiable rewards) compounds gains at the calibrated high-confidence operating point rather than only lifting raw accuracy.
- A domain-finetuned 30B reasoning model can outperform larger frontier general-purpose models on this SOC task, favoring targeted training over raw scale.
- Analysts can be shown accurate reasoning traces on medium-confidence cases while machines handle the high-confidence tail, stretching limited human triage capacity.
Where Pith is reading between the lines
- The same calibrator-over-reasoning pattern is likely required for any high-stakes CoT classifier where auto-action depends on precision targets, not only cybersecurity.
- Temporal drift of the benign high-precision point suggests production systems will need scheduled recalibration or online threshold monitoring, not a one-time fit.
- Extending the recipe beyond binary Windows endpoint labels to multi-class triage or other sensor platforms is the natural operational next test of whether the gains are architecture-general.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper trains a chain-of-thought triage classifier on real, human-labeled Windows endpoint detections using a four-stage recipe (GEPA prompt optimization, AdaSTaR-style self-training, GRPO with format-and-correctness rewards, and a separate LoRA calibrator that reads the full reasoning trace). On a temporally held-out test set the system reaches 82.6% accuracy and, at validation-chosen high-confidence thresholds, reports large lifts over a direct-label SFT baseline in benign and malicious recall. Ablations show that each training stage helps, that an untrained confidence judge collapses high-confidence recall to zero, and that a finetuned 30B model beats stronger zero-shot frontier models on this task.
Significance. If the high-confidence automation claims hold under the SOC’s actual precision constraints, the work is practically important: it targets a real bottleneck (alert fatigue), uses large industrial labels with a temporal split, and shows that CoT plus a trained trace-level calibrator can materially expand the auto-close and prioritization tiers. The factorial ablation establishing that an untrained judge is unusable, and the demonstration that domain finetuning of a 30B model beats larger general-purpose models, are concrete contributions. The pipeline is clearly specified enough to be reproduced in similar industrial settings.
major comments (3)
- [Table 1; Results / Overall Metrics] Table 1, Benign High row: the central automation claim (+43.0 pp benign recall at the tier that “governs automated triage”) is defined by validation thresholds targeting 98% precision, but on the temporally later test set those same thresholds yield only 90.8% precision (vs 98.0% on validation), while the direct-label baseline stays at 98.3%. The paper attributes the drop to a 4.4-point policy accuracy gap and greater calibrator distribution sensitivity. Headline benign recall is therefore not measured at the SOC-required operating point. Please report benign (and malicious) recall after re-selecting or adjusting the threshold on test (or via a nested protocol) so that test precision actually meets 98%/99%, and state whether a material gain over the baseline’s 21.8% remains. Without that, the largest claimed improvement is not established for deployment.
- [Table 2; Methods / Confidence Calibration; Experiments / Training] Table 2 and Confidence Calibration: each row’s calibrator is trained independently and the checkpoint is chosen to maximize benign validation recall at 98% precision. That selection criterion directly optimizes the primary reported metric and, together with training the calibrator on the same policy’s rollouts, can inflate the apparent stage-wise gains and the necessity claim. Please add (i) a fixed early/mid/late checkpoint comparison or selection on a metric that is not the test headline (e.g., ECE or a held-out calibration split), and (ii) results when the calibrator is trained only on an earlier policy or on a disjoint detection subset, so that policy and calibrator improvements are not confounded.
- [Table 1; Table 3; Results] The direct-label SFT baseline isolates “no CoT,” but does not control for total training compute, number of labeled passes, or access to a second model. A stronger control would be a direct-label model given comparable GRPO steps and/or a calibrator that sees only the detection (no CoT). Without that, part of the +43/+18 pp lift may be attributable to extra optimization and the second network rather than to reasoning per se. A compact compute-matched or “calibrator-on-detection-only” ablation would make the CoT contribution load-bearing rather than suggestive.
minor comments (6)
- [Abstract; Introduction] Abstract and Introduction state “improves benign recall by 43.0% and malicious recall by 18.3%”; these are percentage points, not relative percent. Align wording with Table 1.
- [Figure 3] Figure 3 is helpful but the abridged CoT and single example cannot support generality; consider adding failure cases (high-confidence incorrect) in the appendix.
- [Methods; Appendix prompts] Label vocabulary shifts between TRUE_POSITIVE/FALSE_POSITIVE in prompts and malicious/benign in the main text; a single mapping note would reduce confusion.
- [Table 4] Table 4 uses a 1,000-sample subset at T=0 with no confidence threshold; state how the subset was drawn and whether ranks change on the full test set.
- [Table 1] Medium and low tiers are often identical because of bimodal confidence; that is noted in text but should also be flagged in Table 1’s caption so readers do not over-interpret tier granularity.
- [Related Work] Related Work could briefly contrast with other SOC LLM triage / false-positive filtering agents beyond the cited IT/cloud RCA systems, to sharpen the “first CoT-trained detection triage” claim.
Circularity Check
No circularity: empirical SOC triage pipeline evaluated against external human labels on held-out temporal splits.
full rationale
The paper’s load-bearing claims are measured performance numbers (82.6% accuracy; high-confidence benign/malicious recall lifts vs a direct-label SFT baseline) on real, human-labeled Windows endpoint detections with train/val/test collected over consecutive time windows. The CoT policy is optimized with GEPA, AdaSTaR-style self-training, and GRPO under a verifiable reward that is 1 iff the parsed label matches the external ground-truth label and the format is valid—not a quantity defined from the model’s own outputs. The calibrator is supervised on correct/incorrect relative to those same external labels and is scored by precision-targeted recall on validation, then frozen for test; that is ordinary post-hoc thresholding, not a fitted constant renamed as a first-principles prediction. Ablations (untrained judge → zero high-confidence recall; stage-wise gains; frontier zero-shot baselines) are comparative experiments, not self-definitional reductions. Citations (GEPA, AdaSTaR, GRPO, Nemotron, classical calibration) are external methods, not author uniqueness theorems that force the result. Distribution-shift fragility of the benign 98% operating point is a validity concern, not circularity. No step reduces Eq./claim X to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- High/medium/low precision targets (benign 98/95/90%, malicious 99/90/80%) =
98/95/90 benign; 99/90/80 malicious
- Calibrator checkpoint selection criterion =
max val benign recall@98% precision
- GEPA/AdaSTaR/GRPO/LoRA hyperparameters =
as listed in Methods/Appendix
- Instruction-quality judge dimensions and min-aggregation =
metric = accuracy × min(five scores)
axioms (5)
- domain assumption Human SOC labels on Windows endpoint detections are a valid ground truth for malicious vs benign triage.
- domain assumption Temporal consecutive week splits (8/2/1) adequately represent deployment distribution for claimed automation metrics.
- ad hoc to paper Binary format-and-correctness reward (1 iff well-formed CoT+label and label matches) is a sufficient verifiable reward for useful triage reasoning.
- domain assumption A separate LLM judge of correctness from (detection, CoT, label) yields actionable confidence via softmax of correct/incorrect tokens.
- standard math Open-weight Nemotron-3 Nano/Super and GEPA/AdaSTaR/GRPO behave as described in cited works under the authors’ modifications.
invented entities (3)
-
CoT triage policy (finetuned Nemotron-3-Nano-30B)
no independent evidence
-
Trace-level correctness calibrator (second Nano-30B)
no independent evidence
-
GEPA-optimized Charlotte AI system prompt with LLM-as-judge guardrails
no independent evidence
read the original abstract
A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards. We find that CoT reasoning also degrades the label-token probabilities that automated triage relies on, so we separately train a calibrator that reads the full reasoning trace and estimates the probability that the verdict is correct. Our system reaches 82.6% test accuracy and, at the high-confidence operating point that governs automated triage, improves benign recall by 43.0% and malicious recall by 18.3% over a direct-label LLM classifier. We further show that the trained calibrator is necessary - an untrained confidence judge collapses high-confidence recall to zero - and that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.
Figures
Reference graph
Works this paper leans on
-
[1]
2017 , eprint =
Attention Is All You Need , author =. 2017 , eprint =
2017
-
[2]
Zelikman, Eric and Wu, Yuhuai and Mu, Jesse and Goodman, Noah D. , year =. 2203.14465 , archivePrefix =
-
[3]
Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, Y. K. and Wu, Y. and Guo, Daya , year =. 2402.03300 , archivePrefix =
-
[4]
and Axon, Louise and Martinovic, Ivan , booktitle =
Alahmadi, Bushra A. and Axon, Louise and Martinovic, Ivan , booktitle =. 99\. 2022 , publisher =
2022
-
[5]
Hassan, Wajih Ul and Guo, Shengjian and Li, Ding and Chen, Zhengzhang and Jee, Kangkook and Li, Zhichun and Bates, Adam , booktitle =
-
[6]
Proceedings of the 13th USENIX Conference on System Administration (LISA '99) , pages =
Snort: Lightweight Intrusion Detection for Networks , author =. Proceedings of the 13th USENIX Conference on System Administration (LISA '99) , pages =. 1999 , publisher =
1999
-
[7]
2010 IEEE Symposium on Security and Privacy , pages =
Outside the Closed World: On Using Machine Learning for Network Intrusion Detection , author =. 2010 IEEE Symposium on Security and Privacy , pages =. 2010 , publisher =
2010
-
[8]
2019 , publisher =
Pendlebury, Feargus and Pierazzi, Fabio and Jordaney, Roberto and Kinder, Johannes and Cavallaro, Lorenzo , booktitle =. 2019 , publisher =
2019
-
[9]
2025 , eprint =
Large Language Models for Security Operations Centers: A Comprehensive Survey , author =. 2025 , eprint =
2025
-
[10]
Daniel, Nir and Kaiser, Florian Klaus and Giladi, Shay and Sharabi, Sapir and Moyal, Raz and Shpolyansky, Shalev and Murillo, Andres and Elyashar, Aviad and Puzis, Rami , year =. Labeling. 2412.10978 , archivePrefix =
-
[11]
ACM Transactions on Information and System Security , volume =
Clustering Intrusion Detection Alarms to Support Root Cause Analysis , author =. ACM Transactions on Information and System Security , volume =
-
[12]
Recent Advances in Intrusion Detection (RAID) , year =
Using Adaptive Alert Classification to Reduce False Positives in Intrusion Detection , author =. Recent Advances in Intrusion Detection (RAID) , year =
-
[13]
Du, Min and Li, Feifei and Zheng, Guineng and Srikumar, Vivek , booktitle =
-
[14]
van Ede, Thijs and Aghakhani, Hojjat and Spahn, Noah and Bortolameotti, Riccardo and Cova, Marco and Continella, Andrea and van Steen, Maarten and Peter, Andreas and Kruegel, Christopher and Vigna, Giovanni , booktitle =
-
[15]
and Gjomemo, Rigel and Eshete, Birhanu and Sekar, R
Milajerdi, Sadegh M. and Gjomemo, Rigel and Eshete, Birhanu and Sekar, R. and Venkatakrishnan, V. N. , booktitle =
-
[16]
Ali, Tarek and Kostakos, Panos , year =. 2309.16021 , archivePrefix =
-
[17]
Qi, Jiaxing and Huang, Shaohan and Luan, Zhongzhi and Fung, Carol and Yang, Hailong and Qian, Depei , year =. 2309.01189 , archivePrefix =
-
[18]
Albanese, Massimiliano and Ou, Xinming and Lybarger, Kevin and Lende, Daniel and Goldgof, Dmitry , year =. Towards. 2505.06394 , archivePrefix =
-
[19]
Sifting the Noise: A Comparative Study of
Xiong, Yunpeng and Zhang, Ting , year =. Sifting the Noise: A Comparative Study of. 2601.22952 , archivePrefix =
-
[20]
Before You Hand Over the Wheel: Evaluating
Jajodia, Sourov and Sultana, Madeena and Majumdar, Suryadipta and Taylor, Adrian and Vandenberghe, Grant , year =. Before You Hand Over the Wheel: Evaluating. 2603.06422 , archivePrefix =
-
[21]
2023 , eprint =
Automatic Root Cause Analysis via Large Language Models for Cloud Incidents , author =. 2023 , eprint =
2023
-
[22]
Zhang, Dylan and Zhang, Xuchao and Bansal, Chetan and Las-Casas, Pedro and Fonseca, Rodrigo and Rajmohan, Saravan , year =. 2309.05833 , archivePrefix =
-
[23]
Zhang, Jie and Bu, Haoyu and Wen, Hui and Liu, Yongji and Fei, Haiqiang and Xi, Rongrong and Li, Lun and Yang, Yun and Zhu, Hongsong and Meng, Dan , year =. When. 2405.03644 , archivePrefix =
-
[24]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[25]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Large Language Models are Zero-Shot Reasoners , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[26]
International Conference on Learning Representations (ICLR) , year =
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author =. International Conference on Learning Representations (ICLR) , year =
-
[27]
Challenging
Suzgun, Mirac and Scales, Nathan and Sch. Challenging. 2022 , eprint =
2022
-
[28]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =
2023
-
[29]
Xiong, Miao and Hu, Zhiyuan and Lu, Xinyang and Li, Yifei and Fu, Jie and He, Junxian and Hooi, Bryan , booktitle =. Can
-
[30]
International Conference on Learning Representations (ICLR) , year =
Large Language Models Are Human-Level Prompt Engineers , author =. International Conference on Learning Representations (ICLR) , year =
-
[31]
International Conference on Learning Representations (ICLR) , year =
Large Language Models as Optimizers , author =. International Conference on Learning Representations (ICLR) , year =
-
[32]
International Conference on Learning Representations (ICLR) , year =
Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers , author =. International Conference on Learning Representations (ICLR) , year =
-
[33]
and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =
Khattab, Omar and Singhvi, Arnav and Maheshwari, Paridhi and Zhang, Zhiyuan and Santhanam, Keshav and Vardhamanan, Sri and Haq, Saiful and Sharma, Ashutosh and Joshi, Thomas T. and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =. 2310.03714 , archivePrefix =
-
[34]
Agrawal, Lakshya A. and Tan, Shangyin and Soylu, Dilara and Ziems, Noah and Khare, Rishi and Opsahl-Ong, Krista and Singhvi, Arnav and Shandilya, Herumb and Ryan, Michael J. and Jiang, Meng and Potts, Christopher and Sen, Koushik and Dimakis, Alexandros G. and Stoica, Ion and Klein, Dan and Zaharia, Matei and Khattab, Omar , year =. 2507.19457 , archivePrefix =
-
[35]
Zelikman, Eric and Harik, Georges and Shao, Yijia and Jayasiri, Varuna and Haber, Nick and Goodman, Noah D. , year =. 2403.09629 , archivePrefix =
-
[36]
Hosseini, Arian and Yuan, Xingdi and Malkin, Nikolay and Courville, Aaron and Sordoni, Alessandro and Agarwal, Rishabh , year =. 2402.06457 , archivePrefix =
-
[37]
Gulcehre, Caglar and Paine, Tom Le and Srinivasan, Srivatsan and Konyushkova, Ksenia and Weerts, Lotte and Sharma, Abhishek and Siddhant, Aditya and Ahern, Alex and Wang, Miaosen and Gu, Chenjie and Macherey, Wolfgang and Doucet, Arnaud and Firat, Orhan and de Freitas, Nando , year =. Reinforced Self-Training (. 2308.08998 , archivePrefix =
-
[38]
Transactions on Machine Learning Research , year =
Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models , author =. Transactions on Machine Learning Research , year =
-
[39]
2023 , eprint =
Scaling Relationship on Learning Mathematical Reasoning with Large Language Models , author =. 2023 , eprint =
2023
-
[40]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Training Chain-of-Thought via Latent-Variable Inference , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[41]
Koh, Woosung and Oh, Wonbeen and Jang, Jaein and Lee, MinHyung and Kim, Hyeongjin and Kim, Ah Yeon and Kim, Joonkee and Lee, Junghyun and Kim, Taehyeon and Yun, Se-Young , booktitle =
-
[42]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Training Language Models to Follow Instructions with Human Feedback , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[43]
2017 , eprint =
Proximal Policy Optimization Algorithms , author =. 2017 , eprint =
2017
-
[44]
Lambert, Nathan and Morrison, Jacob and Pyatkin, Valentina and Huang, Shengyi and Ivison, Hamish and Brahman, Faeze and Miranda, Lester James V. and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D. and Yang, Jiangjiang and Le Bras, Ronan and Tafjord, Oyvind and Wilhelm, Chris and Soldaini, L...
-
[45]
Back to Basics: Revisiting
Ahmadian, Arash and Cremer, Chris and Gall. Back to Basics: Revisiting. Annual Meeting of the Association for Computational Linguistics (ACL) , year =
-
[46]
Advances in Large Margin Classifiers , pages =
Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods , author =. Advances in Large Margin Classifiers , pages =
-
[47]
International Conference on Machine Learning (ICML) , year =
On Calibration of Modern Neural Networks , author =. International Conference on Machine Learning (ICML) , year =
-
[48]
2022 , eprint =
Language Models (Mostly) Know What They Know , author =. 2022 , eprint =
2022
-
[49]
Transactions on Machine Learning Research , year =
Teaching Models to Express Their Uncertainty in Words , author =. Transactions on Machine Learning Research , year =
-
[50]
International Conference on Learning Representations (ICLR) , year =
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , author =. International Conference on Learning Representations (ICLR) , year =
-
[51]
2021 , eprint =
Training Verifiers to Solve Math Word Problems , author =. 2021 , eprint =
2021
-
[52]
and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =
Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.