Pith. sign in

REVIEW 3 major objections 6 minor 52 references

A chain-of-thought triage model trained on real SOC detections, paired with a separate calibrator that reads its reasoning, sharply raises high-confidence auto-triage recall over direct-label LLMs.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:39 UTC pith:3FJTOY4E

load-bearing objection Solid industrial CoT+calibrator triage paper with real SOC data and clean ablations; the headline +43 pp benign auto-close gain is overstated once you respect the validation-to-test precision drift the authors themselves report. the 3 major comments →

arxiv 2607.28460 v1 pith:3FJTOY4E submitted 2026-07-30 cs.LG cs.CR

Cybersecurity Detection Classification with Reasoning-enabled Language Models

classification cs.LG cs.CR
keywords cybersecurity triagesecurity operations centerchain-of-thoughtlarge language modelsconfidence calibrationreinforcement learning with verifiable rewardsalert fatigueendpoint detections
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Security operations centers drown in more endpoint detections than analysts can review, so automated triage must both classify well and know when it is sure enough to act. This paper argues that training a language model to reason step by step about whether a Windows detection is truly malicious or benign—via prompt optimization, self-training, and reinforcement learning with checkable rewards—beats the usual practice of training the model only to spit out a label. Because that reasoning path ruins the usefulness of ordinary label-token probabilities as confidence, the authors train a second model that reads the full reasoning trace and estimates whether the verdict is correct. On a large human-labeled test set the full system reaches 82.6% accuracy and, at the high-confidence bar that would govern auto-close and auto-prioritize decisions, recovers far more true benign and true malicious cases than a direct-label baseline. They also show that an untrained confidence judge is useless at that bar, and that a specialized 30B model outperforms much larger general-purpose models that were not trained for this task.

Core claim

Explicitly training a language model to produce chain-of-thought reasoning about real Windows endpoint detections, then pairing it with a separately trained calibrator that scores the full reasoning trace, yields 82.6% test accuracy and improves high-confidence benign recall by 43.0 percentage points and malicious recall by 18.3 percentage points over a conventional direct-label LLM classifier; without the trained calibrator, high-confidence recall collapses to zero.

What carries the argument

A four-stage recipe—automated prompt optimization, self-training that keeps only correct or genuinely rationalized reasoning traces, reinforcement learning with a binary format-and-correctness reward, and a second model that maps detection plus reasoning plus label to a correct/incorrect probability—together produce both the verdict and the actionable confidence.

Load-bearing premise

High-precision confidence thresholds chosen on validation, and a calibrator fit on training rollouts, will still mark the right detections as safe to auto-act on when the data comes from a later time period.

What would settle it

On a fresh, later time window of the same SOC stream, measure whether the validation-chosen high-confidence thresholds still deliver near the claimed ~98–99% precision with non-trivial recall; a large drop in precision or a collapse of high-confidence recall would falsify the automation claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • High-confidence tiers of endpoint detections can be auto-closed or auto-prioritized at substantially higher coverage without raising analyst false-positive load.
  • Label-token probabilities alone are not a usable confidence signal once chain-of-thought is used; a trained judge over the reasoning trace becomes mandatory for precision-targeted automation.
  • Each training stage (prompt optimization, self-training, RL with verifiable rewards) compounds gains at the calibrated high-confidence operating point rather than only lifting raw accuracy.
  • A domain-finetuned 30B reasoning model can outperform larger frontier general-purpose models on this SOC task, favoring targeted training over raw scale.
  • Analysts can be shown accurate reasoning traces on medium-confidence cases while machines handle the high-confidence tail, stretching limited human triage capacity.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same calibrator-over-reasoning pattern is likely required for any high-stakes CoT classifier where auto-action depends on precision targets, not only cybersecurity.
  • Temporal drift of the benign high-precision point suggests production systems will need scheduled recalibration or online threshold monitoring, not a one-time fit.
  • Extending the recipe beyond binary Windows endpoint labels to multi-class triage or other sensor platforms is the natural operational next test of whether the gains are architecture-general.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper trains a chain-of-thought triage classifier on real, human-labeled Windows endpoint detections using a four-stage recipe (GEPA prompt optimization, AdaSTaR-style self-training, GRPO with format-and-correctness rewards, and a separate LoRA calibrator that reads the full reasoning trace). On a temporally held-out test set the system reaches 82.6% accuracy and, at validation-chosen high-confidence thresholds, reports large lifts over a direct-label SFT baseline in benign and malicious recall. Ablations show that each training stage helps, that an untrained confidence judge collapses high-confidence recall to zero, and that a finetuned 30B model beats stronger zero-shot frontier models on this task.

Significance. If the high-confidence automation claims hold under the SOC’s actual precision constraints, the work is practically important: it targets a real bottleneck (alert fatigue), uses large industrial labels with a temporal split, and shows that CoT plus a trained trace-level calibrator can materially expand the auto-close and prioritization tiers. The factorial ablation establishing that an untrained judge is unusable, and the demonstration that domain finetuning of a 30B model beats larger general-purpose models, are concrete contributions. The pipeline is clearly specified enough to be reproduced in similar industrial settings.

major comments (3)
  1. [Table 1; Results / Overall Metrics] Table 1, Benign High row: the central automation claim (+43.0 pp benign recall at the tier that “governs automated triage”) is defined by validation thresholds targeting 98% precision, but on the temporally later test set those same thresholds yield only 90.8% precision (vs 98.0% on validation), while the direct-label baseline stays at 98.3%. The paper attributes the drop to a 4.4-point policy accuracy gap and greater calibrator distribution sensitivity. Headline benign recall is therefore not measured at the SOC-required operating point. Please report benign (and malicious) recall after re-selecting or adjusting the threshold on test (or via a nested protocol) so that test precision actually meets 98%/99%, and state whether a material gain over the baseline’s 21.8% remains. Without that, the largest claimed improvement is not established for deployment.
  2. [Table 2; Methods / Confidence Calibration; Experiments / Training] Table 2 and Confidence Calibration: each row’s calibrator is trained independently and the checkpoint is chosen to maximize benign validation recall at 98% precision. That selection criterion directly optimizes the primary reported metric and, together with training the calibrator on the same policy’s rollouts, can inflate the apparent stage-wise gains and the necessity claim. Please add (i) a fixed early/mid/late checkpoint comparison or selection on a metric that is not the test headline (e.g., ECE or a held-out calibration split), and (ii) results when the calibrator is trained only on an earlier policy or on a disjoint detection subset, so that policy and calibrator improvements are not confounded.
  3. [Table 1; Table 3; Results] The direct-label SFT baseline isolates “no CoT,” but does not control for total training compute, number of labeled passes, or access to a second model. A stronger control would be a direct-label model given comparable GRPO steps and/or a calibrator that sees only the detection (no CoT). Without that, part of the +43/+18 pp lift may be attributable to extra optimization and the second network rather than to reasoning per se. A compact compute-matched or “calibrator-on-detection-only” ablation would make the CoT contribution load-bearing rather than suggestive.
minor comments (6)
  1. [Abstract; Introduction] Abstract and Introduction state “improves benign recall by 43.0% and malicious recall by 18.3%”; these are percentage points, not relative percent. Align wording with Table 1.
  2. [Figure 3] Figure 3 is helpful but the abridged CoT and single example cannot support generality; consider adding failure cases (high-confidence incorrect) in the appendix.
  3. [Methods; Appendix prompts] Label vocabulary shifts between TRUE_POSITIVE/FALSE_POSITIVE in prompts and malicious/benign in the main text; a single mapping note would reduce confusion.
  4. [Table 4] Table 4 uses a 1,000-sample subset at T=0 with no confidence threshold; state how the subset was drawn and whether ranks change on the full test set.
  5. [Table 1] Medium and low tiers are often identical because of bimodal confidence; that is noted in text but should also be flagged in Table 1’s caption so readers do not over-interpret tier granularity.
  6. [Related Work] Related Work could briefly contrast with other SOC LLM triage / false-positive filtering agents beyond the cited IT/cloud RCA systems, to sharpen the “first CoT-trained detection triage” claim.

Circularity Check

0 steps flagged

No circularity: empirical SOC triage pipeline evaluated against external human labels on held-out temporal splits.

full rationale

The paper’s load-bearing claims are measured performance numbers (82.6% accuracy; high-confidence benign/malicious recall lifts vs a direct-label SFT baseline) on real, human-labeled Windows endpoint detections with train/val/test collected over consecutive time windows. The CoT policy is optimized with GEPA, AdaSTaR-style self-training, and GRPO under a verifiable reward that is 1 iff the parsed label matches the external ground-truth label and the format is valid—not a quantity defined from the model’s own outputs. The calibrator is supervised on correct/incorrect relative to those same external labels and is scored by precision-targeted recall on validation, then frozen for test; that is ordinary post-hoc thresholding, not a fitted constant renamed as a first-principles prediction. Ablations (untrained judge → zero high-confidence recall; stage-wise gains; frontier zero-shot baselines) are comparative experiments, not self-definitional reductions. Citations (GEPA, AdaSTaR, GRPO, Nemotron, classical calibration) are external methods, not author uniqueness theorems that force the result. Distribution-shift fragility of the benign 98% operating point is a validity concern, not circularity. No step reduces Eq./claim X to its own inputs by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 3 invented entities

Load-bearing premises are standard ML/SOC assumptions plus industrial data and tooling choices, not new physical entities. Free parameters are training and thresholding choices selected on validation. No invented particles or forces; invented artifacts are the trained policy, calibrator, and optimized prompt.

free parameters (4)
  • High/medium/low precision targets (benign 98/95/90%, malicious 99/90/80%) = 98/95/90 benign; 99/90/80 malicious
    Operating points chosen with SOC analysts; they define the headline recall metrics and are not derived from first principles.
  • Calibrator checkpoint selection criterion = max val benign recall@98% precision
    Checkpoint chosen to maximize benign recall at 98% precision on validation among samples labeled correct—directly tunes the main reported operating point.
  • GEPA/AdaSTaR/GRPO/LoRA hyperparameters = as listed in Methods/Appendix
    Search budget, LoRA ranks, learning rates, rollout counts, reward format, and iteration/step selection (e.g., AdaSTaR iter 25, GRPO step 8000) are hand-set and validation-selected.
  • Instruction-quality judge dimensions and min-aggregation = metric = accuracy × min(five scores)
    Five subjective 0–1 scores multiplied via min with accuracy steer prompt search away from brittle thresholds; weights/structure are design choices.
axioms (5)
  • domain assumption Human SOC labels on Windows endpoint detections are a valid ground truth for malicious vs benign triage.
    All accuracy, reward, and calibrator supervision rest on these labels (Task Formulation, Experiments).
  • domain assumption Temporal consecutive week splits (8/2/1) adequately represent deployment distribution for claimed automation metrics.
    Test is later in time; paper notes benign precision drift, so the axiom is partially stressed but still assumed for headline claims.
  • ad hoc to paper Binary format-and-correctness reward (1 iff well-formed CoT+label and label matches) is a sufficient verifiable reward for useful triage reasoning.
    GRPO reward definition; no process reward or analyst preference model.
  • domain assumption A separate LLM judge of correctness from (detection, CoT, label) yields actionable confidence via softmax of correct/incorrect tokens.
    Standard trained-probe calibration assumption (Kadavath/Cobbe lineage), load-bearing for high-confidence tiers.
  • standard math Open-weight Nemotron-3 Nano/Super and GEPA/AdaSTaR/GRPO behave as described in cited works under the authors’ modifications.
    Background tooling and algorithms treated as given.
invented entities (3)
  • CoT triage policy (finetuned Nemotron-3-Nano-30B) no independent evidence
    purpose: Emit <think>…</think>label reasoning and verdict for detections.
    Primary trained artifact; not a new scientific entity beyond a finetuned model.
  • Trace-level correctness calibrator (second Nano-30B) no independent evidence
    purpose: Map detection+CoT+label to P(correct) for precision-targeted tiers.
    Introduced because CoT destroys label-token confidence; evidence is internal ablations only.
  • GEPA-optimized Charlotte AI system prompt with LLM-as-judge guardrails no independent evidence
    purpose: Avoid brittle numeric decision-tree prompts while retaining multi-field reasoning.
    Prompt artifact specific to this pipeline.

pith-pipeline@v1.2.0-daily-grok45 · 21248 in / 3673 out tokens · 68342 ms · 2026-07-31T06:39:07.734916+00:00 · methodology

0 comments
read the original abstract

A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards. We find that CoT reasoning also degrades the label-token probabilities that automated triage relies on, so we separately train a calibrator that reads the full reasoning trace and estimates the probability that the verdict is correct. Our system reaches 82.6% test accuracy and, at the high-confidence operating point that governs automated triage, improves benign recall by 43.0% and malicious recall by 18.3% over a direct-label LLM classifier. We further show that the trained calibrator is necessary - an untrained confidence judge collapses high-confidence recall to zero - and that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.

Figures

Figures reproduced from arXiv: 2607.28460 by Alexandru Apostu, Amol Khanna, Chase Helwig, Chase Midler, Cristian Viorel Popa, Diana Bolocan, Edward Raff, Joan Pujol-Roig, Laura Vasilie, Manu Nandan, Michael Brautbar, Mihaela Gaman, Sven Krasser.

Figure 1
Figure 1. Figure 1: Self-training dynamics across AdaSTaR iterations. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: GRPO reinforcement-learning dynamics over 8,000 training steps; all panels share the training step on the x-axis. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: A representative test-set detection (reasoning trace abridged). The classifier dismisses a severe-sounding [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: RLVR without self-training: GRPO from the [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

52 extracted references · 15 linked inside Pith

  1. [1]

    2017 , eprint =

    Attention Is All You Need , author =. 2017 , eprint =

  2. [2]

    , year =

    Zelikman, Eric and Wu, Yuhuai and Mu, Jesse and Goodman, Noah D. , year =. 2203.14465 , archivePrefix =

  3. [3]

    Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Bi, Xiao and Zhang, Haowei and Zhang, Mingchuan and Li, Y. K. and Wu, Y. and Guo, Daya , year =. 2402.03300 , archivePrefix =

  4. [4]

    and Axon, Louise and Martinovic, Ivan , booktitle =

    Alahmadi, Bushra A. and Axon, Louise and Martinovic, Ivan , booktitle =. 99\. 2022 , publisher =

  5. [5]

    Hassan, Wajih Ul and Guo, Shengjian and Li, Ding and Chen, Zhengzhang and Jee, Kangkook and Li, Zhichun and Bates, Adam , booktitle =

  6. [6]

    Proceedings of the 13th USENIX Conference on System Administration (LISA '99) , pages =

    Snort: Lightweight Intrusion Detection for Networks , author =. Proceedings of the 13th USENIX Conference on System Administration (LISA '99) , pages =. 1999 , publisher =

  7. [7]

    2010 IEEE Symposium on Security and Privacy , pages =

    Outside the Closed World: On Using Machine Learning for Network Intrusion Detection , author =. 2010 IEEE Symposium on Security and Privacy , pages =. 2010 , publisher =

  8. [8]

    2019 , publisher =

    Pendlebury, Feargus and Pierazzi, Fabio and Jordaney, Roberto and Kinder, Johannes and Cavallaro, Lorenzo , booktitle =. 2019 , publisher =

  9. [9]

    2025 , eprint =

    Large Language Models for Security Operations Centers: A Comprehensive Survey , author =. 2025 , eprint =

  10. [10]

    Labeling

    Daniel, Nir and Kaiser, Florian Klaus and Giladi, Shay and Sharabi, Sapir and Moyal, Raz and Shpolyansky, Shalev and Murillo, Andres and Elyashar, Aviad and Puzis, Rami , year =. Labeling. 2412.10978 , archivePrefix =

  11. [11]

    ACM Transactions on Information and System Security , volume =

    Clustering Intrusion Detection Alarms to Support Root Cause Analysis , author =. ACM Transactions on Information and System Security , volume =

  12. [12]

    Recent Advances in Intrusion Detection (RAID) , year =

    Using Adaptive Alert Classification to Reduce False Positives in Intrusion Detection , author =. Recent Advances in Intrusion Detection (RAID) , year =

  13. [13]

    Du, Min and Li, Feifei and Zheng, Guineng and Srikumar, Vivek , booktitle =

  14. [14]

    van Ede, Thijs and Aghakhani, Hojjat and Spahn, Noah and Bortolameotti, Riccardo and Cova, Marco and Continella, Andrea and van Steen, Maarten and Peter, Andreas and Kruegel, Christopher and Vigna, Giovanni , booktitle =

  15. [15]

    and Gjomemo, Rigel and Eshete, Birhanu and Sekar, R

    Milajerdi, Sadegh M. and Gjomemo, Rigel and Eshete, Birhanu and Sekar, R. and Venkatakrishnan, V. N. , booktitle =

  16. [16]

    2309.16021 , archivePrefix =

    Ali, Tarek and Kostakos, Panos , year =. 2309.16021 , archivePrefix =

  17. [17]

    2309.01189 , archivePrefix =

    Qi, Jiaxing and Huang, Shaohan and Luan, Zhongzhi and Fung, Carol and Yang, Hailong and Qian, Depei , year =. 2309.01189 , archivePrefix =

  18. [18]

    Albanese, Massimiliano and Ou, Xinming and Lybarger, Kevin and Lende, Daniel and Goldgof, Dmitry , year =. Towards. 2505.06394 , archivePrefix =

  19. [19]

    Sifting the Noise: A Comparative Study of

    Xiong, Yunpeng and Zhang, Ting , year =. Sifting the Noise: A Comparative Study of. 2601.22952 , archivePrefix =

  20. [20]

    Before You Hand Over the Wheel: Evaluating

    Jajodia, Sourov and Sultana, Madeena and Majumdar, Suryadipta and Taylor, Adrian and Vandenberghe, Grant , year =. Before You Hand Over the Wheel: Evaluating. 2603.06422 , archivePrefix =

  21. [21]

    2023 , eprint =

    Automatic Root Cause Analysis via Large Language Models for Cloud Incidents , author =. 2023 , eprint =

  22. [22]

    2309.05833 , archivePrefix =

    Zhang, Dylan and Zhang, Xuchao and Bansal, Chetan and Las-Casas, Pedro and Fonseca, Rodrigo and Rajmohan, Saravan , year =. 2309.05833 , archivePrefix =

  23. [23]

    Zhang, Jie and Bu, Haoyu and Wen, Hui and Liu, Yongji and Fei, Haiqiang and Xi, Rongrong and Li, Lun and Yang, Yun and Zhu, Hongsong and Meng, Dan , year =. When. 2405.03644 , archivePrefix =

  24. [24]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  25. [25]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Large Language Models are Zero-Shot Reasoners , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  26. [26]

    International Conference on Learning Representations (ICLR) , year =

    Self-Consistency Improves Chain of Thought Reasoning in Language Models , author =. International Conference on Learning Representations (ICLR) , year =

  27. [27]

    Challenging

    Suzgun, Mirac and Scales, Nathan and Sch. Challenging. 2022 , eprint =

  28. [28]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

    Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback , author =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) , year =

  29. [29]

    Xiong, Miao and Hu, Zhiyuan and Lu, Xinyang and Li, Yifei and Fu, Jie and He, Junxian and Hooi, Bryan , booktitle =. Can

  30. [30]

    International Conference on Learning Representations (ICLR) , year =

    Large Language Models Are Human-Level Prompt Engineers , author =. International Conference on Learning Representations (ICLR) , year =

  31. [31]

    International Conference on Learning Representations (ICLR) , year =

    Large Language Models as Optimizers , author =. International Conference on Learning Representations (ICLR) , year =

  32. [32]

    International Conference on Learning Representations (ICLR) , year =

    Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers , author =. International Conference on Learning Representations (ICLR) , year =

  33. [33]

    and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =

    Khattab, Omar and Singhvi, Arnav and Maheshwari, Paridhi and Zhang, Zhiyuan and Santhanam, Keshav and Vardhamanan, Sri and Haq, Saiful and Sharma, Ashutosh and Joshi, Thomas T. and Moazam, Hanna and Miller, Heather and Zaharia, Matei and Potts, Christopher , year =. 2310.03714 , archivePrefix =

  34. [34]

    and Tan, Shangyin and Soylu, Dilara and Ziems, Noah and Khare, Rishi and Opsahl-Ong, Krista and Singhvi, Arnav and Shandilya, Herumb and Ryan, Michael J

    Agrawal, Lakshya A. and Tan, Shangyin and Soylu, Dilara and Ziems, Noah and Khare, Rishi and Opsahl-Ong, Krista and Singhvi, Arnav and Shandilya, Herumb and Ryan, Michael J. and Jiang, Meng and Potts, Christopher and Sen, Koushik and Dimakis, Alexandros G. and Stoica, Ion and Klein, Dan and Zaharia, Matei and Khattab, Omar , year =. 2507.19457 , archivePrefix =

  35. [35]

    , year =

    Zelikman, Eric and Harik, Georges and Shao, Yijia and Jayasiri, Varuna and Haber, Nick and Goodman, Noah D. , year =. 2403.09629 , archivePrefix =

  36. [36]

    2402.06457 , archivePrefix =

    Hosseini, Arian and Yuan, Xingdi and Malkin, Nikolay and Courville, Aaron and Sordoni, Alessandro and Agarwal, Rishabh , year =. 2402.06457 , archivePrefix =

  37. [37]

    Reinforced Self-Training (

    Gulcehre, Caglar and Paine, Tom Le and Srinivasan, Srivatsan and Konyushkova, Ksenia and Weerts, Lotte and Sharma, Abhishek and Siddhant, Aditya and Ahern, Alex and Wang, Miaosen and Gu, Chenjie and Macherey, Wolfgang and Doucet, Arnaud and Firat, Orhan and de Freitas, Nando , year =. Reinforced Self-Training (. 2308.08998 , archivePrefix =

  38. [38]

    Transactions on Machine Learning Research , year =

    Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models , author =. Transactions on Machine Learning Research , year =

  39. [39]

    2023 , eprint =

    Scaling Relationship on Learning Mathematical Reasoning with Large Language Models , author =. 2023 , eprint =

  40. [40]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Training Chain-of-Thought via Latent-Variable Inference , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  41. [41]

    Koh, Woosung and Oh, Wonbeen and Jang, Jaein and Lee, MinHyung and Kim, Hyeongjin and Kim, Ah Yeon and Kim, Joonkee and Lee, Junghyun and Kim, Taehyeon and Yun, Se-Young , booktitle =

  42. [42]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Training Language Models to Follow Instructions with Human Feedback , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  43. [43]

    2017 , eprint =

    Proximal Policy Optimization Algorithms , author =. 2017 , eprint =

  44. [44]

    and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D

    Lambert, Nathan and Morrison, Jacob and Pyatkin, Valentina and Huang, Shengyi and Ivison, Hamish and Brahman, Faeze and Miranda, Lester James V. and Liu, Alisa and Dziri, Nouha and Lyu, Shane and Gu, Yuling and Malik, Saumya and Graf, Victoria and Hwang, Jena D. and Yang, Jiangjiang and Le Bras, Ronan and Tafjord, Oyvind and Wilhelm, Chris and Soldaini, L...

  45. [45]

    Back to Basics: Revisiting

    Ahmadian, Arash and Cremer, Chris and Gall. Back to Basics: Revisiting. Annual Meeting of the Association for Computational Linguistics (ACL) , year =

  46. [46]

    Advances in Large Margin Classifiers , pages =

    Probabilistic Outputs for Support Vector Machines and Comparisons to Regularized Likelihood Methods , author =. Advances in Large Margin Classifiers , pages =

  47. [47]

    International Conference on Machine Learning (ICML) , year =

    On Calibration of Modern Neural Networks , author =. International Conference on Machine Learning (ICML) , year =

  48. [48]

    2022 , eprint =

    Language Models (Mostly) Know What They Know , author =. 2022 , eprint =

  49. [49]

    Transactions on Machine Learning Research , year =

    Teaching Models to Express Their Uncertainty in Words , author =. Transactions on Machine Learning Research , year =

  50. [50]

    International Conference on Learning Representations (ICLR) , year =

    Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation , author =. International Conference on Learning Representations (ICLR) , year =

  51. [51]

    2021 , eprint =

    Training Verifiers to Solve Math Word Problems , author =. 2021 , eprint =

  52. [52]

    and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =

    Hu, Edward J. and Shen, Yelong and Wallis, Phillip and Allen-Zhu, Zeyuan and Li, Yuanzhi and Wang, Shean and Wang, Lu and Chen, Weizhu , booktitle =