Pith. sign in

REVIEW 3 major objections 5 minor 96 references

Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that advisory prompts distilled from a model's own thinking traces recover much of thinking mode's accuracy at 58.6% fewer output tokens.

desk verdict Clever, honest, under-ablated: the paper's held-out gains are plausible but the reasoning-trace mechanism is unproven. read the letter →

arxiv 2608.00437 v1 pith:4FJSJ2WL submitted 2026-08-01 cs.SE cs.AI

classification cs.SEcs.AI
keywords advisorypromptspromptdistillationreasoningtracesthinkingmodecodegenerationsoftwareengineeringtaskscommon-modefailuressmalllanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Language models that can "think" before answering make fewer coding mistakes, but thinking costs tokens. This paper asks whether the lessons hidden in those thinking traces can be compressed into a short advisory prompt that the model follows without thinking. It proposes an inference-only pipeline: find cases where thinking flips a small model's answer from wrong to right (or right to wrong), have a larger model explain why, and turn those explanations into generic instructions to prepend to the task prompt. On held-out software-engineering tasks, the distilled prompt improves non-thinking accuracy in 19 of 20 model-task comparisons (mean +3.1 points) while using 58.6% fewer output tokens than full thinking. The gains do not fully match thinking mode, but the method is training-free and works on modestly sized models.

What carries the argument

The load-bearing mechanism is the paired correctness delta between a hybrid model's thinking and non-thinking modes on the same example. The delta identifies where explicit reasoning changes the outcome; the paired responses and reasoning trace give a teacher model evidence for why, and the teacher's diagnoses are aggregated into improvement and guard instructions that form candidate advisory prompts. The selected prompt is fixed across the whole task, not generated per problem, and it preserves the original task placeholders.

What would settle it

Run the pipeline again, but feed the teacher the correct and incorrect answers without the reasoning trace; if the resulting prompts produce the same held-out gains, the reasoning trace is not carrying the effect. Likewise, compare the distilled prompt against a generic 'think carefully' control prompt: if the control matches its gains, the specific diagnoses are not carrying the effect.

Watch

Extended reading notes

Core claim

The central discovery is that a fixed, task-level advisory prompt can recover a meaningful share of the accuracy benefit of thinking mode without paying for the full reasoning trace. The authors define a correctness delta for each example: the same student model passes under thinking but fails without it (improvement case), or passes without thinking but fails with it (regression case). A teacher model diagnoses each delta and aggregates the diagnoses into two reusable instructions, an improvement instruction and a guard instruction, which are composed with the original prompt and selected on a validation split. This procedure is purely inference-time, with no weight updates. Across 20 model-task comparisons, the distilled non-thinking prompt beats the original non-thinking baseline in 19 cases and cuts output tokens by 58.6% on average; it does not beat the thinking-mode baseline, which remains more accurate in 17 of 20 cases.

Load-bearing premise

The teacher's post-hoc natural-language diagnoses accurately capture why thinking mode helped or hurt, even though those diagnoses are not independently validated and are not ablated from the pipeline.

Editorial extensions

If this is right

  • Adding a distilled advisory prompt to a small hybrid model is a cheap, training-free way to raise accuracy on software-engineering tasks in non-thinking mode.
  • A distilled prompt does not replace thinking mode: thinking still outperforms distilled non-thinking inference in most comparisons, so the practical operating point is a cheaper default with escalation to thinking for hard cases.
  • Distilled prompts transfer between models on three of five task families, but transfer is directional and can hurt, so a prompt should be validated on the target model before deployment.
  • Shared failure between two models does not by itself predict whether one model's distilled prompt will help the other, because shared failures are rescued less often than target-only failures.
  • The token savings are not merely shorter outputs: in most held-out sets, distilled non-thinking outputs are actually longer than the non-thinking baseline, suggesting the prompt elicits brief visible working instead of a full trace.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not test is applying the same delta-mining procedure to non-code reasoning tasks with a pass/fail oracle, such as math word problems; the mechanism only requires a hybrid student and a larger teacher.
  • The authors' results hint that guidance with discrete semantic checks, such as type compatibility and API existence, transfers better across models, but a controlled experiment varying prompt content would be needed to confirm this.
  • If the teacher's diagnoses are later shown to be faithful, the method suggests a broader principle: a fixed verbal rule can act as amortized reasoning, trading per-instance deliberation for a one-time distilled summary.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes an inference-only, training-free pipeline that distills the behavioral delta between a hybrid reasoning student model's thinking and non-thinking modes into a short advisory prompt. Using paired student responses and reasoning traces on cases where thinking flips correctness, a frontier teacher diagnoses each delta, aggregates diagnoses into improvement and guard instructions, and selects among k=3 candidate prompts on a validation split. The selected prompt is then evaluated on held-out test folds (or LiveCodeBench v1-v4 releases) on five software-engineering task families and four small student models. The main reported results are a +3.1 point mean non-thinking pass-rate gain in 19 of 20 model-task comparisons, a 58.6% average output-token saving relative to the thinking-mode baseline, smaller and less consistent thinking-mode gains, positive but directional cross-model transfer, and a weak overall association between pairwise failure correlation and transfer gain.

Significance. If the held-out gains hold, the complete pipeline is a practical, training-free way to improve small models on execution-checked software-engineering tasks, with a favorable accuracy-cost tradeoff on the Gemma students and a clean empirical study of cross-model transfer and common-mode failures. The held-out design is a genuine strength: distillation and validation sets are disjoint from test sets, the teacher never sees test examples, and results are reported on externally defined benchmarks. I also agree with the reader's assessment that the evaluation is not circular, since prompts are selected on validation and evaluated on held-out splits. However, the central mechanistic claim—that reasoning traces are the source of the gains—is not yet supported, because the pipeline is never ablated against a no-trace or generic-advice condition, and the teacher also receives the correct answer and both student responses; the paper's own Limitations concede this. Statistical reporting also needs tightening: one stochastic sample per example, only 10 of 20 comparisons nominally significant, and no multiple-comparison correction.

major comments (3)
  1. [Prompt Distillation Pipeline; Limitations] The central claim of the paper is that thinking-mode reasoning traces can be distilled into advisory prompts that transfer part of thinking's benefit to non-thinking inference. The pipeline, however, bundles the reasoning trace together with the correct answer, both student responses, the delta-mining criterion, the aggregation prompt, and validation-based selection; the teacher also sees the correct answer and is asked to produce summary instructions. The paper's own Limitations state that "we do not ablate the teacher's inputs, our results measure the effect of the complete pipeline rather than of the reasoning trace alone." This is load-bearing: the keyword audit shows that many distilled prompts are generic instructions (49% ask to track state, 40% to dry-run examples, 25% to check edge cases), so the observed +3.1 point non-thinking gain could plausibly be produced by a generic "be careful and check edge cases" prompt effect rather than by any information extracted from the trace. To support the title-level claim, the paper needs an ablation that removes the trace (e.g., the teacher receives paired responses and the correct answer but not the trace, or a trace-shuffled control) and a generic-advice baseline constructed without delta mining. Until then, RQ1/RQ2 establish only that the complete pipeline helps, not that reasoning-trace distillation is the mechanism.
  2. [Results, Table 1; Protocol] The empirical basis for RQ1 is the headline result of +3.1 points in 19 of 20 comparisons in Table 1. The protocol uses a single stochastic generation per example at temperature 0.7, and only 10 of the 20 comparisons are nominally significant under paired exact McNemar tests. With 20 comparisons and no multiple-comparison correction, a substantial fraction of the positive cells could be noise, and the mean effect is not accompanied by confidence intervals or an adjusted significance threshold. Please report corrected p-values (e.g., Benjamini-Hochberg), confidence intervals for the mean deltas, and either repeated sampling or a sensitivity analysis showing how many samples per example are needed to stabilize the binary pass-rate estimates. This is particularly important because the 20 cells are not independent: they share the same four student models and five task families.
  3. [Evaluation Metrics (RQ2); Discussion] The RQ2 token-savings claim needs a sharper decomposition. The paper reports that distilled non-thinking inference uses 58.6% fewer output tokens than the thinking-mode baseline, but the Discussion also states that in 89 of 116 model-specific held-out sets the distilled outputs are longer than the non-thinking baseline outputs. The headline saving is therefore driven by disabling thinking rather than by the distilled prompt shortening outputs; the distilled prompt adds 54.5 prompt-template tokens and, by the paper's own account, elicits visible working in the answer. Please report the full 2x2 table (baseline/distilled by thinking/non-thinking) for both accuracy and output tokens, and define the accuracy-cost operating point accordingly, separating the effect of the prompt from the effect of turning off the thinking mode.
minor comments (5)
  1. [Discussion] The exploratory analysis of 116 model-specific held-out sets reports a Spearman correlation of 0.39 between "the gain due to thinking per se" and the non-thinking distillation gain, but "gain due to thinking per se" is never formally defined; please define it explicitly (presumably bp0,1 minus bp0,0).
  2. [Table 1] The table caption does not define the abbreviation "savings / Γ"; the row would be clearer if it read "output-token savings / Γ (distilled non-thinking minus baseline thinking, percentage points)", matching the definitions in the Evaluation Metrics section.
  3. [Evaluation Metrics (RQ1, RQ2)] There is a stray closing parenthesis after "introduces regressions.)" in the paragraph defining thinking-mode gains; this appears to be a typo.
  4. [Prompt Distillation Pipeline] The design choice that the original template is not included as a candidate is mentioned only in a supplementary reference; since it affects the interpretation of validation selection, a sentence of rationale in the main text would be helpful.
  5. [General] No code, data, or full list of distilled prompts is provided in the manuscript. Given that the paper is an empirical study of a prompt-engineering pipeline, releasing the exact prompts, the teacher prompts, and the selection code would materially improve reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: candidate prompts are selected on a validation split and evaluated on held-out test data, so the reported gains are not forced by construction.

full rationale

The derivation chain is self-contained and non-circular. The pipeline separates distillation, validation, and test sets: 'Each dataset is split into disjoint distillation, validation, and test sets. Only the distillation set mines deltas and synthesizes prompt instructions; validation selects candidate prompts. The test set is reserved for final evaluation.' The selected prompt is chosen by validation pass rate and the original template is deliberately excluded as a candidate ('The original template is not included as a candidate'), so the RQ1/RQ2 comparisons of distilled versus baseline on held-out folds measure generalization rather than refitting. The token-savings figure (58.6%) is a directly measured comparison of emitted output-token counts between conditions, not a fitted parameter renamed as a prediction. The student's own reasoning trace is the intended input of the distillation mechanism, not a self-referential output. The paper's Limitations statement that 'we do not ablate the teacher's inputs, our results measure the effect of the complete pipeline rather than of the reasoning trace alone' is an internal-validity caveat about attributing the gain to the trace specifically; it does not show that any reported prediction is equivalent by construction to its inputs. References to prior work by the authors (e.g., Spiess, Devanbu, and Barr 2026) supply benchmark datasets and are not used as load-bearing justification of the method's claims. No uniqueness theorem or ansatz is imported from the authors' earlier work. Therefore the paper exhibits no circularity under the defined patterns.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the teacher's diagnostic validity, the representativeness of the correctness deltas, and single-sample stability. The selected prompt is a validation-fitted artifact rather than a theory-derived constant. No new physical or conceptual entities are introduced.

free parameters (5)
  • Selected advisory prompt pi* = not enumerated; one per model, task, fold
    Chosen from k=3 candidates by validation pass rate; data-dependent selection.
  • Candidate prompt count k = 3
    Hand-chosen; affects the diversity of the candidate set.
  • Sampling temperature = 0.7
    Hand-chosen for all student and teacher generations; no sweep reported.
  • Output token budget = 14,000 tokens
    Hand-chosen cap on student outputs; may truncate long thinking traces.
  • Data split ratios = 70/10/20
    Distillation, validation, and test proportions chosen by hand.
assumptions (5)
  • domain assumption Teacher diagnoses accurately explain why thinking mode helped or hurt.
    The mechanism relies on post-hoc natural-language diagnoses; the paper states they were not independently validated (Limitations).
  • domain assumption Correctness deltas mined on the distillation split are representative of held-out errors.
    Generalizing from an 80% split to test assumes error types recur; no distribution shift analysis is provided.
  • domain assumption Single-sample pass rate at temperature 0.7 is a stable performance estimate.
    The paper notes that a single stochastic generation may reflect sampling variability (Limitations).
  • domain assumption LiveCodeBench v5/v6 are free of training leakage for the student models.
    The paper chooses the latest releases 'to avoid training leakage'; contamination status is assumed.
  • domain assumption Perturbed CRUXEval variants measure execution reasoning reliably.
    Adopted from Spiess, Devanbu, and Barr (2026); the paper relies on that prior result.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks." pith.science (2026). https://pith.science/paper/4FJSJ2WL

@misc{pith2026260800437,
  author       = {Pith},
  title        = {Pith review of: Distilling Reasoning Traces into Advisory Prompts for Software Engineering Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4FJSJ2WL}},
  note         = {Machine review of arXiv:2608.00437}
}
read the original abstract

Language models are widely used for generating and otherwise processing code (e.g., identifying code hallucinations, possible inputs, or predicting outputs); however, LLMs can make mistakes, which can be serious. One key issue is that models are trained on (still) largely human-written, and thus imperfect, code; it's not easy to find sufficiently large code corpora that are entirely free of bugs. Thus, other inference-time ways of reducing LLM errors, without additional training, are desirable. "Reasoning" or "thinking" modes, exposed as a togglable feature by hybrid reasoning models, do reduce errors; however, reasoning consumes additional resources. This paper asks if better performance can be achieved without always incurring the cost of reasoning. Human students of programming learn to avoid mistakes by (a) identifying them, (b) reflecting upon the cognitive lapses that led to them (essentially, "thinking through" the errors), (c) inferring general rules or lessons from these reflections, and (d) internalizing these lessons into rules. In tutorial sessions with an instructor, this is a common Socratic interaction. Examples of such internalizable rules might include the nugget "Before coding, restate the requirements to clarify them." Inspired by this process, this paper describes an approach where we first identify examples in which "thinking mode" in a (low-resource) LLM avoids errors. These errors, and their avoidance via "thinking" in the same LLM, are then examined by a bigger LLM to generate summary explanations; these are then summarized by a large LLM into brief advisory prompts. This approach works on many modest-sized models; in some cases, the "advisory prompts" thus learned can also be gainfully transferred to other models. We also present investigations into the nature of coding errors that language models make, and a characterization of when this approach can be helpful.

Figures

Figures reproduced from arXiv: 2608.00437 by the authors.

Figure 1
Figure 1. Exception-prediction example: Stage 1 diagnoses a [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The prompt-distillation pipeline: correctness-delta [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 76 canonical work pages

  1. [1]

    Second Conference on Language Modeling , year =

    Pranjal Aggarwal and Sean Welleck , title =. Second Conference on Language Modeling , year =

  2. [2]

    International Conference on Learning Representations , year =

    Lakshya A Agrawal and Shangyin Tan and Dilara Soylu and others , title =. International Conference on Learning Representations , year =

  3. [3]

    International Conference on Learning Representations , year =

    Rico Angell and Jannik Brinkmann and He He , title =. International Conference on Learning Representations , year =

  4. [4]

    Claude 3.7 Sonnet System Card , year =

  5. [5]

    Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) , year =

    Sanket Badhe and Deep Shah , title =. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track) , year =

  6. [6]

    The Reversal Curse:

    Berglund, Lukas and Tong, Meg and Kaufmann, Max and others , booktitle=. The Reversal Curse:

  7. [7]

    Digest of Papers, 8th Annual International Symposium on Fault-Tolerant Computing (FTCS-8) , year =

    Liming Chen and Algirdas Avizienis , title =. Digest of Papers, 8th Annual International Symposium on Fault-Tolerant Computing (FTCS-8) , year =

  8. [8]

    Proceedings, Mining Software Repositories, 2026 , year=

    Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs , author=. Proceedings, Mining Software Repositories, 2026 , year=

Show all 96 references
  1. [9]

    2021 , eprint =

    Mark Chen and Jerry Tworek and Heewoo Jun and others , title =. 2021 , eprint =

  2. [10]

    2024 , eprint =

    Xingyu Chen and Jiahao Xu and Tian Liang and Zhiwei He and Jianhui Pang and Dian Yu and Linfeng Song and Qiuzhi Liu and Mengfei Zhou and Zhuosheng Zhang and Rui Wang and Zhaopeng Tu and Haitao Mi and Dong Yu , title =. 2024 , eprint =

  3. [11]

    2026 , eprint =

    Josef Chen , title =. 2026 , eprint =

  4. [12]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

    Daixuan Cheng and Shaohan Huang and Junyu Bi and Yuefeng Zhan and Jianfeng Liu and Yujing Wang and Hao Sun and Furu Wei and Weiwei Deng and Qi Zhang , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

  5. [13]

    Gonzalez , title =

    Alejandro Cuadron and Dacheng Li and Wenjie Ma and Xingyao Wang and Yichuan Wang and Siyuan Zhuang and Shu Liu and Luis Gaspar Schroeder and Tian Xia and Huanzhi Mao and Nicholas Thumiger and Aditya Desai and Ion Stoica and Ana Klimovic and Graham Neubig and Joseph E. Gonzalez...

  6. [14]

    Nature , volume =

    Daya Guo and Dejian Yang and Haowei Zhang and Junxiao Song and Peiyi Wang and Qihao Zhu and others , title =. Nature , volume =. 2025 , doi =

  7. [15]

    2023 , eprint =

    Yuntian Deng and Kiran Prasad and Roland Fernandez and Paul Smolensky and Vishrav Chaudhary and Stuart Shieber , title =. 2023 , eprint =

  8. [16]

    2025 , eprint =

    Aniket Didolkar and Nicolas Ballas and Sanjeev Arora and Anirudh Goyal , title =. 2025 , eprint =

  9. [17]

    Dyagin and Nikita I

    Ernest A. Dyagin and Nikita I. Kulin and Artur R. Khairullin and Viktor N. Zhuravlev and Alena N. Sitkina , title =. 2025 , eprint =

  10. [18]

    Eckhardt and Larry D

    Dave E. Eckhardt and Larry D. Lee , title =. IEEE Transactions on Software Engineering , year =

  11. [19]

    2024 , eprint =

    Aryaz Eghbali and Michael Pradel , title =. 2024 , eprint =

  12. [20]

    Advances in Neural Information Processing Systems 37 , year =

    Yao Fu and Dong-Ki Kim and Jaekyeom Kim and Sungryull Sohn and Lajanugen Logeswaran and Kyunghoon Bae and Honglak Lee , title =. Advances in Neural Information Processing Systems 37 , year =

  13. [21]

    Inverse Scaling in

    Aryo Pradipta Gema and Alexander H. Inverse Scaling in. Transactions on Machine Learning Research , year =

  14. [22]

    2607.02770 , archivePrefix =

    Gemma 4 Technical Report , year =. 2607.02770 , archivePrefix =

  15. [23]

    Great Models Think Alike and this Undermines

    Shashwat Goel and Joschka Str. Great Models Think Alike and this Undermines. Proceedings of the 42nd International Conference on Machine Learning , year =

  16. [24]

    Proceedings of the 41st International Conference on Machine Learning , year =

    Alex Gu and Baptiste Rozi. Proceedings of the 41st International Conference on Machine Learning , year =

  17. [25]

    Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , year =

    Priyanshu Gupta and Shashank Kirtania and Ananya Singha and Sumit Gulwani and Arjun Radhakrishna and Gustavo Soares and Sherry Shi , title =. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , year =

  18. [26]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

    Namgyu Ho and Laura Schmid and Se-Young Yun , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  19. [27]

    Wang and Chenhui Zhang and Zhangheng Li and Bo Li and Zhangyang Wang , title =

    Junyuan Hong and Jiachen T. Wang and Chenhui Zhang and Zhangheng Li and Bo Li and Zhangyang Wang , title =. The Twelfth International Conference on Learning Representations , year =

  20. [28]

    Bowman and Omer Levy , title =

    Or Honovich and Uri Shaham and Samuel R. Bowman and Omer Levy , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  21. [29]

    Findings of the Association for Computational Linguistics: ACL 2023 , year =

    Cheng-Yu Hsieh and Chun-Liang Li and Chih-kuan Yeh and others , title =. Findings of the Association for Computational Linguistics: ACL 2023 , year =

  22. [30]

    The Twelfth International Conference on Learning Representations , year =

    Jie Huang and Xinyun Chen and Swaroop Mishra and others , title =. The Twelfth International Conference on Learning Representations , year =

  23. [31]

    International Conference on Learning Representations , year =

    Naman Jain and King Han and Alex Gu and others , title =. International Conference on Learning Representations , year =

  24. [32]

    Jimenez and John Yang and Alexander Wettig and others , title =

    Carlos E. Jimenez and John Yang and Alexander Wettig and others , title =. The Twelfth International Conference on Learning Representations , year =

  25. [33]

    Joshi and Hanna Moazam and Heather Miller and Matei Zaharia and Christopher Potts , title =

    Omar Khattab and Arnav Singhvi and Paridhi Maheshwari and Zhiyuan Zhang and Keshav Santhanam and Sri Vardhamanan and Saiful Haq and Ashutosh Sharma and Thomas T. Joshi and Hanna Moazam and Heather Miller and Matei Zaharia and Christopher Potts , title =. International Conferen...

  26. [34]

    Proceedings of the 42nd International Conference on Machine Learning , year =

    Elliot Myunghoon Kim and Avi Garg and Kenny Peng and Nikhil Garg , title =. Proceedings of the 42nd International Conference on Machine Learning , year =

  27. [35]

    Knight and Nancy G

    John C. Knight and Nancy G. Leveson , title =. IEEE Transactions on Software Engineering , year =

  28. [36]

    Le and others , title =

    Derek Koh and Jinghui Mo and Benjamin H. Le and others , title =. 2026 , eprint =

  29. [37]

    Advances in Neural Information Processing Systems 35 , year =

    Takeshi Kojima and Shixiang Shane Gu and Machel Reid and Yutaka Matsuo and Yusuke Iwasawa , title =. Advances in Neural Information Processing Systems 35 , year =

  30. [38]

    Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

    Liunian Harold Li and Jack Hessel and Youngjae Yu and Xiang Ren and Kai-Wei Chang and Yejin Choi , title =. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  31. [39]

    Advances in Neural Information Processing Systems 38 , year =

    Xiaomin Li and Zhou Yu and Zhiwei Zhang and others , title =. Advances in Neural Information Processing Systems 38 , year =

  32. [40]

    Findings of the Association for Computational Linguistics: ACL 2026 , year =

    Mingqi Li and Karan Aggarwal and Yong Xie and Aitzaz Ahmad and Stephen Lau , title =. Findings of the Association for Computational Linguistics: ACL 2026 , year =

  33. [41]

    Miller , title =

    Bev Littlewood and Douglas R. Miller , title =. IEEE Transactions on Software Engineering , year =

  34. [42]

    Advances in Neural Information Processing Systems 36 , year =

    Jiawei Liu and Chunqiu Steven Xia and Yuyao Wang and Lingming Zhang , title =. Advances in Neural Information Processing Systems 36 , year =

  35. [43]

    Wu and Ilia Sucholutsky and Tania Lombrozo and Thomas L

    Ryan Liu and Jiayi Geng and Addison J. Wu and Ilia Sucholutsky and Tania Lombrozo and Thomas L. Griffiths , title =. Proceedings of the 42nd International Conference on Machine Learning , year =

  36. [44]

    IEEE Transactions on Software Engineering , year =

    Fang Liu and Yang Liu and Lin Shi and Zhen Yang and Li Zhang and Xiaoli Lian and Zhongqi Li and Yuchi Ma , title =. IEEE Transactions on Software Engineering , year =

  37. [45]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

    Yao Lu and Max Bartolo and Alastair Moore and Sebastian Riedel and Pontus Stenetorp , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  38. [46]

    2024 , eprint =

    Zhenyan Lu and Xiang Li and Dongqi Cai and others , title =. 2024 , eprint =

  39. [47]

    2025 , eprint =

    Wenjie Ma and Jingxuan He and Charlie Snell and Tyler Griggs and Sewon Min and Matei Zaharia , title =. 2025 , eprint =

  40. [48]

    Howie Huang and Enric Boix-Adser

    Rimon Melamed and Lucas Hurley McCabe and Tanay Wakhare and Yejin Kim and H. Howie Huang and Enric Boix-Adser. Prompts Have Evil Twins , booktitle =. 2024 , pages =

  41. [49]

    Transactions of the Association for Computational Linguistics , year =

    Moran Mizrahi and Guy Kaplan and Dan Malkin and Rotem Dror and Dafna Shahaf and Gabriel Stanovsky , title =. Transactions of the Association for Computational Linguistics , year =

  42. [50]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

    Niklas Muennighoff and Zitong Yang and Weijia Shi and others , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

  43. [51]

    Goodman and Judith E

    Linas Nasvytis and Simon Jerome Han and Ben Prystawski and Satchel Grant and Noah D. Goodman and Judith E. Fan , title =. 2026 , eprint =

  44. [52]

    Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing , year =

    Chien Van Nguyen and Xuan Shen and Ryan Aponte and others , title =. Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing , year =

  45. [53]

    2026 , eprint =

    Yansong Ning and Mianpeng Liu and Jingwen Ye and Weidong Zhang and Hao Liu , title =. 2026 , eprint =

  46. [54]

    A Systematic Methodology for Evaluating Failure Independence in

    Rodrigo Pato Nogueira and Karthik Pattabiraman and Marco Vieira and Jo. A Systematic Methodology for Evaluating Failure Independence in. 2026 , eprint =

  47. [55]

    Olausson and Jeevana Priya Inala and Chenglong Wang and Jianfeng Gao and Armando Solar-Lezama , title =

    Theo X. Olausson and Jeevana Priya Inala and Chenglong Wang and Jianfeng Gao and Armando Solar-Lezama , title =. The Twelfth International Conference on Learning Representations , year =

  48. [56]

    2025 , eprint =

    Julian Aron Prenner and Romain Robbes , title =. 2025 , eprint =

  49. [57]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

    Reid Pryzant and Dan Iter and Jerry Li and Yin Tat Lee and Chenguang Zhu and Michael Zeng , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

  50. [58]

    Can discrete information extraction prompts generalize across language models? , booktitle =

    Nathana. Can discrete information extraction prompts generalize across language models? , booktitle =

  51. [59]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

    Kiran Ramnath and Kang Zhou and Sheng Guan and others , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

  52. [60]

    2nd International Conference on Foundation and Large Language Models , year =

    Matthew Renze and Erhan Guven , title =. 2nd International Conference on Foundation and Large Language Models , year =

  53. [61]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

    Martin Riddell and Ansong Ni and Arman Cohan , title =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , year =

  54. [62]

    2026 , eprint =

    Rishav Rishav and Pushpak Pujari and Pushpendre Rastogi , title =. 2026 , eprint =

  55. [63]

    The Twelfth International Conference on Learning Representations , year =

    Manley Roberts and Himanshu Thakur and Christine Herlihy and Colin White and Samuel Dooley , title =. The Twelfth International Conference on Learning Representations , year =

  56. [64]

    2026 , eprint =

    Javier Ron and Benoit Baudry and Martin Monperrus , title =. 2026 , eprint =

  57. [65]

    The Twelfth International Conference on Learning Representations , year =

    Melanie Sclar and Yejin Choi and Yulia Tsvetkov and Alane Suhr , title =. The Twelfth International Conference on Learning Representations , year =

  58. [66]

    ICLR 2026 Workshop on Logical Reasoning of Large Language Models , year =

    Deep Shah and Sanket Badhe and Nehal Kathrotia and Priyanka Tiwari , title =. ICLR 2026 Workshop on Logical Reasoning of Large Language Models , year =

  59. [67]

    Advances in Neural Information Processing Systems 36 , year =

    Noah Shinn and Federico Cassano and Ashwin Gopinath and Karthik Narasimhan and Shunyu Yao , title =. Advances in Neural Information Processing Systems 36 , year =

  60. [68]

    2022 , eprint =

    Charlie Snell and Dan Klein and Ruiqi Zhong , title =. 2022 , eprint =

  61. [69]

    The Thirteenth International Conference on Learning Representations , year =

    Charlie Victor Snell and Jaehoon Lee and Kelvin Xu and Aviral Kumar , title =. The Thirteenth International Conference on Learning Representations , year =

  62. [70]

    Barr , title =

    Claudio Spiess and Prem Devanbu and Earl T. Barr , title =. Proceedings of the 3rd ACM International Conference on AI-Powered Software , year =

  63. [71]

    Proceedings of the 34th USENIX Security Symposium , year =

    Joseph Spracklen and Raveen Wijewickrama and A H M Nazmus Sakib and Anindya Maiti and Bimal Viswanath and Murtuza Jadliwala , title =. Proceedings of the 34th USENIX Security Symposium , year =

  64. [72]

    International Conference on Learning Representations , year =

    Zayne Rea Sprague and Fangcong Yin and Juan Diego Rodriguez and others , title =. International Conference on Learning Representations , year =

  65. [73]

    Transactions on Machine Learning Research , year =

    Yang Sui and Yu-Neng Chuang and Guanchu Wang and others , title =. Transactions on Machine Learning Research , year =

  66. [74]

    Proceedings of the AAAI Conference on Artificial Intelligence , year =

    Yuchen Tian and Weixiang Yan and Qian Yang and Xuandong Zhao and Qian Chen and Wen Wang and Ziyang Luo and Lei Ma and Dawn Song , title =. Proceedings of the AAAI Conference on Artificial Intelligence , year =

  67. [75]

    Zhang and Mark Harman and Helen Yannakoudakis , title =

    Lukas Twist and Jie M. Zhang and Mark Harman and Helen Yannakoudakis , title =. 2025 , eprint =

  68. [76]

    2025 , eprint =

    Shouren Wang and Wang Yang and Xianxuan Long and Qifan Wang and Vipin Chaudhary and Xiaotian Han , title =. 2025 , eprint =

  69. [77]

    2025 , eprint =

    Yaxuan Wang and Quan Liu and Zhenting Wang and others , title =. 2025 , eprint =

  70. [78]

    Advances in Neural Information Processing Systems 35 , year =

    Jason Wei and Xuezhi Wang and Dale Schuurmans and others , title =. Advances in Neural Information Processing Systems 35 , year =

  71. [79]

    2024 , eprint =

    Xiaohan Xu and Ming Li and Chongyang Tao and others , title =. 2024 , eprint =

  72. [80]

    2025 , eprint =

    Silei Xu and Wenhao Xie and Lingxiao Zhao and Pengcheng He , title =. 2025 , eprint =

  73. [81]

    Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

    Zeyuan Yang and Peng Li and Yang Liu , title =. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , year =

  74. [82]

    Gonzalez and Bin Cui , title =

    Ling Yang and Zhaochen Yu and Tianjun Zhang and Shiyi Cao and Minkai Xu and Wentao Zhang and Joseph E. Gonzalez and Bin Cui , title =. Advances in Neural Information Processing Systems 37 , year =

  75. [83]

    Le and Denny Zhou and Xinyun Chen , title =

    Chengrun Yang and Xuezhi Wang and Yifeng Lu and Hanxiao Liu and Quoc V. Le and Denny Zhou and Xinyun Chen , title =. International Conference on Learning Representations , year =

  76. [84]

    2024 , eprint =

    Ping Yu and Jing Xu and Jason Weston and Ilia Kulikov , title =. 2024 , eprint =

  77. [85]

    2025 , eprint =

    Ye Yu and Yaoning Yu and Haohan Wang , title =. 2025 , eprint =

  78. [86]

    Goodman , title =

    Eric Zelikman and Yuhuai Wu and Jesse Mu and Noah D. Goodman , title =. Advances in Neural Information Processing Systems 35 , year =

  79. [87]

    Proceedings of the 41st International Conference on Machine Learning , year =

    Tianjun Zhang and Aman Madaan and Luyu Gao and others , title =. Proceedings of the 41st International Conference on Machine Learning , year =

  80. [88]

    Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

    Jiajie Zhang and Nianyi Lin and Lei Hou and Ling Feng and Juanzi Li , title =. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , year =

  81. [89]

    Proceedings of the AAAI Conference on Artificial Intelligence , year =

    Andrew Zhao and Daniel Huang and Quentin Xu and Matthieu Lin and Yong-Jin Liu and Gao Huang , title =. Proceedings of the AAAI Conference on Artificial Intelligence , year =

  82. [90]

    International Conference on Learning Representations , year =

    Yongchao Zhou and Andrei Ioan Muresanu and Ziwen Han and others , title =. International Conference on Learning Representations , year =

  83. [91]

    Efficient Knowledge Injection in

    Kalle Kujanp. Efficient Knowledge Injection in. Transactions on Machine Learning Research , year =

  84. [92]

    Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year =

    Lei Li and Yongfeng Zhang and Li Chen , title =. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year =

  85. [93]

    Zico Kolter and Matt Fredrikson , title =

    Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson , title =. 2023 , eprint =

  86. [94]

    2026 , howpublished =

  87. [95]

    2026 , howpublished =

    Introducing. 2026 , howpublished =

  88. [96]

    2025 , eprint =

    Naizhu Jin and Zhong Li and Guang Yang and Tian Zhang and Qingkai Zeng , title =. 2025 , eprint =

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.