REVIEW 5 major objections 5 minor 32 references
Prompting as Scientific Inquiry
T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Prompting is behavioral science, not a workaround, if LLMs are trained organisms.
desk verdict A coherent, honest position paper that frames prompt science vs. prompt engineering; the central thesis is already in the cited literature, but the detailed comparison with mechanistic interpretability is a useful contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the distinction between prompt science and prompt engineering. Prompt science uses natural-language inputs and observed outputs as an intervention-and-measurement instrument: exploratory prompting discovers novel behavior, and prompting studies confirm hypotheses through systematic variation. The argument is carried by three mechanisms: language as a native interface whose structure mirrors what the model learned; a three-level analysis (computational, algorithmic, implementational) that places prompting at the computational level; and productive vagueness—the idea that linguistic prompts can specify hypotheses at varying precision while remaining falsifiable. Together these make prompting a scalable, testable, and accessible probe of LLM behavior.
What would settle it
Train or fine-tune a model so a known function is encoded in its internal representations, detectable by activation probes, and then search systematically for any prompt that makes the model output that function. If no prompt can elicit it, prompting alone fails to reveal a real capability and the completeness assumption behind the paper's argument would be falsified.
Extended reading notes
Core claim
The paper's central claim is that prompting, understood as behavioral experimentation, is the most impactful scientific method we have for finding out what lies inside large language models. On this view, language is the model's native interface, and probing with language intervenes on the same input distribution the model was trained on. Prompt science is falsifiable: claims such as 'adding a reasoning request changes accuracy on a benchmark' can be tested by varying prompts and checking whether the effect generalizes. Mechanistic interpretability and prompting operate at different levels of analysis—implementation versus computational behavior—and complement each other at the algorithmic level. The paper also contends that rebranding successful prompting as training or inference-time compute obscures a shared methodology and should stop.
Load-bearing premise
The load-bearing premise is that relevant computations inside an LLM can always be expressed in, and observed through, the probability distributions over language that some prompt elicits. If a capability leaves no prompt-elicitable behavioral trace, prompt science would miss it.
Editorial extensions
If this is right
- Prompting studies become a legitimate source of scientific evidence about LLM capabilities, alongside or ahead of circuit-level analysis.
- Prompt brittleness is reframed as a signal of what factors influence model behavior, not just a nuisance to engineer away.
- Successful prompting methods no longer need to be rebranded as training or inference-time compute to be taken seriously.
- Prompting for discovery and mechanistic interpretability for confirmation can be combined into a single research loop.
- Systematic, falsifiable prompt experiments—varying prompts and testing across models—become the expected standard for behavioral claims.
Reading between the lines
- If the argument holds, behavioral probing should be the primary evaluation mode for new LLM releases, with capability reports coming from prompt experiments before weight-level analysis.
- The argument rests on an elicitation completeness assumption that could be tested by comparing the set of behaviors reachable by prompts with the set detectable by activation probes on the same model.
- The same logic extends beyond text: whatever interface is optimized during training—vision, audio, tool use—becomes the natural probe for behavioral science of that system.
- A practical consequence is methodological: prompt studies should adopt preregistration, systematic variation, and negative-result reporting borrowed from behavioral science.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that prompting—interacting with LLMs through natural language and observing their outputs—should be considered a legitimate scientific method, analogous to behavioral science, rather than dismissed as engineering or alchemy. The authors distinguish prompt science from prompt engineering, present three case studies (Sparks of AGI, chain-of-thought, Constitutional AI) as evidence that prompting has been the primary discovery mechanism for LLM capabilities, compare prompting with mechanistic interpretability along Marr's levels of analysis, and rebut six common critiques of prompting. The central claim is that prompting is "the most impactful method we have for finding out what lies inside LLMs."
Significance. If the central claim is accepted, the paper would legitimately reframe prompting as a core scientific tool for studying and controlling LLMs, complementary to mechanistic interpretability, with practical implications for how the field allocates effort. The paper is clearly written, makes a useful distinction between prompt science and prompt engineering, and provides a structured rebuttal to common objections, which makes it a valuable position piece. However, the argument relies on several asserted empirical claims and does not yet provide an operational criterion for separating genuine model capabilities from elicitation artifacts, which is essential to the central claim.
major comments (5)
- [Section 4 and Section 6 ("Prompts are brittle")] The paper asserts that prompt sensitivity is informative, but it does not provide an operational criterion for distinguishing a stable model capability from an elicitation artifact. Since the paper's central claim is that prompting is "the most impactful method we have for finding out what lies inside LLMs," this omission is load-bearing. The authors should specify a verification protocol, e.g., requiring convergent evidence across paraphrastic prompts, across models, or through mechanistic follow-up, before a behavior is attributed to the model rather than the prompt.
- [Section 5.1 ("Different abstraction languages")] The claim that "the structure of language mirrors aspects of the model itself" is asserted without argument or evidence. This is load-bearing for the paper's view that language is the model's native interface and that prompting gives faithful access. The authors should provide a derivation from the training objective or weaken the claim to a testable hypothesis.
- [Section 1] The claim that "interpretability has, so far, largely confirmed hypotheses we've already formulated" is a universal empirical assertion with no supporting citation or survey. Because it is used to argue that prompting is the primary discovery method, the paper should either substantiate it with a systematic review or soften it to a more defensible claim.
- [Section 5.1 ("Different abstraction languages")] The inference that difficulty in prompting a task is "a powerful signal that the LLM does not have an accurate representation of the necessary concepts" conflates the expressive limits of the prompting interface with the model's internal representations. The paper should address the alternative explanation that the human, not the model, lacks the right linguistic handle.
- [Section 4 (reason 2)] The claim that language models were "discovered not designed" is cited to Holtzman et al. (2023), but the argument is not summarized. This phrase carries a load-bearing role in the paper's argument that prompting is the most direct discovery method. The authors should either lay out the reasoning or rely on independent evidence.
minor comments (5)
- [Figure 1 caption] The word "overivew" should be "overview."
- [Section 2, Case Study: Chain-of-Thought] The phrase "an method to engineer" should be "a method to engineer."
- [Section 2, Case Study: Chain-of-Thought] "GPT-4o1" appears to be a typo; the reference list cites "Introducing OpenAI o1," so the in-text name should be "OpenAI o1."
- [Section 5.1] "where as" should be "whereas."
- [References] The Marr (1982) citation is incomplete as given ("The philosophy and the approach. Visual perception: Essential readings"); the standard book citation is Marr, D. (1982). Vision. W. H. Freeman.
Circularity Check
No circularity: the argument is a position essay anchored in external case studies and falsifiability, not a derivation that reduces to its inputs.
full rationale
The central claim—that prompting is a scientifically legitimate and productive way to study LLMs—is supported by historical case studies (in-context learning, chain-of-thought, constitutional AI) and by independent external work (Bubeck et al., McCoy et al., Sclar et al., Makelov et al., Paulo & Belrose). The paper contains no fitted parameter later renamed as a prediction, no equation-level equivalence between premise and conclusion, and no uniqueness theorem imported from the authors' prior work. The only self-citations are to Holtzman et al. (2023) for the proposition that 'language models were discovered not designed' and for a critique of anisotropy worries; these are supporting background claims rather than the sole load-bearing justification. Even if that premise were set aside, the paper's case for prompt science rests on the falsifiability of prompt-based claims and on structured variation across contexts and models. Thus any concern that the 'most impactful method' assertion overreaches is a substantive correctness or skepticism question, not a circularity in the paper's reasoning.
Assumptions & free parameters
assumptions (3)
- domain assumption LLMs are best understood as a new kind of organism that is trained rather than programmed.
- domain assumption Language is the 'native interface' of LLMs, so probing in language gives direct and sufficient access to their capabilities.
- domain assumption Behavioral evidence from prompting generalizes across model families at scale.
Cite this review
Pith. "Pith review of Prompting as Scientific Inquiry." pith.science (2026). https://pith.science/paper/UAFPP255
@misc{pith2026250700163,
author = {Pith},
title = {Pith review of: Prompting as Scientific Inquiry},
year = {2026},
howpublished = {\url{https://pith.science/paper/UAFPP255}},
note = {Machine review of arXiv:2507.00163}
}
read the original abstract
Prompting is the primary method by which we study and control large language models. It is also one of the most powerful: nearly every major capability attributed to LLMs-few-shot learning, chain-of-thought, constitutional AI-was first unlocked through prompting. Yet prompting is rarely treated as science and is frequently frowned upon as alchemy. We argue that this is a category error. If we treat LLMs as a new kind of complex and opaque organism that is trained rather than programmed, then prompting is not a workaround: it is behavioral science. Mechanistic interpretability peers into the neural substrate, prompting probes the model in its native interface: language. We contend that prompting is not inferior, but rather a key component in the science of LLMs.
Figures
Reference graph
Works this paper leans on
-
[3]
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reversal curse: Llms trained on" a is b" fail to learn" b is a". arXiv preprint arXiv:2309.12288,
-
[5]
Trenton Bricken, Aditya Templeton, Jonathan Batson, Brittany Chen, Adam Jermyn, et al
doi: 10.1145/3531146.3533083. Trenton Bricken, Aditya Templeton, Jonathan Batson, Brittany Chen, Adam Jermyn, et al. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread (Anthropic),
-
[6]
Available online at transformer-circuits.pub/2023/monosemantic- features. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901,
work page 2023
-
[11]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948,
-
[12]
We can’t understand ai using our existing vocabulary
John Hewitt, Robert Geirhos, and Been Kim. We can’t understand ai using our existing vocabulary. arXiv preprint arXiv:2502.07586,
-
[13]
Ari Holtzman, Peter West, and Luke Zettlemoyer. Generative models as a complex systems science: How can we make sense of large language model behavior? arXiv preprint arXiv:2308.00189,
-
[15]
Prompt-Hacking: The New p-Hacking?
Thomas Kosch and Sebastian Feger. Prompt-hacking: The new p-hacking? arXiv preprint arXiv:2504.14571,
-
[16]
Llms get lost in multi-turn conversation
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. Llms get lost in multi-turn conversation. arXiv preprint arXiv:2505.06120,
Show all 32 references
-
[17]
Mitigating the alignment tax of rlhf
Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, et al. Mitigating the alignment tax of rlhf. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 580–606,
2024
-
[19]
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig
doi: 10.1145/3317287.3328534. Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys, 55(9):1–35,
-
[21]
Prompting science report 1: Prompt engineering is complicated and contingent
Lennart Meincke, Ethan Mollick, Lilach Mollick, and Dan Shapiro. Prompting science report 1: Prompt engineering is complicated and contingent. arXiv preprint arXiv:2503.04818,
-
[22]
Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. ...
2022
-
[23]
URL https://arxiv.org/abs/2301.05217. Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Os- azuwa Ness, Hoifung Poon, Tao Qin, Naoto Usuyama, Chr...
-
[24]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al
Accessed: 2025-5-21. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information proces...
2025
-
[25]
Sparse autoencoders trained on the same data learn different features
Gonçalo Paulo and Nora Belrose. Sparse autoencoders trained on the same data learn different features. arXiv preprint arXiv:2501.16615,
-
[26]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman
Accessed: 2025-05-20. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling,
2025
-
[27]
org/abs/2207.13243
URL https://arxiv. org/abs/2207.13243. Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference...
-
[28]
Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young. Machine learning: The high interest credit card of technical debt. In SE4ML: Software Engineering for Machine Learning (NIPS 2014 Workshop),
2014
-
[29]
doi: 10.1145/3709599
ISSN 0001-0782. doi: 10.1145/3709599. URL https://doi.org/10. 1145/3709599. Lewis Smith, Sen Rajamanoharan, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, and Neel Nanda. Negative results for sparse autoen- coders on downstream tasks and deprioritising...
-
[30]
14 Chenhao Tan
URL https://deepmindsafetyresearch.medium.com/negative-results-for-sparse- autoencoders-on-downstream-tasks-and-deprioritising-sae-research- 6cadcfc125b9. 14 Chenhao Tan. On the diversity and limits of human explanations. arXiv preprint arXiv:2106.11988,
-
[31]
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530,
-
[32]
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Universal adversarial triggers for attacking and analyzing nlp. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natur...
2019
-
[2006]
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, Georg Lange, and Neel Nanda. Towards principled evaluations of sparse autoencoders for interpretability and control. InSeT-LLM Workshop at the International Conference on Learning Representations (ICLR) 2024,
2024
-
[2017]
Language models use trigonometry to do addition
Subhash Kantamneni and Max Tegmark. Language models use trigonometry to do addition. arXiv preprint arXiv:2502.00873,
-
[2018]
doi: 10.1145/3236386. 3241340. Zachary Chase Lipton and Jacob Steinhardt. Troubling trends in machine learning scholarship. Queue, 17(1):45–77,
-
[2019]
Language models as agent models
Jacob Andreas. Language models as agent models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 5769–5779,
2022
-
[2020]
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712,
-
[2021]
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093,
-
[2022]
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073,
-
[2023]
The values encoded in machine learning research
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 173–184,
2022
-
[2024]
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al
arXiv:2309.08600 [cs.LG]. Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12,
-
[2025]
Thomas L Griffiths, Jian-Qiao Zhu, Erin Grant, and R Thomas McCoy
URL https://arxiv.org/abs/2505.00038. Thomas L Griffiths, Jian-Qiao Zhu, Erin Grant, and R Thomas McCoy. Bayes in the age of intelligent machines. Current Directions in Psychological Science, 33(5):283–291,
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.