Pith. sign in

REVIEW 5 major objections 5 minor 32 references

Prompting as Scientific Inquiry

T0 review · 5 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Prompting is behavioral science, not a workaround, if LLMs are trained organisms.

desk verdict A coherent, honest position paper that frames prompt science vs. prompt engineering; the central thesis is already in the cited literature, but the detailed comparison with mechanistic interpretability is a useful contribution. read the letter →

arxiv 2507.00163 v2 pith:UAFPP255 submitted 2025-06-30 cs.CL

classification cs.CL
keywords promptsciencelargelanguagemodelsbehavioralmechanisticinterpretabilityfalsifiabilityscientificinquirychain-of-thoughtpromptingengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that prompting is not a stopgap for interacting with large language models but a genuine scientific method. Its central claim is that if LLMs are treated as complex, opaque organisms that are trained rather than programmed, then prompting is behavioral science: it probes the model through its native interface, language. The authors distinguish prompt science—using structured inputs and outputs to discover and confirm regularities—from prompt engineering, which optimizes prompts for a specific model, task, and dataset without making falsifiable claims about why they work. They argue that prompting has produced the field's major capability discoveries, from in-context learning to chain-of-thought, and that it is complementary to mechanistic interpretability, not inferior. A reader should care because this reframes how the community evaluates which methods count as real insight into LLMs.

What carries the argument

The central object is the distinction between prompt science and prompt engineering. Prompt science uses natural-language inputs and observed outputs as an intervention-and-measurement instrument: exploratory prompting discovers novel behavior, and prompting studies confirm hypotheses through systematic variation. The argument is carried by three mechanisms: language as a native interface whose structure mirrors what the model learned; a three-level analysis (computational, algorithmic, implementational) that places prompting at the computational level; and productive vagueness—the idea that linguistic prompts can specify hypotheses at varying precision while remaining falsifiable. Together these make prompting a scalable, testable, and accessible probe of LLM behavior.

What would settle it

Train or fine-tune a model so a known function is encoded in its internal representations, detectable by activation probes, and then search systematically for any prompt that makes the model output that function. If no prompt can elicit it, prompting alone fails to reveal a real capability and the completeness assumption behind the paper's argument would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that prompting, understood as behavioral experimentation, is the most impactful scientific method we have for finding out what lies inside large language models. On this view, language is the model's native interface, and probing with language intervenes on the same input distribution the model was trained on. Prompt science is falsifiable: claims such as 'adding a reasoning request changes accuracy on a benchmark' can be tested by varying prompts and checking whether the effect generalizes. Mechanistic interpretability and prompting operate at different levels of analysis—implementation versus computational behavior—and complement each other at the algorithmic level. The paper also contends that rebranding successful prompting as training or inference-time compute obscures a shared methodology and should stop.

Load-bearing premise

The load-bearing premise is that relevant computations inside an LLM can always be expressed in, and observed through, the probability distributions over language that some prompt elicits. If a capability leaves no prompt-elicitable behavioral trace, prompt science would miss it.

Editorial extensions

If this is right

  • Prompting studies become a legitimate source of scientific evidence about LLM capabilities, alongside or ahead of circuit-level analysis.
  • Prompt brittleness is reframed as a signal of what factors influence model behavior, not just a nuisance to engineer away.
  • Successful prompting methods no longer need to be rebranded as training or inference-time compute to be taken seriously.
  • Prompting for discovery and mechanistic interpretability for confirmation can be combined into a single research loop.
  • Systematic, falsifiable prompt experiments—varying prompts and testing across models—become the expected standard for behavioral claims.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the argument holds, behavioral probing should be the primary evaluation mode for new LLM releases, with capability reports coming from prompt experiments before weight-level analysis.
  • The argument rests on an elicitation completeness assumption that could be tested by comparing the set of behaviors reachable by prompts with the set detectable by activation probes on the same model.
  • The same logic extends beyond text: whatever interface is optimized during training—vision, audio, tool use—becomes the natural probe for behavioral science of that system.
  • A practical consequence is methodological: prompt studies should adopt preregistration, systematic variation, and negative-result reporting borrowed from behavioral science.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. This position paper argues that prompting—interacting with LLMs through natural language and observing their outputs—should be considered a legitimate scientific method, analogous to behavioral science, rather than dismissed as engineering or alchemy. The authors distinguish prompt science from prompt engineering, present three case studies (Sparks of AGI, chain-of-thought, Constitutional AI) as evidence that prompting has been the primary discovery mechanism for LLM capabilities, compare prompting with mechanistic interpretability along Marr's levels of analysis, and rebut six common critiques of prompting. The central claim is that prompting is "the most impactful method we have for finding out what lies inside LLMs."

Significance. If the central claim is accepted, the paper would legitimately reframe prompting as a core scientific tool for studying and controlling LLMs, complementary to mechanistic interpretability, with practical implications for how the field allocates effort. The paper is clearly written, makes a useful distinction between prompt science and prompt engineering, and provides a structured rebuttal to common objections, which makes it a valuable position piece. However, the argument relies on several asserted empirical claims and does not yet provide an operational criterion for separating genuine model capabilities from elicitation artifacts, which is essential to the central claim.

major comments (5)
  1. [Section 4 and Section 6 ("Prompts are brittle")] The paper asserts that prompt sensitivity is informative, but it does not provide an operational criterion for distinguishing a stable model capability from an elicitation artifact. Since the paper's central claim is that prompting is "the most impactful method we have for finding out what lies inside LLMs," this omission is load-bearing. The authors should specify a verification protocol, e.g., requiring convergent evidence across paraphrastic prompts, across models, or through mechanistic follow-up, before a behavior is attributed to the model rather than the prompt.
  2. [Section 5.1 ("Different abstraction languages")] The claim that "the structure of language mirrors aspects of the model itself" is asserted without argument or evidence. This is load-bearing for the paper's view that language is the model's native interface and that prompting gives faithful access. The authors should provide a derivation from the training objective or weaken the claim to a testable hypothesis.
  3. [Section 1] The claim that "interpretability has, so far, largely confirmed hypotheses we've already formulated" is a universal empirical assertion with no supporting citation or survey. Because it is used to argue that prompting is the primary discovery method, the paper should either substantiate it with a systematic review or soften it to a more defensible claim.
  4. [Section 5.1 ("Different abstraction languages")] The inference that difficulty in prompting a task is "a powerful signal that the LLM does not have an accurate representation of the necessary concepts" conflates the expressive limits of the prompting interface with the model's internal representations. The paper should address the alternative explanation that the human, not the model, lacks the right linguistic handle.
  5. [Section 4 (reason 2)] The claim that language models were "discovered not designed" is cited to Holtzman et al. (2023), but the argument is not summarized. This phrase carries a load-bearing role in the paper's argument that prompting is the most direct discovery method. The authors should either lay out the reasoning or rely on independent evidence.
minor comments (5)
  1. [Figure 1 caption] The word "overivew" should be "overview."
  2. [Section 2, Case Study: Chain-of-Thought] The phrase "an method to engineer" should be "a method to engineer."
  3. [Section 2, Case Study: Chain-of-Thought] "GPT-4o1" appears to be a typo; the reference list cites "Introducing OpenAI o1," so the in-text name should be "OpenAI o1."
  4. [Section 5.1] "where as" should be "whereas."
  5. [References] The Marr (1982) citation is incomplete as given ("The philosophy and the approach. Visual perception: Essential readings"); the standard book citation is Marr, D. (1982). Vision. W. H. Freeman.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the argument is a position essay anchored in external case studies and falsifiability, not a derivation that reduces to its inputs.

full rationale

The central claim—that prompting is a scientifically legitimate and productive way to study LLMs—is supported by historical case studies (in-context learning, chain-of-thought, constitutional AI) and by independent external work (Bubeck et al., McCoy et al., Sclar et al., Makelov et al., Paulo & Belrose). The paper contains no fitted parameter later renamed as a prediction, no equation-level equivalence between premise and conclusion, and no uniqueness theorem imported from the authors' prior work. The only self-citations are to Holtzman et al. (2023) for the proposition that 'language models were discovered not designed' and for a critique of anisotropy worries; these are supporting background claims rather than the sole load-bearing justification. Even if that premise were set aside, the paper's case for prompt science rests on the falsifiability of prompt-based claims and on structured variation across contexts and models. Thus any concern that the 'most impactful method' assertion overreaches is a substantive correctness or skepticism question, not a circularity in the paper's reasoning.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's argument rests on three domain assumptions: LLMs are organism-like, language is a sufficient probe, and prompting results generalize across models. No free parameters are fitted, and the paper introduces no new physical or mathematical entities. The conceptual distinction between prompt science and prompt engineering is a framing device, not an invented entity.

assumptions (3)
  • domain assumption LLMs are best understood as a new kind of organism that is trained rather than programmed.
    The paper's central analogy to behavioral science rests on treating LLMs as opaque organisms; this is asserted in the abstract and Section 1 without proof.
  • domain assumption Language is the 'native interface' of LLMs, so probing in language gives direct and sufficient access to their capabilities.
    The argument that prompting is not inferior assumes that input-output language interfaces reveal the model's internal functioning; stated in Section 4 and Section 5.1.
  • domain assumption Behavioral evidence from prompting generalizes across model families at scale.
    The rebuttal to 'prompts will not generalize across models' in Section 6 relies on examples (ICL, RAG, CoT) being robust across families, presented without systematic evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Prompting as Scientific Inquiry." pith.science (2026). https://pith.science/paper/UAFPP255

@misc{pith2026250700163,
  author       = {Pith},
  title        = {Pith review of: Prompting as Scientific Inquiry},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UAFPP255}},
  note         = {Machine review of arXiv:2507.00163}
}
read the original abstract

Prompting is the primary method by which we study and control large language models. It is also one of the most powerful: nearly every major capability attributed to LLMs-few-shot learning, chain-of-thought, constitutional AI-was first unlocked through prompting. Yet prompting is rarely treated as science and is frequently frowned upon as alchemy. We argue that this is a category error. If we treat LLMs as a new kind of complex and opaque organism that is trained rather than programmed, then prompting is not a workaround: it is behavioral science. Mechanistic interpretability peers into the neural substrate, prompting probes the model in its native interface: language. We contend that prompting is not inferior, but rather a key component in the science of LLMs.

Figures

Figures reproduced from arXiv: 2507.00163 by the authors.

Figure 1
Figure 1. An overview of key arguments in this work. A. It is important to distinguish [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. An illustration of the timeline related to the success of prompting (highlighted in green) [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. An overview of the comparison between prompt science and interpretability. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

32 extracted references · 14 canonical work pages

  1. [3]

    a is b" fail to learn

    Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reversal curse: Llms trained on" a is b" fail to learn" b is a". arXiv preprint arXiv:2309.12288,

  2. [5]

    Trenton Bricken, Aditya Templeton, Jonathan Batson, Brittany Chen, Adam Jermyn, et al

    doi: 10.1145/3531146.3533083. Trenton Bricken, Aditya Templeton, Jonathan Batson, Brittany Chen, Adam Jermyn, et al. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread (Anthropic),

  3. [6]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

    Available online at transformer-circuits.pub/2023/monosemantic- features. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901,

  4. [11]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948,

  5. [12]

    We can’t understand ai using our existing vocabulary

    John Hewitt, Robert Geirhos, and Been Kim. We can’t understand ai using our existing vocabulary. arXiv preprint arXiv:2502.07586,

  6. [13]

    Generative Models as a Complex Systems Science: How can we make sense of large language model behavior?

    Ari Holtzman, Peter West, and Luke Zettlemoyer. Generative models as a complex systems science: How can we make sense of large language model behavior? arXiv preprint arXiv:2308.00189,

  7. [15]

    Prompt-Hacking: The New p-Hacking?

    Thomas Kosch and Sebastian Feger. Prompt-hacking: The new p-hacking? arXiv preprint arXiv:2504.14571,

  8. [16]

    Llms get lost in multi-turn conversation

    Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. Llms get lost in multi-turn conversation. arXiv preprint arXiv:2505.06120,

Show all 32 references
  1. [17]

    Mitigating the alignment tax of rlhf

    Yong Lin, Hangyu Lin, Wei Xiong, Shizhe Diao, Jianmeng Liu, Jipeng Zhang, Rui Pan, Haoxiang Wang, Wenbin Hu, Hanning Zhang, et al. Mitigating the alignment tax of rlhf. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 580–606,

  2. [19]

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig

    doi: 10.1145/3317287.3328534. Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys, 55(9):1–35,

  3. [21]

    Prompting science report 1: Prompt engineering is complicated and contingent

    Lennart Meincke, Ethan Mollick, Lilach Mollick, and Dan Shapiro. Prompting science report 1: Prompt engineering is complicated and contingent. arXiv preprint arXiv:2503.04818,

  4. [22]

    Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp

    Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. ...

  5. [23]

    URL https://arxiv.org/abs/2301.05217. Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Os- azuwa Ness, Hoifung Poon, Tao Qin, Naoto Usuyama, Chr...

  6. [24]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al

    Accessed: 2025-5-21. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information proces...

  7. [25]

    Sparse autoencoders trained on the same data learn different features

    Gonçalo Paulo and Nora Belrose. Sparse autoencoders trained on the same data learn different features. arXiv preprint arXiv:2501.16615,

  8. [26]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman

    Accessed: 2025-05-20. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling,

  9. [27]

    org/abs/2207.13243

    URL https://arxiv. org/abs/2207.13243. Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference...

  10. [28]

    Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young

    D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young. Machine learning: The high interest credit card of technical debt. In SE4ML: Software Engineering for Machine Learning (NIPS 2014 Workshop),

  11. [29]

    doi: 10.1145/3709599

    ISSN 0001-0782. doi: 10.1145/3709599. URL https://doi.org/10. 1145/3709599. Lewis Smith, Sen Rajamanoharan, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, and Neel Nanda. Negative results for sparse autoen- coders on downstream tasks and deprioritising...

  12. [30]

    14 Chenhao Tan

    URL https://deepmindsafetyresearch.medium.com/negative-results-for-sparse- autoencoders-on-downstream-tasks-and-deprioritising-sae-research- 6cadcfc125b9. 14 Chenhao Tan. On the diversity and limits of human explanations. arXiv preprint arXiv:2106.11988,

  13. [31]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530,

  14. [32]

    Universal adversarial triggers for attacking and analyzing nlp

    Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Universal adversarial triggers for attacking and analyzing nlp. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natur...

  15. [2006]

    Towards principled evaluations of sparse autoencoders for interpretability and control

    Aleksandar Makelov, Georg Lange, and Neel Nanda. Towards principled evaluations of sparse autoencoders for interpretability and control. InSeT-LLM Workshop at the International Conference on Learning Representations (ICLR) 2024,

  16. [2017]

    Language models use trigonometry to do addition

    Subhash Kantamneni and Max Tegmark. Language models use trigonometry to do addition. arXiv preprint arXiv:2502.00873,

  17. [2018]

    doi: 10.1145/3236386. 3241340. Zachary Chase Lipton and Jacob Steinhardt. Troubling trends in machine learning scholarship. Queue, 17(1):45–77,

  18. [2019]

    Language models as agent models

    Jacob Andreas. Language models as agent models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 5769–5779,

  19. [2020]

    Sparks of artificial general intelligence: Early experiments with gpt-4

    Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712,

  20. [2021]

    Scaling and evaluating sparse autoencoders

    Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093,

  21. [2022]

    Constitutional ai: Harmlessness from ai feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073,

  22. [2023]

    The values encoded in machine learning research

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. The values encoded in machine learning research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pp. 173–184,

  23. [2024]

    Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al

    arXiv:2309.08600 [cs.LG]. Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1(1):12,

  24. [2025]

    Thomas L Griffiths, Jian-Qiao Zhu, Erin Grant, and R Thomas McCoy

    URL https://arxiv.org/abs/2505.00038. Thomas L Griffiths, Jian-Qiao Zhu, Erin Grant, and R Thomas McCoy. Bayes in the age of intelligent machines. Current Directions in Psychological Science, 33(5):283–291,

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.