Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Prompt-Hacking: The New p-Hacking?

T0 review · 3 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper argues that LLMs' inherent biases, non-determinism, and opacity make them unsuitable for data analysis tasks that demand rigor, impartiality, and reproducibility, so prompt-hacking should be treated like p-hacking.

desk verdict A clear, well-intentioned opinion piece that repackages existing LLM-integrity critiques under a catchy label, but overstates its case by claiming LLMs are categorically unsuitable while also recommending practices that presuppose they can be made reliable. read the letter →

arxiv 2504.14571 v2 pith:N7UPQJJL submitted 2025-04-20 cs.HC

classification cs.HC
keywords largelanguagemodelsprompt-hackingp-hackingPARKingreproducibilityresearchintegritydataanalysispreregistration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This opinion paper argues that using large language models for scientific data analysis threatens research integrity in the same way p-hacking does, but with a crucial difference: the bias and variability are in the tool itself. The authors ask researchers to treat prompt-hacking, the reworking of prompts until an LLM produces the desired result, as a serious integrity risk and to generally avoid LLM-based analysis unless its use is necessary and justified. The paper introduces PARKing (Prompt Adjustments to Reach Known Outcomes) to name the practice of steering prompts toward pre-existing hypotheses. A sympathetic reader would take away a clear recommendation: document prompts, preregister them, test prompt stability, and default to traditional methods.

What carries the argument

The paper's central device is the analogy between prompt-hacking and p-hacking, operationalized through the named concept PARKing (Prompt Adjustments to Reach Known Outcomes), defined as systematically reworking prompts until the LLM's output matches a pre-existing hypothesis. This analogy transfers the known integrity harms of undisclosed analytical flexibility from statistical testing to LLM prompting, and the paper's additional claim that the tool itself is biased by design makes the risk stronger than in classic p-hacking.

What would settle it

Run one fixed statistical analysis task with a preregistered prompt, temperature set to 0, a fixed random seed, and independent validation against a standard statistical package across 100 repeated executions; if all outputs are identical and correct, the claim that LLMs cannot be relied on for reproducible analysis would be contradicted in that regime.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLMs are not neutral analytical instruments: their outputs inherit training-data biases, vary with prompt phrasing, and resist reproduction, so even a perfectly documented and honestly used LLM analysis cannot guarantee validity. The paper draws a direct parallel between prompt-hacking and p-hacking, and it sharpens the analogy with PARKing (Prompt Adjustments to Reach Known Outcomes), the practice of systematically modifying prompts until they yield results that align with a pre-existing hypothesis. Unlike p-hacking, which misuses inherently neutral statistical techniques, prompt-hacking exploits tools that are not impartial by design. Therefore, the authors conclude, the central question for most data analysis tasks is not how to use LLMs responsibly but whether to use them at all.

Load-bearing premise

The categorical conclusion rests on the premise that LLM outputs remain too variable, biased, and opaque for any documentation, stability testing, or validation to make them trustworthy for analysis; if seeded, low-temperature, repeatedly validated LLM runs prove stable and accurate, the blanket recommendation to avoid LLMs loses its empirical grounding.

Editorial extensions

If this is right

  • Researchers should generally avoid LLMs for data analysis unless they can establish that the LLM is necessary and its benefits outweigh its risks.
  • Prompt stability checks, repeated runs, and transparent documentation of the full prompt creation process should become routine before any LLM output is relied on.
  • Prompts and their revision history should be preregistered, with publication venues, funders, and infrastructure providers expecting such documentation as part of reproducible research practice.
  • Even fully documented and honestly used LLM analysis cannot be assumed valid, because non-determinism and training-data bias are inherent to the tool rather than correctable by good practice.
  • Treating prompt-hacking like p-hacking implies that undisclosed prompt flexibility should be seen as a research-integrity violation, not merely a methodological weakness.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit testable extension: prompt stability, measured as output variance across repeated identical runs, could become a standard reproducibility metric analogous to inter-rater reliability for human coders.
  • The analogy predicts that fields with low preregistration and high analyst flexibility will be the first to see integrity failures from LLM-driven analysis; auditing submitted prompt logs would reveal how often prompts were tuned to match hypotheses.
  • A softer policy conclusion than the paper's blanket avoidance: LLMs could be confined to clearly labeled exploratory uses, with confirmatory statistics always run in traditional software, preserving the integrity argument while allowing some practical benefit.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This opinion paper argues that large language models (LLMs) are unsuitable for most data analysis tasks in empirical research because of their inherent biases, non-determinism, and opacity. It draws an analogy between 'prompt-hacking'—iteratively adjusting prompts to elicit desirable outputs—and p-hacking, and introduces the term PARKing (Prompt Adjustments to Reach Known Outcomes) as a parallel to HARKing. The paper recommends that researchers avoid LLM-based analysis except in limited, justified, documented cases, and proposes practical measures such as prompt preregistration, stability checks, and transparent documentation.

Significance. The paper's value is primarily argumentative: it names a risk and offers actionable, if generic, safeguards for LLM use in research. Its strongest contributions are the explicit parallel to p-hacking and the concrete checklist in Section 5, which could inform future guidelines and preregistration templates. However, the central categorical claim—that even correct LLM use cannot guarantee validity—is not supported by direct evidence, and the paper does not quantify the variability or bias it invokes. The significance is therefore conditional: it is a useful prompt for community debate and a set of testable claims, not a settled scientific finding.

major comments (3)
  1. [Abstract and Section 6] The abstract and Section 6 assert that 'even the correct use of LLMs in analysis cannot guarantee validity,' but the paper offers no measurement of LLM output variability, bias, or opacity on a concrete analysis task, and no comparison with human analysis or traditional statistical pipelines. The cited support (e.g., Refs. [3], [4], [5]) consists of prior commentary rather than data. This universal negative claim exceeds what an opinion piece can establish; I recommend reformulating it as a risk claim or adding a small, task-specific demonstration (e.g., repeated prompting of a fixed analysis prompt) to turn the claim into a falsifiable observation.
  2. [Section 5 vs. Section 6] There is a tension between the categorical unsuitability claim and the recommended protocol in Section 5. The recommendations to 'repeat prompts and assess the stability of generated results over time,' to document prompt history, and to preregister prompts presuppose that stability checking and documentation can provide reproducibility and validity. If, as Section 6 states, even correct use cannot guarantee validity, then the Section 5 protocol cannot deliver its stated purpose; if it can, the categorical claim is too strong. The authors should either soften the conclusion to 'correct use reduces but does not eliminate risk' or explicitly argue why documentation and stability checks are insufficient.
  3. [Section 4] The definition of prompt-hacking mixes deliberate and inadvertent prompt adjustment, whereas p-hacking is typically understood as intentional manipulation of analysis to achieve significance. The analogy is weakened if inadvertent variation is included, because the ethical and causal structure differs. The paper should separate accidental variability from intentional 'hacking' and state which of these the recommendations target, or justify why the same caution applies to both.
minor comments (4)
  1. [Section 5] The sentence 'Reproducible, stable, and repeatable outputs are important data analysis components for ensuring reproducibility' is tautological; consider rewording to describe the specific properties needed (e.g., 'stable outputs across replications are a condition for reproducibility').
  2. [Reference [10]] The reference to Yilmaz (2013) contains a malformed URL prefix 'arXiv:https://onlinelibrary.wiley.com/...' that should be cleaned up.
  3. [Section 4] The claim that 'the prompting space is infinite' is rhetorically strong but imprecise; more accurate is 'the space of possible prompts is effectively unbounded.'
  4. [Section 4] PARKing is introduced with 'may arise' but is not used elsewhere in the paper; consider integrating it into the p-hacking discussion or removing it for clarity.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the paper's claims are argued from prior cited commentary, not from fitted inputs or self-referential equations.

full rationale

This is an opinion paper rather than a formal derivation or empirical study, so the standard circularity patterns do not directly apply. There are no equations, fitted parameters, or quantitative predictions whose values reduce to inputs by construction. The central claim that LLMs are unsuitable for rigorous data analysis rests on cited prior commentary, including Gibney [3], Morris [5], and the authors' own prior piece [4]. The self-citation [4] is used to support the assertions that LLMs inherit biases and limitations and that prompt-hacking was introduced recently, but these claims are also supported by external references and by the paper's own qualitative argument, so the self-citation is not load-bearing in a way that makes the conclusion equivalent to its input. The internal tension between the categorical unsuitability claim in the abstract and Section 6 and the Section 5 recommendation that repeated prompts and documentation can improve stability is a substantive argumentative weakness, but it is not circularity: the paper does not define its conclusion into existence. The lack of direct measurement of LLM variability and the absence of a comparison to human analysis pipelines are evidentiary gaps, which are correctness risks rather than circularity. Consistent with the rule that absence of standard consensus or assertion-based support is not itself grounds for a circularity finding, no specific circular step can be exhibited from the text.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

This opinion paper adds no fitted parameters or formal derivations. It depends on several domain assumptions about LLM behavior and research practice, and it coins the construct PARKing without independent validation. The ledger reflects that the paper's contribution is conceptual and advisory, not empirical.

assumptions (4)
  • domain assumption LLMs exhibit inherent biases, hallucinations, and non-deterministic outputs.
    Stated in Section 4 and relied on throughout; supported only by citations to prior commentary, not by data in this paper.
  • domain assumption Non-determinism and bias in an analysis tool render it unsuitable for tasks requiring rigor and impartiality.
    Central normative premise in Sections 4 and 6; asserted without arguing that these properties cannot be adequately mitigated by documentation and validation.
  • domain assumption The p-hacking analogy directly transfers to prompt-hacking.
    The parallel between p-hacking and prompt-hacking is asserted in Sections 1 and 4 but not empirically established, for example by showing that prompt modifications systematically shift outcomes toward significance.
  • domain assumption LLMs cannot understand or evaluate data context as a human researcher would.
    Section 4 states this without evidence, and it underlies the claim that LLMs generate 'plausible but factually incorrect outputs' in analysis.
invented entities (1)
  • PARKing (Prompt Adjustments to Reach Known Outcomes)
    purpose: To name the behavior of systematically modifying prompts until results match pre-existing hypotheses, analogous to HARKing.
    Introduced in Section 4. No empirical evidence is presented that this practice occurs or is widespread; it is offered as a conceptual label based on analogy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Prompt-Hacking: The New p-Hacking?." pith.science (2026). https://pith.science/paper/N7UPQJJL

@misc{pith2026250414571,
  author       = {Pith},
  title        = {Pith review of: Prompt-Hacking: The New p-Hacking?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/N7UPQJJL}},
  note         = {Machine review of arXiv:2504.14571}
}
read the original abstract

As Large Language Models (LLMs) become increasingly embedded in empirical research workflows, their use as analytical tools for quantitative or qualitative data raises pressing concerns for scientific integrity. This opinion paper draws a parallel between "prompt-hacking", the strategic tweaking of prompts to elicit desirable outputs from LLMs, and the well-documented practice of "p-hacking" in statistical analysis. We argue that the inherent biases, non-determinism, and opacity of LLMs make them unsuitable for data analysis tasks demanding rigor, impartiality, and reproducibility. We emphasize how researchers may inadvertently, or even deliberately, adjust prompts to confirm hypotheses while undermining research validity. We advocate for a critical view of using LLMs in research, transparent prompt documentation, and clear standards for when LLM use is appropriate. We discuss how LLMs can replace traditional analytical methods, whereas we recommend that LLMs should only be used with caution, oversight, and justification.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prompting as Scientific Inquiry

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Position paper arguing that prompting LLMs is a form of behavioral science and should be recognized as a core scientific method alongside mechanistic interpretability.

Reference graph

Works this paper leans on

10 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [3]

    Elizabeth Gibney. 2022. Is AI fuelling a reproducibility crisis in science. Nature 608, 7922 (2022), 250–1. doi:10.1038/ d41586-022-02035-w

  2. [4]

    Thomas Kosch and Sebastian Feger. 2024. Risk or Chance? Large Language Models and Reproducibility in HCI Research. Interactions 31, 6 (Oct. 2024), 44–49. doi:10.1145/3695765

  3. [5]

    Meredith Ringel Morris. 2024. Prompting Considered Harmful. Commun. ACM 67, 12 (Nov. 2024), 28–30. doi:10.1145/ 3673861

  4. [1]

    Arriaga, and Adam Tauman Kalai

    Gati V Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara , Vol. 1, No. 1, Articl...

  5. [2]

    Andy Cockburn, Carl Gutwin, and Alan Dix. 2018. HARK No More: On the Preregistration of CHI Experiments. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–12. doi:10.1145/3173574.3173715

  6. [6]

    Albrecht Schmidt, Passant Elagroudy, Fiona Draxler, Frauke Kreuter, and Robin Welsch. 2024. Simulating the Human in HCD with ChatGPT: Redesigning Interaction Design with AI. Interactions 31, 1 (Jan. 2024), 24–31. doi:10.1145/3637436

  7. [7]

    Simmons, Leif D

    Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn. 2011. False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant. Psychological Science 22, 11 (2011), 1359–1366. doi:10.1177/0956797611417632

  8. [8]

    Stefan and Felix D

    Angelika M. Stefan and Felix D. Schönbrodt. 2023. Big little lies: a compendium and simulation of p-hacking strategies. Royal Society Open Science 10, 2 (2023), 220346. doi:10.1098/rsos.220346

Show all 10 references
  1. [9]

    Wilbert Tabone and Joost de Winter. 2023. Using ChatGPT for human–computer interaction research: a primer. Royal Society Open Science 10, 9 (2023), 231053. doi:10.1098/rsos.231053

  2. [10]

    Kaya Yilmaz. 2013. Comparison of Quantitative and Qualitative Research Traditions: epistemological, theoreti- cal, and methodological differences. European Journal of Education 48, 2 (2013), 311–325. doi:10.1111/ejed.12014 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.