{"id":"10e88e81-9840-4460-9932-22e57c8d0f16","arxiv_id":"2504.14571","paper_version":2,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Prompt-hacking, the strategic adjustment of prompts to get desired LLM outputs, is likened to p-hacking, and the authors argue LLMs are unreliable analysts that should be avoided for rigorous science.","lead":"This opinion paper argues that repeatedly tweaking prompts to make LLMs produce desired answers, a practice it calls prompt-hacking, damages research integrity much like p-hacking. It recommends that researchers avoid LLMs for most scientific data analysis and document and preregister prompts when LLMs are used.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is asserted, not measured; the paper's own recommendations in §5 presuppose that LLM outputs can be stabilized and validated, which undercuts the categorical unsuitability claim in the abstract and §6.","rationale":"The reader's weakest assumption is that LLM outputs are so variable, biased, and opaque as to be unsuitable for analysis regardless of documentation and validation. This is exactly the load-bearing empirical premise: the strongest claim depends on it, and it is supported only by citations to prior commentary (Gibney 2022; Morris 2024; Kosch & Feger 2024) and by informal extrapolation from the existence of prompt sensitivity. My stress-test confirms that premise is not measured in this paper: there is no new data, no controlled comparison, and no formal argument that prompt documentation/stability checks cannot restore reproducibility. The internal tension between §5 (which prescribes repeat-and-validate procedures that would make outputs reproducible) and §6 (which says correct use cannot guarantee validity) makes the categorical conclusion depend on a substantive empirical claim about what validation can or cannot achieve. The correct scholarly response is therefore to treat the paper as an opinion piece whose central claim is unverified, which matches the reader's UNVERDICTED verdict. I found no independent flaw in the logic of the p-hacking analogy itself; the analogy is reasonable as a rhetorical device, but the strength of the conclusion exceeds the evidence. The concrete test I propose would settle the empirical premise: it measures variability and bias on fixed tasks with controls, including temperature 0, and compares against human/traditional baselines and against the paper's own §5 stability protocol. I recommend UNCHANGED because the reader's verdict already marks the paper UNVERDICTED with high confidence; my concern reinforces that verdict rather than altering it. A REJECT verdict would be too strong because the paper is explicitly an opinion/position paper and does not claim to present new measurements; a CONDITIONAL ACCEPT would be inappropriate because no result is being accepted. The paper's recommendations (documentation, preregistration, caution) are reasonable and independently supported by the cited literature, so the paper retains value as a call for standards even though its strongest claim is unproven.","tokens_in":4619,"tokens_out":2097,"duration_ms":17594,"concrete_test":"Design a small evaluation battery with three standard analysis tasks (e.g., a t-test interpretation with a known effect size, a thematic coding of 20 short interview excerpts with a pre-registered codebook, and a descriptive statistics report on a fixed dataset). For each task, run a fixed prompt N=100 times at the provider's default temperature and also at temperature 0, across 3 model versions. Measure (a) semantic answer variability (e.g., percent of runs whose conclusion changes), (b) accuracy against ground truth, (c) effect of a deliberately varied prompt synonym set on conclusions. Then run the same tasks with 20 human analysts or 20 independent traditional pipelines.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (abstract, §4, §6) is categorical: LLMs are 'fundamentally unreliable' and 'unsuitable' for data analysis tasks demanding rigor, with 'even the correct use' unable to guarantee validity. This is the load-bearing premise behind the recommendation that LLMs be avoided for most analysis tasks. The stated support is a set of citations to prior opinion/commentary (Gibney 2022; Morris 2024; Kosch & Feger 2024) describing risks such as variability, bias, hallucination, and opacity. The paper reports no measurement of these properties on any concrete analysis task, and it also provides no baseline: it never characterizes how often traditional human analysis or conventional statistical pipelines produce unstable, biased, or opaque results. Without a task-specific measurement of LLM output variability/bias and a comparison against the status quo, the argument cannot establish the categorical conclusion; it establishes at most that LLM use has risks that require caution. Additionally, the paper's own §5 recommendations (repeat prompts, assess stability over time, preregister prompts, document variations) presuppose that prompt stability and documented validation are achievable and would provide reproducibility. That is in tension with the abstract's claim that 'even the correct use of LLMs in analysis cannot guarantee validity.' If correct use cannot guarantee validity, then the recommended 'correct use' protocol in §5 is insufficient by the paper's own standard; if the §5 protocol does provide reproducibility through prompt stability checks and preregistration, then the categorical claim in §6 is overstated. The paper does not resolve this internal tension, and the reader's UNVERDICTED verdict appropriately reflects that the central empirical premise is neither verified nor falsified here.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This opinion paper argues that large language models (LLMs) are unsuitable for most data analysis tasks in empirical research because of their inherent biases, non-determinism, and opacity. It draws an analogy between 'prompt-hacking'—iteratively adjusting prompts to elicit desirable outputs—and p-hacking, and introduces the term PARKing (Prompt Adjustments to Reach Known Outcomes) as a parallel to HARKing. The paper recommends that researchers avoid LLM-based analysis except in limited, justified, documented cases, and proposes practical measures such as prompt preregistration, stability checks, and transparent documentation.","tokens_in":4858,"tokens_out":4824,"duration_ms":38836,"significance":"The paper's value is primarily argumentative: it names a risk and offers actionable, if generic, safeguards for LLM use in research. Its strongest contributions are the explicit parallel to p-hacking and the concrete checklist in Section 5, which could inform future guidelines and preregistration templates. However, the central categorical claim—that even correct LLM use cannot guarantee validity—is not supported by direct evidence, and the paper does not quantify the variability or bias it invokes. The significance is therefore conditional: it is a useful prompt for community debate and a set of testable claims, not a settled scientific finding.","major_comments":[{"comment":"The abstract and Section 6 assert that 'even the correct use of LLMs in analysis cannot guarantee validity,' but the paper offers no measurement of LLM output variability, bias, or opacity on a concrete analysis task, and no comparison with human analysis or traditional statistical pipelines. The cited support (e.g., Refs. [3], [4], [5]) consists of prior commentary rather than data. This universal negative claim exceeds what an opinion piece can establish; I recommend reformulating it as a risk claim or adding a small, task-specific demonstration (e.g., repeated prompting of a fixed analysis prompt) to turn the claim into a falsifiable observation.","section":"Abstract and Section 6"},{"comment":"There is a tension between the categorical unsuitability claim and the recommended protocol in Section 5. The recommendations to 'repeat prompts and assess the stability of generated results over time,' to document prompt history, and to preregister prompts presuppose that stability checking and documentation can provide reproducibility and validity. If, as Section 6 states, even correct use cannot guarantee validity, then the Section 5 protocol cannot deliver its stated purpose; if it can, the categorical claim is too strong. The authors should either soften the conclusion to 'correct use reduces but does not eliminate risk' or explicitly argue why documentation and stability checks are insufficient.","section":"Section 5 vs. Section 6"},{"comment":"The definition of prompt-hacking mixes deliberate and inadvertent prompt adjustment, whereas p-hacking is typically understood as intentional manipulation of analysis to achieve significance. The analogy is weakened if inadvertent variation is included, because the ethical and causal structure differs. The paper should separate accidental variability from intentional 'hacking' and state which of these the recommendations target, or justify why the same caution applies to both.","section":"Section 4"}],"minor_comments":[{"comment":"The sentence 'Reproducible, stable, and repeatable outputs are important data analysis components for ensuring reproducibility' is tautological; consider rewording to describe the specific properties needed (e.g., 'stable outputs across replications are a condition for reproducibility').","section":"Section 5"},{"comment":"The reference to Yilmaz (2013) contains a malformed URL prefix 'arXiv:https://onlinelibrary.wiley.com/...' that should be cleaned up.","section":"Reference [10]"},{"comment":"The claim that 'the prompting space is infinite' is rhetorically strong but imprecise; more accurate is 'the space of possible prompts is effectively unbounded.'","section":"Section 4"},{"comment":"PARKing is introduced with 'may arise' but is not used elsewhere in the paper; consider integrating it into the p-hacking discussion or removing it for clarity.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"This is an opinion paper, and the manuscript's self-citation to the authors' earlier Interactions piece (Ref. [4]) is prominent. The editorial process should weigh the genre when considering the strength of the empirical claims. The paper could be a reasonable 'views' contribution if the categorical wording is softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The punchline first: this is an opinion paper, not an empirical study, so it should be read as argumentative commentary. Its central claim—that LLMs are categorically unsuitable for rigorous data analysis, and that even correct use cannot guarantee validity—is much stronger than the evidence it brings to bear. But the paper is also clearly written, honest about its opinion status, and its practical recommendations are sensible.\n\nWhat is actually new is the label PARKing (Prompt Adjustments to Reach Known Outcomes), which neatly extends HARKing into the LLM context. That is a genuinely useful addition to the vocabulary. The paper also does a decent job of synthesizing earlier critiques from Morris (2024) and the authors' own 2024 Interactions piece into a compact five-page argument. For an audience that has not been following this literature closely, it is a readable and thought-provoking summary.\n\nThe soft spots are in proportion to the paper's ambitions. No empirical evidence is offered for the load-bearing claim that LLM outputs are so variable, biased, and opaque that no amount of documentation or validation can make them reliable for analysis. The argument leans on citations to other opinion pieces and to the authors' own prior work. Worse, the paper's own Section 5 recommendations—repeat prompts, assess stability over time, preregister prompts, document variations—presuppose that prompt stability and validation are achievable and would meaningfully reduce risk. That sits in direct tension with the abstract and Section 6 claim that even correct use cannot guarantee validity. If the recommended protocol can provide reproducibility, the categorical claim is overstated; if it cannot, the paper has not explained why. The paper does not resolve this, and it should.\n\nMinor point: the p-hacking analogy is evocative but imperfect. P-hacking is a documented practice with known prevalence and systematic consequence. Prompt-hacking is, so far, a hypothesized risk. Calling it \"the new p-hacking\" oversells the analogy before the empirical case is made.\n\nThe self-citation is fine—it is a natural extension of their own prior argument, not a hidden agenda.\n\nWho does this paper serve? HCI and adjacent empirical researchers weighing whether and how to use LLMs in analysis, and editors or funders thinking about reproducibility standards. It consolidates existing concerns and gives them a memorable name, but it does not break new empirical ground.\n\nMy recommendation: send it to peer review at a venue that explicitly accepts opinion or position papers. A serious referee should push the authors to soften the categorical language, acknowledge the internal tension, and clarify under what conditions (if any) they believe prompt stability checks and preregistration are sufficient. If the venue expects empirical results, desk reject; otherwise, it deserves a proper read.","headline":"A clear, well-intentioned opinion piece that repackages existing LLM-integrity critiques under a catchy label, but overstates its case by claiming LLMs are categorically unsuitable while also recommending practices that presuppose they can be made reliable.","tokens_in":5460,"tokens_out":2012,"would_cite":false,"duration_ms":20911,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that LLMs' inherent biases, non-determinism, and opacity make them unsuitable for data analysis tasks that demand rigor, impartiality, and reproducibility, so prompt-hacking should be treated like p-hacking.","keywords":["large language models","prompt-hacking","p-hacking","PARKing","reproducibility","research integrity","data analysis","preregistration"],"falsifier":"Run one fixed statistical analysis task with a preregistered prompt, temperature set to 0, a fixed random seed, and independent validation against a standard statistical package across 100 repeated executions; if all outputs are identical and correct, the claim that LLMs cannot be relied on for reproducible analysis would be contradicted in that regime.","tokens_in":4387,"feed_emoji":"⚠️","tokens_out":4538,"duration_ms":40062,"temperature":0.7,"pith_summary":"This opinion paper argues that using large language models for scientific data analysis threatens research integrity in the same way p-hacking does, but with a crucial difference: the bias and variability are in the tool itself. The authors ask researchers to treat prompt-hacking, the reworking of prompts until an LLM produces the desired result, as a serious integrity risk and to generally avoid LLM-based analysis unless its use is necessary and justified. The paper introduces PARKing (Prompt Adjustments to Reach Known Outcomes) to name the practice of steering prompts toward pre-existing hypotheses. A sympathetic reader would take away a clear recommendation: document prompts, preregister them, test prompt stability, and default to traditional methods.","feed_headline":"LLM analysis is the new p-hacking, paper argues","feed_subtitle":"Even honest, documented prompt use cannot guarantee valid results; researchers should default to traditional methods.","key_machinery":"The paper's central device is the analogy between prompt-hacking and p-hacking, operationalized through the named concept PARKing (Prompt Adjustments to Reach Known Outcomes), defined as systematically reworking prompts until the LLM's output matches a pre-existing hypothesis. This analogy transfers the known integrity harms of undisclosed analytical flexibility from statistical testing to LLM prompting, and the paper's additional claim that the tool itself is biased by design makes the risk stronger than in classic p-hacking.","core_discovery":"The paper's central claim is that LLMs are not neutral analytical instruments: their outputs inherit training-data biases, vary with prompt phrasing, and resist reproduction, so even a perfectly documented and honestly used LLM analysis cannot guarantee validity. The paper draws a direct parallel between prompt-hacking and p-hacking, and it sharpens the analogy with PARKing (Prompt Adjustments to Reach Known Outcomes), the practice of systematically modifying prompts until they yield results that align with a pre-existing hypothesis. Unlike p-hacking, which misuses inherently neutral statistical techniques, prompt-hacking exploits tools that are not impartial by design. Therefore, the authors conclude, the central question for most data analysis tasks is not how to use LLMs responsibly but whether to use them at all.","pith_inferences":["An implicit testable extension: prompt stability, measured as output variance across repeated identical runs, could become a standard reproducibility metric analogous to inter-rater reliability for human coders.","The analogy predicts that fields with low preregistration and high analyst flexibility will be the first to see integrity failures from LLM-driven analysis; auditing submitted prompt logs would reveal how often prompts were tuned to match hypotheses.","A softer policy conclusion than the paper's blanket avoidance: LLMs could be confined to clearly labeled exploratory uses, with confirmatory statistics always run in traditional software, preserving the integrity argument while allowing some practical benefit."],"forward_implications":["Researchers should generally avoid LLMs for data analysis unless they can establish that the LLM is necessary and its benefits outweigh its risks.","Prompt stability checks, repeated runs, and transparent documentation of the full prompt creation process should become routine before any LLM output is relied on.","Prompts and their revision history should be preregistered, with publication venues, funders, and infrastructure providers expecting such documentation as part of reproducible research practice.","Even fully documented and honestly used LLM analysis cannot be assumed valid, because non-determinism and training-data bias are inherent to the tool rather than correctable by good practice.","Treating prompt-hacking like p-hacking implies that undisclosed prompt flexibility should be seen as a research-integrity violation, not merely a methodological weakness."],"supporting_citations":[{"why":"introduces HARKing, the preregistration-motivated contrast that PARKing parallels.","marker":"[2]"},{"why":"the report that AI and LLM use is fueling a reproducibility crisis, grounding the paper's core concern.","marker":"[3]"},{"why":"prior work by the same authors framing LLMs and reproducibility in HCI as a risk.","marker":"[4]"},{"why":"the critique that prompting is a poor interface and that unrecorded prompt histories break replicability.","marker":"[5]"},{"why":"the classic demonstration that undisclosed analytical flexibility lets researchers present anything as significant.","marker":"[7]"},{"why":"the compendium of p-hacking strategies that makes the p-hacking comparison concrete.","marker":"[8]"},{"why":"a primer proposing ChatGPT for HCI research, the target practice the paper argues against.","marker":"[9]"}],"fun_headline_variants":["Prompt-hacking is the new p-hacking","LLM analysis: the new p-hacking","AI prompt-hacking: the new p-hacking","Prompt-hacking parallels p-hacking in research","LLMs are not neutral; prompt-hacking is p-hacking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The categorical conclusion rests on the premise that LLM outputs remain too variable, biased, and opaque for any documentation, stability testing, or validation to make them trustworthy for analysis; if seeded, low-temperature, repeatedly validated LLM runs prove stable and accurate, the blanket recommendation to avoid LLMs loses its empirical grounding.","fun_headline_variants_meta":{"raw":{"variants":["Prompt-hacking is the new p-hacking","LLM analysis: the new p-hacking","AI prompt-hacking: the new p-hacking","Prompt-hacking parallels p-hacking in research","LLMs are not neutral; prompt-hacking is p-hacking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000941,"raw_usage":{"total_tokens":3971,"prompt_tokens":842,"completion_tokens":3129,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":3053}},"tokens_in":458,"tokens_out":3129,"duration_ms":20347,"temperature":1.0,"reasoning_tokens":3053,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:45:16.526897+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one fixed statistical analysis task with a preregistered prompt, temperature set to 0, a fixed random seed, and independent validation against a standard statistical package across 100 repeated executions; if all outputs are identical and correct, the claim that LLMs cannot be relied on for reproducible analysis would be contradicted in that regime.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the report that AI and LLM use is fueling a reproducibility crisis, grounding the paper's core concern."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"prior work by the same authors framing LLMs and reproducibility in HCI as a risk."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"the critique that prompting is a poor interface and that unrecorded prompt histories break replicability."}],"review_version":1}