{"id":"4ac436e1-c14d-4a87-83a5-ef0748c2ebce","arxiv_id":"2507.18638","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 243 users finds most agree that clearer prompts improve AI results, yet the study never links real prompting behavior to measured productivity gains.","lead":"This paper surveys 243 AI users about how they write prompts and whether they feel the results help their work. Most respondents say clear, structured prompts improve AI output, but the study measures opinions, not actual productivity.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never tests the claimed link: no statistical comparison connects Table 8 prompting behavior to productivity outcomes, so the central claim rests on circular self-report.","rationale":"The reader's weakest assumption identifies the same gap: self-reported beliefs are treated as direct evidence of a causal link between prompt behavior and productivity. My independent review of the manuscript finds no analysis that would rescue the claim. The paper does present useful descriptive statistics about who uses AI and which prompting techniques they report, and those descriptions appear internally consistent. However, the central causal claim about prompt engineering improving productivity is never actually tested. Because the missing step is a simple crosstab that the survey could have supported, this is an internal evidentiary gap rather than an external controversy. This supports the reader's REJECT verdict, so I recommend no change.","tokens_in":10711,"tokens_out":2663,"duration_ms":28829,"concrete_test":"Request the anonymized raw survey responses from the authors. Construct a binary predictor 'structured prompting' from Table 8 (respondent selected role, chain-of-thought, or instruction prompting) and an outcome from Table 13 (respondent 'agrees' or 'strongly agrees' that AI speeds task completion). Compute the 2×2 contingency table and run a chi-square test with Cramér's V, plus a logistic regression controlling for AI usage frequency and education level. If the association is not statistically significant, or if it disappears after controls, the paper's central claim is unsupported. If the raw data are not available, rerun the same test on an independent sample of comparable size and recruitment strategy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that users who employ clear, structured, context-aware prompts report higher task efficiency and better outcomes. For that claim to hold, the data must show an association between actual prompting behavior and productivity-related outcomes. The manuscript never provides such a test. Table 8 records which prompting techniques respondents say they use, and Table 9 records revision frequency, but Tables 11–14 only record satisfaction and self-reported beliefs. No crosstabulation, correlation, regression, or significance test links Table 8 or Table 9 with any productivity variable. The most direct evidence offered, Table 12, is simply agreement with the statement 'clearer prompts lead to better results'; using that as evidence is circular because respondents are endorsing the hypothesis. Table 14 reports means only, with no disaggregation by prompting behavior. Thus the abstract's causal-sounding conclusion is not derived from the analysis. This is a claim-without-test, not merely a disagreement with external consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a self-report survey of 243 AI users, with sections covering respondent demographics, AI usage patterns, prompting techniques, prompt revision frequency, satisfaction, and perceived productivity benefits. The authors' central claim is that users who employ clear, structured, and context-aware prompts experience higher task efficiency and better outcomes, and they conclude that prompt engineering is a critical determinant of AI-assisted productivity. The analysis, however, is almost entirely descriptive: it reports frequencies and means but never statistically connects reported prompting behavior to any productivity outcome.","tokens_in":10899,"tokens_out":4941,"duration_ms":45914,"significance":"If the central claim were properly supported, the paper would offer practical guidance for AI literacy programs and workplace prompt-training initiatives. The manuscript provides a clear taxonomy of manual versus automatic prompting techniques (Table 1) and a diverse, internationally distributed convenience sample, which is useful as a descriptive snapshot. However, the core relationship between prompting behavior and productivity is never tested. The paper's main 'finding' is essentially a restatement of respondents' agreement with a survey item, so the study's current contribution is limited to documenting perceptions rather than establishing the claimed effect. No reproducible code or machine-checked derivations are included.","major_comments":[{"comment":"The central claim that users who employ clear, structured, and context-aware prompts report higher task efficiency and better outcomes is not tested anywhere in the paper. Table 8 records which prompting techniques respondents say they use and Table 9 records revision frequency, while Tables 11–14 report satisfaction and perceived benefits. No cross-tabulation, correlation, regression, or significance test connects Table 8 or Table 9 with any productivity variable. Section 3.6 explicitly promises 'Correlation and Trend Analysis' using scipy.stats, but the Results chapter contains no such analysis. As a result, the abstract's conclusion is not derivable from the data presented.","section":"§4.3–§4.4 and Abstract"},{"comment":"The key evidence for 'clear prompts lead to better outcomes' is respondents' agreement with the statement 'more specific and clearer prompts lead to better AI results.' This is circular: the conclusion is effectively the survey item itself. Self-reported belief is neither a measure of prompting behavior nor an objective measure of task outcomes. To support the central claim, the authors would need to compare outcome ratings across groups that differ in reported prompting behavior, such as users of particular techniques versus non-users, or high-frequency versus low-frequency revisers.","section":"§4.4, Table 12"},{"comment":"The crosstabulation of education level and prompting techniques in Table 10 counts mentions rather than unique respondents, so its rows are not coherent with the sample sizes reported in Table 3. For example, the Bachelor's degree row sums to 129, yet only 83 respondents hold a Bachelor's degree. Consequently, the statement in §4.5 that 'educational background influences prompting strategy diversity' is unsupported by the table as presented.","section":"§4.3, Table 10"},{"comment":"The paper repeatedly describes its design as 'descriptive quantitative' and 'exploratory,' yet the Abstract and Section 4.5 use causal and evaluative language such as 'lead to clearer outcomes' and 'critical factor.' A descriptive survey of beliefs and self-reported satisfaction cannot support causal claims about the effect of prompt structure on productivity; either the conclusions must be tempered or the missing inferential analysis must be provided.","section":"§3.3, §4.5"}],"minor_comments":[{"comment":"The phrase 'rapid engineering using scalable language models' appears to be a typo; it should presumably read 'prompt engineering.'","section":"Abstract"},{"comment":"The Data Collection Procedure states that data collection spanned six weeks (05 January 2025 to 10 February 2025), but later says 'At the end of the three-week period.' This is internally inconsistent.","section":"§3.5"},{"comment":"The text says 'more than 66% of respondents (160 out of 243)' use AI at least twice a week, but 160/243 ≈ 65.8%, so 'approximately 66%' would be accurate.","section":"§4.5"},{"comment":"Table 14 reports means and standard deviations only; the interpretive sentence that these results 'affirm the hypothesis' is overconfident without any inferential tests or confidence intervals.","section":"§4.4, Table 14"},{"comment":"The survey relies entirely on self-report, but the manuscript does not discuss common-method bias, social desirability, or self-selection of respondents, all of which are salient for the validity of the stated conclusions.","section":"§4.4"}],"recommendation":"reject","confidential_remarks":"The manuscript could be suitable as a descriptive workshop paper, but as a full journal article the gap between the analysis and the central claim is too large. The authors would need to supply the missing cross-tabulations and inferential tests from their raw data, and even then the conclusions would remain limited by the self-report nature of the measures. Given that the current write-up presents a claim without a corresponding test, I recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a survey of 243 LLM users' prompting habits and attitudes. The descriptive data are fine: the tables on usage frequency, task types, and prompting technique adoption are clear, and the related-work section competently summarizes manual and automatic prompting methods. There is real value in the observation that role prompting and chain-of-thought are the most commonly used techniques among this convenience sample, and that respondents revise prompts often.\n\nThe problem is the headline claim. The abstract says users who employ clear, structured, context-aware prompts report higher task efficiency and better outcomes. The data never show that. There is no cross-tabulation between prompting technique use (Table 8) and any productivity measure, no correlation, no regression. The paper says in §3.6 that it used scipy for correlation analysis, but the Results section contains only frequencies, percentages, and means. The closest thing to evidence, Table 12, asks respondents if they agree that clearer prompts lead to better results, and 83.7% agree. That's circular: the \"finding\" is the survey item itself. Table 13 is the same for speed. So the central conclusion is not derived from the analysis; it is a restatement of the respondents' beliefs.\n\nThis is a load-bearing flaw, not a cosmetic one. If the paper were framed as an exploratory study of user perceptions of prompt engineering, the conclusion would be defensible. But as written, it claims to show a relationship between behavior and outcomes, and it doesn't. The convenience sample and lack of shared data are secondary issues; the primary issue is that the analysis never tests the relationship the paper is about.\n\nThe citation pattern is mostly fine; it draws on the right papers for few-shot, CoT, and prompt tuning, though a few references are Medium posts or industry whitepapers, which is common in this space.\n\nI would desk reject this version. The descriptive content could support a short research note on prompting practices, but the current manuscript's framing overreaches, and the missing analysis is not a minor revision. Not worth a full referee cycle until the author either provides a behavioral test or explicitly limits the claims to perception.","headline":"A survey that describes prompting habits but never tests its central claim, which rests on circular self-report.","tokens_in":11376,"tokens_out":3919,"would_cite":false,"duration_ms":40531,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Survey of 243 users ties clear prompts to higher perceived productivity.","keywords":["prompt engineering","large language models","human productivity","human-AI interaction","survey research","self-reported productivity","prompt clarity","user satisfaction"],"falsifier":"Run the same set of writing or coding tasks with users randomly assigned to vague prompts versus structured, context-rich prompts, and compare objective completion time and independently scored output quality; if structured prompts show no measurable advantage, the paper's central causal claim is refuted.","tokens_in":10546,"feed_emoji":"📊","tokens_out":5580,"duration_ms":50049,"temperature":0.7,"pith_summary":"This paper argues that prompt engineering — how clearly and structurally a user phrases requests to a large language model — is a major determinant of how much productivity users get out of tools like ChatGPT, Gemini, and DeepSeek. Drawing on 243 survey responses, it reports that a large majority of users believe specific, context-rich prompts improve AI output quality, and that most users revise prompts to get better results. The paper's core claim is that structured prompting, even when learned by trial and error, translates into higher task efficiency and better outcomes. If correct, this makes prompt literacy a practical workplace and classroom skill worth teaching.","feed_headline":"Clear prompts lift perceived AI productivity, survey of 243 finds","feed_subtitle":"Respondents who use structured, context-aware prompts report faster task completion and better LLM outputs.","key_machinery":"The central instrument is a structured online questionnaire with four sections — demographics, AI usage, prompting awareness, and perceived productivity — analyzed with descriptive statistics and simple correlational checks. The claim is carried by three self-report measures: belief that prompt clarity improves results (Table 12, mean 4.01/5), belief that AI speeds up work (Table 13, mean 3.87/5), and satisfaction with AI output (Table 11, mean 3.24/5). The paper uses these Likert-scale responses as the bridge between prompting behavior and productivity.","core_discovery":"The paper claims to show that users who employ clear, structured, and context-aware prompts report higher task efficiency and better outcomes with LLMs. In the survey, 83.7% of respondents agreed or strongly agreed that clearer and more specific prompts lead to better AI results, and 75.7% agreed that AI helps them complete tasks faster. The study also reports that role prompting, chain-of-thought prompting, and instruction prompting are the most-used techniques, and that over half of respondents revise their prompts often or occasionally. The author interprets these patterns as evidence that human input quality, not just model capability, is decisive in realizing productivity gains from generative AI.","pith_inferences":["Extending beyond the paper, a controlled experiment comparing vague and structured prompts on identical writing or coding tasks could show whether the perceived gains match objective gains in speed and output quality.","The finding that education level tracks prompting breadth suggests a testable extension: short prompt-training sessions may narrow the productivity gap between less- and more-educated users.","If perceived productivity is what drives continued AI adoption, then even subjective gains could be economically meaningful, because they shape user engagement and acceptance."],"forward_implications":["If the claim holds, teaching prompt design as a basic digital skill should raise the value people get from LLMs in education and work.","Tool designers could reasonably add prompt-guiding interfaces, templates, and revision feedback to improve outcomes for users who do not prompt well spontaneously.","The finding implies employers and educators should invest in prompt literacy programs rather than leaving users to learn by trial and error.","Users' habit of iterative revision suggests that human-in-the-loop refinement is part of how LLM productivity is actually realized."],"supporting_citations":[{"why":"supplies the definition of zero-shot prompting used to frame the survey's technique categories.","marker":"[1]"},{"why":"supplies the definition of few-shot prompting used to frame the survey's technique categories.","marker":"[2]"},{"why":"provides the chain-of-thought reasoning technique that anchors one of the most-used prompting strategies.","marker":"[3]"},{"why":"gives guidelines for instruction prompting that shape the paper's account of clear prompt design.","marker":"[4]"},{"why":"supplies the role prompting technique that was the most frequently reported in the survey.","marker":"[5]"},{"why":"offers prior evidence that generative AI plus prompt engineering improves learning outcomes.","marker":"[10]"},{"why":"reports up to 30% faster turnaround times for workers using prompting strategies, supporting the productivity claim.","marker":"[11]"},{"why":"provides prior evidence that structured prompting enhances human decision-making quality.","marker":"[13]"}],"fun_headline_variants":["Survey: Structured prompts boost perceived AI output quality","83% report clearer prompts yield better LLM results","Prompt clarity linked to higher task efficiency in AI use","243 users: precise prompts make AI more productive","Better prompts, better AI: survey of 243 users"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that respondents' agreement that clear prompts improve results and that AI speeds up their work is a reliable measure of actual prompting behavior and real productivity gains, never linking reported technique use to objective task outcomes.","fun_headline_variants_meta":{"raw":{"variants":["Survey: Structured prompts boost perceived AI output quality","83% report clearer prompts yield better LLM results","Prompt clarity linked to higher task efficiency in AI use","243 users: precise prompts make AI more productive","Better prompts, better AI: survey of 243 users"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1257,"prompt_tokens":775,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":391,"completion_tokens_details":{"reasoning_tokens":407}},"tokens_in":391,"tokens_out":482,"duration_ms":5478,"temperature":1.0,"reasoning_tokens":407,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:34:07.916078+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same set of writing or coding tasks with users randomly assigned to vague prompts versus structured, context-rich prompts, and compare objective completion time and independently scored output quality; if structured prompts show no measurable advantage, the paper's central causal claim is refuted.","supporting_citations":[{"cited_title":"Fairness-guided Few-shot Prompting for Large Language Models,","cited_arxiv_id":null,"evidence_quote":"supplies the definition of few-shot prompting used to frame the survey's technique categories."},{"cited_title":"Guidelines for Prompting Large Language Models,","cited_arxiv_id":null,"evidence_quote":"gives guidelines for instruction prompting that shape the paper's account of clear prompt design."},{"cited_title":"Assigning Roles to Chatbots,","cited_arxiv_id":null,"evidence_quote":"supplies the role prompting technique that was the most frequently reported in the survey."},{"cited_title":"Enhancing English Comprehension through Generative AI and Prompt Engineering: A Study on Undergraduate Learning Outcomes,","cited_arxiv_id":null,"evidence_quote":"offers prior evidence that generative AI plus prompt engineering improves learning outcomes."},{"cited_title":"Mastering generative AI: Why effective prompting is the key to success at work,","cited_arxiv_id":null,"evidence_quote":"reports up to 30% faster turnaround times for workers using prompting strategies, supporting the productivity claim."},{"cited_title":"The Cognitive Effects of AI-Human Collaboration: A Behavioral Study on Prompt Design and Decision-Making,","cited_arxiv_id":null,"evidence_quote":"provides prior evidence that structured prompting enhances human decision-making quality."}],"review_version":1}