Pith. sign in

REVIEW 4 major objections 5 minor 13 references

Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Survey of 243 users ties clear prompts to higher perceived productivity.

desk verdict A survey that describes prompting habits but never tests its central claim, which rests on circular self-report. read the letter →

arxiv 2507.18638 v2 pith:BVSLZOLJ submitted 2025-05-10 cs.HC cs.AI

classification cs.HCcs.AI
keywords promptengineeringlargelanguagemodelshumanproductivityhuman-AIinteractionsurveyresearchself-reportedclarityusersatisfaction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that prompt engineering — how clearly and structurally a user phrases requests to a large language model — is a major determinant of how much productivity users get out of tools like ChatGPT, Gemini, and DeepSeek. Drawing on 243 survey responses, it reports that a large majority of users believe specific, context-rich prompts improve AI output quality, and that most users revise prompts to get better results. The paper's core claim is that structured prompting, even when learned by trial and error, translates into higher task efficiency and better outcomes. If correct, this makes prompt literacy a practical workplace and classroom skill worth teaching.

What carries the argument

The central instrument is a structured online questionnaire with four sections — demographics, AI usage, prompting awareness, and perceived productivity — analyzed with descriptive statistics and simple correlational checks. The claim is carried by three self-report measures: belief that prompt clarity improves results (Table 12, mean 4.01/5), belief that AI speeds up work (Table 13, mean 3.87/5), and satisfaction with AI output (Table 11, mean 3.24/5). The paper uses these Likert-scale responses as the bridge between prompting behavior and productivity.

What would settle it

Run the same set of writing or coding tasks with users randomly assigned to vague prompts versus structured, context-rich prompts, and compare objective completion time and independently scored output quality; if structured prompts show no measurable advantage, the paper's central causal claim is refuted.

Watch

Extended reading notes

Core claim

The paper claims to show that users who employ clear, structured, and context-aware prompts report higher task efficiency and better outcomes with LLMs. In the survey, 83.7% of respondents agreed or strongly agreed that clearer and more specific prompts lead to better AI results, and 75.7% agreed that AI helps them complete tasks faster. The study also reports that role prompting, chain-of-thought prompting, and instruction prompting are the most-used techniques, and that over half of respondents revise their prompts often or occasionally. The author interprets these patterns as evidence that human input quality, not just model capability, is decisive in realizing productivity gains from generative AI.

Load-bearing premise

The study assumes that respondents' agreement that clear prompts improve results and that AI speeds up their work is a reliable measure of actual prompting behavior and real productivity gains, never linking reported technique use to objective task outcomes.

Editorial extensions

If this is right

  • If the claim holds, teaching prompt design as a basic digital skill should raise the value people get from LLMs in education and work.
  • Tool designers could reasonably add prompt-guiding interfaces, templates, and revision feedback to improve outcomes for users who do not prompt well spontaneously.
  • The finding implies employers and educators should invest in prompt literacy programs rather than leaving users to learn by trial and error.
  • Users' habit of iterative revision suggests that human-in-the-loop refinement is part of how LLM productivity is actually realized.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Extending beyond the paper, a controlled experiment comparing vague and structured prompts on identical writing or coding tasks could show whether the perceived gains match objective gains in speed and output quality.
  • The finding that education level tracks prompting breadth suggests a testable extension: short prompt-training sessions may narrow the productivity gap between less- and more-educated users.
  • If perceived productivity is what drives continued AI adoption, then even subjective gains could be economically meaningful, because they shape user engagement and acceptance.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript reports a self-report survey of 243 AI users, with sections covering respondent demographics, AI usage patterns, prompting techniques, prompt revision frequency, satisfaction, and perceived productivity benefits. The authors' central claim is that users who employ clear, structured, and context-aware prompts experience higher task efficiency and better outcomes, and they conclude that prompt engineering is a critical determinant of AI-assisted productivity. The analysis, however, is almost entirely descriptive: it reports frequencies and means but never statistically connects reported prompting behavior to any productivity outcome.

Significance. If the central claim were properly supported, the paper would offer practical guidance for AI literacy programs and workplace prompt-training initiatives. The manuscript provides a clear taxonomy of manual versus automatic prompting techniques (Table 1) and a diverse, internationally distributed convenience sample, which is useful as a descriptive snapshot. However, the core relationship between prompting behavior and productivity is never tested. The paper's main 'finding' is essentially a restatement of respondents' agreement with a survey item, so the study's current contribution is limited to documenting perceptions rather than establishing the claimed effect. No reproducible code or machine-checked derivations are included.

major comments (4)
  1. [§4.3–§4.4 and Abstract] The central claim that users who employ clear, structured, and context-aware prompts report higher task efficiency and better outcomes is not tested anywhere in the paper. Table 8 records which prompting techniques respondents say they use and Table 9 records revision frequency, while Tables 11–14 report satisfaction and perceived benefits. No cross-tabulation, correlation, regression, or significance test connects Table 8 or Table 9 with any productivity variable. Section 3.6 explicitly promises 'Correlation and Trend Analysis' using scipy.stats, but the Results chapter contains no such analysis. As a result, the abstract's conclusion is not derivable from the data presented.
  2. [§4.4, Table 12] The key evidence for 'clear prompts lead to better outcomes' is respondents' agreement with the statement 'more specific and clearer prompts lead to better AI results.' This is circular: the conclusion is effectively the survey item itself. Self-reported belief is neither a measure of prompting behavior nor an objective measure of task outcomes. To support the central claim, the authors would need to compare outcome ratings across groups that differ in reported prompting behavior, such as users of particular techniques versus non-users, or high-frequency versus low-frequency revisers.
  3. [§4.3, Table 10] The crosstabulation of education level and prompting techniques in Table 10 counts mentions rather than unique respondents, so its rows are not coherent with the sample sizes reported in Table 3. For example, the Bachelor's degree row sums to 129, yet only 83 respondents hold a Bachelor's degree. Consequently, the statement in §4.5 that 'educational background influences prompting strategy diversity' is unsupported by the table as presented.
  4. [§3.3, §4.5] The paper repeatedly describes its design as 'descriptive quantitative' and 'exploratory,' yet the Abstract and Section 4.5 use causal and evaluative language such as 'lead to clearer outcomes' and 'critical factor.' A descriptive survey of beliefs and self-reported satisfaction cannot support causal claims about the effect of prompt structure on productivity; either the conclusions must be tempered or the missing inferential analysis must be provided.
minor comments (5)
  1. [Abstract] The phrase 'rapid engineering using scalable language models' appears to be a typo; it should presumably read 'prompt engineering.'
  2. [§3.5] The Data Collection Procedure states that data collection spanned six weeks (05 January 2025 to 10 February 2025), but later says 'At the end of the three-week period.' This is internally inconsistent.
  3. [§4.5] The text says 'more than 66% of respondents (160 out of 243)' use AI at least twice a week, but 160/243 ≈ 65.8%, so 'approximately 66%' would be accurate.
  4. [§4.4, Table 14] Table 14 reports means and standard deviations only; the interpretive sentence that these results 'affirm the hypothesis' is overconfident without any inferential tests or confidence intervals.
  5. [§4.4] The survey relies entirely on self-report, but the manuscript does not discuss common-method bias, social desirability, or self-selection of respondents, all of which are salient for the validity of the stated conclusions.

Circularity Check

2 steps flagged · score 7.0 of 10

Central claim reduces to respondents' agreement with the same claim; no behavioral test links prompt use to productivity.

  1. self definitional [§4.4 Perceived Productivity and Effectiveness, Table 12]
    "To assess whether user input plays a critical role in output quality, participants were asked if more specific and clearer prompts lead to better responses. Table 12 presents these results. ... A significant 83.7% (203 out of 243) of respondents agreed or strongly agreed that clearer and more specific prompts lead to better AI results. This supports the premise that prompt engineering is not just a technical skill but a determinant of successful AI-human collaboration."

    The survey item records belief in the hypothesis ('clearer and more specific prompts lead to better AI results'), and that same belief is then presented as evidence for the hypothesis. No independent outcome measure is compared across groups with different prompting behavior, so the finding is the survey question itself restated as a result.

  2. other [§4.4 and §4.5 Summary of Findings]
    "The high average ratings for prompt effectiveness (mean = 4.01) and work efficiency (mean = 3.87) affirm the hypothesis that user prompting strategies directly impact productivity outcomes."

    These means are aggregated self-reports of perceived effectiveness and efficiency, not measurements of productivity tied to actual prompting strategies. Table 8 (technique usage) and Table 9 (revision frequency) are never cross-tabulated with any productivity variable, so the claimed association between employing clear structured prompts and higher efficiency is not derived from the data; it is only a restatement of average opinions.

full rationale

The paper's central conclusion is that users who employ clear, structured, context-aware prompts achieve higher task efficiency and better outcomes. The only evidence offered for this link is respondents' belief that clearer prompts produce better results (Table 12) and their self-reported efficiency gains (Tables 13-14). No statistical comparison connects actual prompting behavior (Table 8) or revision behavior (Table 9) to any productivity outcome; the planned correlational analysis in §3.6 is not reported. The first step is circular by construction: the survey item asks for agreement with the conclusion, and agreement with the conclusion is then used as confirmation of the conclusion. The second step adds an unsupported causal reading of averages. There is no self-citation chain or imported uniqueness theorem; the circularity is the use of the dependent variable as the independent variable. Score 7 because the central claim reduces to a self-report of the same claim, though some peripheral descriptive findings are independently described.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The study treats self-reported perceptions as outcomes and assumes a convenience sample is representative. No free parameters or invented entities are present; the main burden is the measurement assumption that beliefs about prompt utility are valid proxies for actual productivity effects.

assumptions (3)
  • domain assumption Self-reported perceptions are valid measures of productivity and prompt effectiveness
    The analysis uses Likert-scale self-reports (Tables 11-14) as the outcome variables for productivity and output quality, without objective validation (see §4.4).
  • domain assumption Convenience sample is representative of AI users
    The paper generalizes to 'users' in §4.5 despite using non-probability convenience sampling described in §3.2.
  • domain assumption Respondents accurately recall and report prompting behavior
    Data on prompting techniques and revision frequency in §4.3 are retrospective self-reports, never verified against logs or observations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity." pith.science (2026). https://pith.science/paper/BVSLZOLJ

@misc{pith2026250718638,
  author       = {Pith},
  title        = {Pith review of: Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BVSLZOLJ}},
  note         = {Machine review of arXiv:2507.18638}
}
read the original abstract

The widespread adoption of large language models (LLMs) such as ChatGPT, Gemini, and DeepSeek has significantly changed how people approach tasks in education, professional work, and creative domains. This paper investigates how the structure and clarity of user prompts impact the effectiveness and productivity of LLM outputs. Using data from 243 survey respondents across various academic and occupational backgrounds, we analyze AI usage habits, prompting strategies, and user satisfaction. The results show that users who employ clear, structured, and context-aware prompts report higher task efficiency and better outcomes. These findings emphasize the essential role of prompt engineering in maximizing the value of generative AI and provide practical implications for its everyday use.

Figures

Figures reproduced from arXiv: 2507.18638 by the authors.

Figure 1
Figure 1. Distribution of respondents by education level [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 10 canonical work pages

  1. [1]

    A Practical Survey on Zero-shot Prompt Design for In-context Learning,

    Z. Liu et al., “A Practical Survey on Zero-shot Prompt Design for In-context Learning,” arXiv preprint arXiv:2309.13205, 2023. [Online]. Available: https://arxiv.org/abs/2309.13205

  2. [2]

    Fairness-guided Few-shot Prompting for Large Language Models,

    Y . Zhang et al., “Fairness-guided Few-shot Prompting for Large Language Models,” Tencent AI Lab, 2023. [Online]. Available: https://ailab.tencent.com/ailab/media/publications/Fairness-guided_Few-shot_ Prompting_for_Large_Language_Models.pdf

  3. [3]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

    J. Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” arXiv preprint arXiv:2201.11903, 2022. [Online]. Available: https://arxiv.org/abs/2201.11903

  4. [4]

    Guidelines for Prompting Large Language Models,

    P. Pandey, “Guidelines for Prompting Large Language Models,” Medium, 2023. [Online]. Available:https:// medium.com/@pankaj_pandey/guidelines-for-prompting-large-language-models-b598189abed5

  5. [5]

    Assigning Roles to Chatbots,

    Learn Prompting, “Assigning Roles to Chatbots,” LearnPrompting.org. [Online]. Available: https:// learnprompting.org/docs/basics/roles

  6. [6]

    Automatic Prompt Engineer (APE),

    T. Yao et al., “Automatic Prompt Engineer (APE),” Prompt Engineering Guide. [Online]. Available: https: //www.promptingguide.ai/techniques/ape

  7. [7]

    Prompt Tuning, Hard Prompts and Soft Prompts,

    C. Greyling, “Prompt Tuning, Hard Prompts and Soft Prompts,” Medium, 2023. [Online]. Available: https: //cobusgreyling.medium.com/prompt-tuning-hard-prompts-soft-prompts-49740de6c64c

  8. [8]

    RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning,

    C. Li et al., “RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning,” Carnegie Mellon University, 2023. [Online]. Available: https://blog.ml.cmu.edu/2023/02/24/ rlprompt-optimizing-discrete-text-prompts-with-reinforcement-learning

Show all 13 references
  1. [9]

    Automatic Prompt Optimization with ‘Gradient Descent’ and Beam Search,

    S. Yao et al., “Automatic Prompt Optimization with ‘Gradient Descent’ and Beam Search,” arXiv preprint arXiv:2305.03495, 2023. [Online]. Available: https://arxiv.org/abs/2305.03495

  2. [10]

    Enhancing English Comprehension through Generative AI and Prompt Engineering: A Study on Undergraduate Learning Outcomes,

    H. Lee and Q. Zhang, “Enhancing English Comprehension through Generative AI and Prompt Engineering: A Study on Undergraduate Learning Outcomes,” Education Sciences, vol. 14, no. 2, p. 199, 2024. [Online]. Available: https://www.mdpi.com/2227-7102/14/2/199 14 Prompt Engineering...

  3. [11]

    Mastering generative AI: Why effective prompting is the key to success at work,

    G. Morrison and Y . Lee, “Mastering generative AI: Why effective prompting is the key to success at work,” Business Horizons, vol. 67, no. 1, pp. 23–34, 2024. [Online]. Available: https://www.sciencedirect.com/ science/article/pii/S0007681324000533

  4. [12]

    The Critical Role of Prompt Engineering in the Workplace,

    Moor Insights & Strategy , “The Critical Role of Prompt Engineering in the Workplace,” Whitepaper, 2024. [Online]. Available: https://moorinsightsstrategy.com/research-notes/ the-critical-role-of-prompt-engineering-in-the-workplace

  5. [13]

    The Cognitive Effects of AI-Human Collaboration: A Behavioral Study on Prompt Design and Decision-Making,

    Y . Wang and J. Kim, “The Cognitive Effects of AI-Human Collaboration: A Behavioral Study on Prompt Design and Decision-Making,” SSRN Electronic Journal, 2024. [Online]. Available: https://papers.ssrn.com/sol3/ papers.cfm?abstract_id=5140787 15

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.