REVIEW 4 major objections 3 minor 2 cited by
Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement
T0 review · 4 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This paper claims that a closed loop of automated scoring and manual reference checks makes ChatGPT write better research proposals with fewer fabricated citations.
desk verdict A small empirical study of iterative prompting for research proposal writing, but the unvalidated grading loop and tiny sample cannot support the objective-quality claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-metric evaluation loop, the dual-metric framework: content quality is computed as the average score from three AI grading platforms (Study Fetch, QuillBot, Grammarly), and reference validity is computed by the authors manually verifying each generated citation for authenticity and format. The iterative prompting method then treats these scores as feedback: the AI grader's detailed scores on clarity, fluency, and organization are re-prompted into ChatGPT, and a reference guide listing author, title, year, journal, volume, page, and publisher corrections is given to the model to learn from. The loop repeats until scores stabilize, which for references the results place a
What would settle it
Ask domain experts to rank the before and after versions of the same proposal without knowing which is which, and independently verify each corrected reference against a bibliographic database. If expert rankings do not track the automated score gains, or if a substantial share of 'corrected' citations still cannot be found, the iterative loop is not doing what the paper claims.
Extended reading notes
Core claim
The central discovery is that ChatGPT-4o's research-proposal writing can be improved in a closed loop: generate a proposal, score its content with an automated grading system that averages outputs from Study Fetch, QuillBot, and Grammarly, manually check every reference for accuracy and fabrication, then feed both sets of findings back as prompts telling the model what to fix. Using this loop, the paper reports average content scores rising from 81–85 to 86–91 across three topics, with the largest gain an 11.02% increase on one assisted-writing proposal, and reference correctness improving from as low as 38.89% to 44.4%, 80%, or 100% depending on the proposal. The paper also reports that in-
Load-bearing premise
The paper assumes the three commercial grading platforms' average is a true measure of research-proposal content quality, and that the two authors' manual citation checks are error-free ground truth; if the graders reward style over substance or the checks miss errors, the reported improvements are not real improvements.
Editorial extensions
If this is right
- Researchers using LLMs to draft proposals can deploy the same loop—automated content scoring plus a manual citation check, then re-prompting—to nudge quality upward.
- The results imply that asking LLMs to self-check or merely instructing them to cite 'primary sources' is not enough; external scoring and fact-checking are needed to keep drafts honest.
- Reference fabrication in LLM-generated academic drafts can be reduced to near zero in small-scale proposal-writing tasks, taking on a specific ethical concern in academic publishing.
- The loop offers a quantitative benchmark for comparing writing strategies (GPT-only vs. GPT-assisted with human-provided references) in terms of both content quality and reference validity.
- Improvement is bounded and saturates: reference correctness stabilizes by the third prompting round, suggesting a practical stopping rule for refinement.
Reading between the lines
- The same loop should transfer to other structured writing tasks that demand verifiable citations, such as grant applications, literature reviews, and clinical summaries, because the feedback vocabulary (grammar, fluency, clarity, citation accuracy) is similar.
- The reported gains depend on three grading tools whose internal weights are opaque; a direct head-to-head comparison with expert human reviewers would reveal whether the loop improves substantive argument quality or only surface polish.
- The manual reference-guide step could plausibly be automated with retrieval-based citation checkers that query digital libraries, making the loop faster and removing the human ground-truth assumption.
- The stabilization after three rounds suggests an adaptive stopping rule for user-facing tools: run the loop until the score delta falls below a threshold, then stop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two evaluation metrics for LLM-generated research proposals—content quality, computed by averaging scores from Study Fetch, QuillBot, and Grammarly, and reference validity, established by manual fact-checking by the first two authors—and an iterative prompting method that feeds these scores back into ChatGPT to refine the proposals. The experiments cover six proposals (three education topics × two writing strategies: GPT-only and GPT-assisted). The authors claim that the dual-metric framework is an “objective, quantitative” way to assess ChatGPT’s writing performance and that iterative prompting improves content quality and reduces reference fabrications. The paper includes Table 2 with content-quality scores, Table 3 with reference-correctness rates, and Figures 2–3 showing improvement over prompting rounds.
Significance. If the claims held, the practical contribution would be modest but useful: a case study showing that an automated feedback loop can move commercial text-quality scores and that iterative reference checking can clean citation lists. The paper addresses a real problem—hallucinated references in LLM-assisted academic writing—and the authors document their manual fact-checking effort and provide the query templates in the appendix. However, the central claims are not supported by the evidence as presented. The content-quality metric is never validated against expert human judgment, the improvement loop is evaluated on the same scores used as feedback, and the empirical basis is six proposals with no statistical analysis. The paper is better read as an exploratory case study than as a demonstration of an “objective, quantitative framework.”
major comments (4)
- [Metrics Assessing GPT's Writing Capability] The content-quality metric averages scores from Study Fetch, QuillBot, and Grammarly without any validation that these platforms measure research-proposal content quality. No evidence is reported that their scores correlate with expert human judgments of full-length proposals, that they are calibrated for this genre, or that averaging the three yields a meaningful construct. The abstract’s claim of an “objective, quantitative framework” therefore conflates reproducibility with validity. Because this score is the outcome variable for RQ1 and the feedback signal for RQ2, the central results in Table 2 and Figure 2 are unanchored. A human-validated gold standard or external criterion is required before the metric can support the paper’s claims.
- [Iterative Prompting Method / Evaluation of Iterative Prompting Method] The improvement loop feeds the same three commercial scores back into ChatGPT and then measures improvement on those same scores. This creates a circularity problem: the loop may simply optimize the graders’ proprietary rubrics, which need not correspond to substantive proposal quality. There is no control condition—for example, repeated generation without feedback, or feedback from expert human raters—so the gains in Figure 2 cannot be attributed to the method. An independent human-evaluation arm and a control prompting condition are necessary.
- [Experimental Settings / Performance Evaluation] The empirical basis is six proposals (3 topics × 2 strategies), one run each, with no confidence intervals, no repeated trials, and no significance tests. Percentages in Table 3 are computed from very small counts (4–18 references), and observed differences such as 40% versus 38.89% are within sampling noise. The statement that results “consistently indicate a clear enhancement” overstates what can be inferred from Figure 2, which has no error bars. At minimum, repeated generations and bootstrapped intervals or expert-rated independent samples are needed before quantitative claims can be made.
- [Validity of References] Reference validity is established by manual fact-checking by the first two authors, but no inter-rater reliability is reported, no coding protocol is given, and the raw counts are small. The manual check is treated as error-free ground truth, which is questionable. Additionally, the in-text “file name” errors (e.g., citing “Document 1”) appear to be induced by the AskYourPDF workflow in the GPT-assisted condition, not necessarily by the model’s reference generation; this conflates an input-format artifact with the claimed reduction of fabrication. The reference-correctness findings should be re-analyzed with a more rigorous annotation procedure and with the input format controlled.
minor comments (3)
- [Throughout] Typos and formatting issues: “belows” in the contributions list, “itrative” in Methodology, “LLMs offers” in the opening sentence, “AgencyandIdentity” missing spaces, “Table 2)” stray parenthesis, and reference [6] appears incomplete. The phrase “extensive experiments” is an overstatement for six proposals.
- [Figures 2 and 3] Figure 2 has no numeric labels on the Before/After axis and no error bars. Figure 3’s legend omits Student Agency (GPT-only), although the text discusses that condition. Both figures should be labeled to show the number of runs and the variance.
- [Appendix] The iterative prompting prompts are not included; the appendix only gives the initial GPT-only and GPT-assisted generation queries. For reproducibility, the authors should provide the actual feedback prompts, the stopping criterion, the number of rounds, and the model temperature/settings used in iteration.
Circularity Check
Iterative prompting is evaluated with the same AI scores used as feedback, and reference-validity ground truth is the same authors who supply corrections.
-
fitted input called prediction
[Methodology > Iterative Prompting Method; Performance Evaluation > Evaluation of Iterative Prompting Method > Content Quality (Figure 2)]
"Content quality was evaluated using an AI-based grading process, with scores averaged across three professional platforms: Study Fetch, QuillBot [18], and Grammarly [19]. ... the scores obtained from the AI scoring systems are fed back to GPT as targeted feedback, focusing on specific writing weaknesses such as clarity and fluency. ... The revised proposal is then re-evaluated using the same AI scoring systems."
The content-quality score is operationally defined as the average of the same three AI graders (Study Fetch, QuillBot, Grammarly) that are used both as the feedback signal and as the outcome measure. The iterative loop feeds these scores to GPT and then reports improvement on these exact scores (Figure 2). Thus the reported 'enhancement in content quality' is the model's success at optimizing the same reward function it was prompted to maximize; the evaluation is not independent of the intervention, so the improvement is at least partly self-fulfilling rather than evidence of externally validated proposal quality.
-
self definitional
[Methodology > Metrics Assessing GPT's Writing Capability (Reference Validity); Methodology > Iterative Prompting Method; Performance Evaluation > Evaluation of Iterative Prompting Method > Reference V]
"Reference validity was assessed through a manual fact-checking process conducted by the first two authors. ... After each research proposal was generated, the authors provided a reference guide for review. This iterative process was applied to refine GPT’s outputs by correcting errors in format, author, title, year, journal, volume, page number, and publisher, identifying fabricated references..."
Reference validity is established by the two authors' manual fact-check, and the same authors' corrections are the intervention that 'improves' reference correctness. The authors supply GPT with a reference guide that corrects errors and removes fabrications, then re-check the final output with the same manual procedure. The reported increase in correctness (e.g., Teacher Agency GPT-only rising from 40% to 80%) counts the authors' own corrections as correct. The reduction in fabrications is therefore by construction an artifact of the ground-truth generator being identical to the correction source, not an independent model capability.
full rationale
The paper's central claims—that iterative prompting improves content quality and reduces reference fabrication—are supported by evaluations that are not independent of the interventions. Content quality is measured by the same AI grading platforms whose scores are fed back to GPT; the reported gains are improvements on the optimization target, so they do not validate real-world writing quality beyond the metric. Reference validity relies on the two authors' manual fact-check both as the ground truth and as the source of corrections; the reported error reduction is essentially the authors correcting the references and then verifying their own corrections. No self-citation chain or imported uniqueness theorem is present. These are not mere validity concerns: the outcome metrics are constructed from the same inputs that drive the refinement loop, making the claimed improvements partially self-fulfilling. Because the central claims reduce in part to the self-referential evaluation setup, a score of 6 is appropriate.
Assumptions & free parameters
free parameters (2)
- Stopping threshold for quality =
not specified
- Number of prompting rounds =
3
assumptions (3)
- ad hoc to paper Study Fetch, QuillBot, and Grammarly scores reflect research proposal content quality
- domain assumption Manual fact-checking by the two authors is accurate and unbiased
- domain assumption ChatGPT-4o's outputs on different days are stable enough for comparison
Cite this review
Pith. "Pith review of Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement." pith.science (2026). https://pith.science/paper/5SS2PUAP
@misc{pith2026250909709,
author = {Pith},
title = {Pith review of: Assisting Research Proposal Writing with Large Language Models: Evaluation and Refinement},
year = {2026},
howpublished = {\url{https://pith.science/paper/5SS2PUAP}},
note = {Machine review of arXiv:2509.09709}
}
read the original abstract
Large language models (LLMs) like ChatGPT are increasingly used in academic writing, yet issues such as incorrect or fabricated references raise ethical concerns. Moreover, current content quality evaluations often rely on subjective human judgment, which is labor-intensive and lacks objectivity, potentially compromising the consistency and reliability. In this study, to provide a quantitative evaluation and enhance research proposal writing capabilities of LLMs, we propose two key evaluation metrics--content quality and reference validity--and an iterative prompting method based on the scores derived from these two metrics. Our extensive experiments show that the proposed metrics provide an objective, quantitative framework for assessing ChatGPT's writing performance. Additionally, iterative prompting significantly enhances content quality while reducing reference inaccuracies and fabrications, addressing critical ethical challenges in academic contexts.
Forward citations
Cited by 2 Pith papers
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review
In a controlled test, three mid-2025 LLMs shared under 6% of literature references with physics experts, and 64% of their real references had at least one metadata error.
-
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
AI-generated one-page research proposals are scored about the same as human-written ones by human reviewers, but AI reviewers favor AI-written proposals by roughly one point and detect authorship perfectly.
Reference graph
Works this paper leans on
-
[1]
C. Song and Y . Song, “Enhancing academic writing skills and motivation: assessing the efficacy of chat- gpt in ai-assisted language learning for efl students,” Frontiers in Psychology, vol. 14, p. 1260843, 2023. 6 Assisting Research Proposal Writing with LLMs May 2025 THEME/FEATURE/DEPARTMENT
work page 2023
-
[2]
O. D. Awosanya, A. Harris, A. Creecy, X. Qiao, A. J. Toepp, T. McCune, M. A. Kacena, and M. V. Ozanne, “The utility of ai in writing a scientific review article on the impacts of covid-19 on musculoskeletal health,” Current Osteoporosis Reports, vol. 22, no. 1, pp. 146–151, 2024
work page 2024
-
[3]
Academic writing with gpt-3.5 (chatgpt): reflections on practices, efficacy and transparency,
O. Buruk, “Academic writing with gpt-3.5 (chatgpt): reflections on practices, efficacy and transparency,” inProceedings of the 26th International Academic Mindtrek Conference, 2023, pp. 144–153
work page 2023
-
[4]
The use of artificial intelligence in writing scien- tific review articles,
M. A. Kacena, L. I. Plotkin, and J. C. Fehrenbacher, “The use of artificial intelligence in writing scien- tific review articles,”Current Osteoporosis Reports, vol. 22, no. 1, pp. 115–121, 2024
work page 2024
-
[5]
J. Bell and S. Waters,Ebook: doing your research project: a guide for first-time researchers. McGraw- hill education (UK), 2018
work page 2018
-
[6]
“Reports outline management research from osma- nia university (effective strategies for crafting re- search proposals in higher education),” pp. 554–, 2024
work page 2024
-
[7]
Exploring chatgpt as a writing assessment tool,
J. L. Bucol and N. Sangkawong, “Exploring chatgpt as a writing assessment tool,”Innovations in Educa- tion and Teaching International, pp. 1–16, 2024
work page 2024
-
[8]
Z. Du and K. Hashimoto, “Exploring sentence-level revision capabilities of llms in english for academic purposes writing assistance,” 2024
work page 2024
Show all 20 references
-
[9]
Brainstorming will never be the same again—a human group supported by arti- ficial intelligence,
F . Lavriˇc and A. Škraba, “Brainstorming will never be the same again—a human group supported by arti- ficial intelligence,”Machine Learning and Knowledge Extraction, vol. 5, no. 4, pp. 1282–1301, 2023
2023
-
[10]
Graduate teacher education students use and evaluate chatgpt as an essay-writing tool
A. G. Picciano, “Graduate teacher education students use and evaluate chatgpt as an essay-writing tool.” Online Learning, vol. 28, no. 2, p. n2, 2024
2024
-
[11]
Chatgpt revisited: Using chatgpt-4 for finding references and editing language in medical scientific articles,
O. M. Alyasiri, A. M. Salman, S. Salisuet al., “Chatgpt revisited: Using chatgpt-4 for finding references and editing language in medical scientific articles,”Jour- nal of Stomatology, Oral and Maxillofacial Surgery, p. 101842, 2024
2024
-
[12]
Exploring the potential of using an ai language model for automated essay scoring,
A. Mizumoto and M. Eguchi, “Exploring the potential of using an ai language model for automated essay scoring,”Research Methods in Applied Linguistics, vol. 2, no. 2, p. 100050, 2023
2023
-
[13]
Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities,
M. Lee, P . Liang, and Q. Y ang, “Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities,” inProceedings of the 2022 CHI conference on human factors in computing systems, 2022, pp. 1–19
2022
-
[14]
What is the impact of chatgpt on edu- cation? a rapid review of the literature,
C. K. Lo, “What is the impact of chatgpt on edu- cation? a rapid review of the literature,”Education Sciences, vol. 13, no. 4, p. 410, 2023
2023
-
[15]
Man is to computer programmer as woman is to homemaker? debiasing word embed- dings,
T. Bolukbasi, K.-W. Chang, J. Y . Zou, V. Saligrama, and A. T. Kalai, “Man is to computer programmer as woman is to homemaker? debiasing word embed- dings,”Advances in neural information processing systems, vol. 29, 2016
2016
-
[16]
Today’s academic research: The role of chatgpt writing,
J. E. Chukwuere, “Today’s academic research: The role of chatgpt writing,”Journal of Information Sys- tems and Informatics, vol. 6, no. 1, pp. 30–46, 2024
2024
-
[17]
Chat- ting and cheating: Ensuring academic integrity in the era of chatgpt,
D. R. Cotton, P . A. Cotton, and J. R. Shipway, “Chat- ting and cheating: Ensuring academic integrity in the era of chatgpt,”Innovations in education and teaching international, vol. 61, no. 2, pp. 228–239, 2024
2024
-
[18]
The impact of quillbot as an automated writing evaluation tool on efl learners,
N. Gürbüz, “The impact of quillbot as an automated writing evaluation tool on efl learners,”Journal of Ed- ucational Studies and Multidisciplinary Approaches, vol. 4, no. 2, 2024
2024
-
[19]
Ding and D
L. Ding and D. Zou, “Automated writing evaluation systems: A systematic review of grammarly, pigai, and criterion with a perspective on future directions in the age of generative artificial intelligence,”Education and Information Technologies, vol. 29, no. 11, pp. 14 151–14 203, 2024
2024
-
[20]
What is agency? conceptualizing pro- fessional agency at work,
A. Eteläpelto, K. Vähäsantanen, P . Hökkä, and S. Paloniemi, “What is agency? conceptualizing pro- fessional agency at work,”Educational research re- view, vol. 10, pp. 45–65, 2013. Jing Renis a PhD student in Education from Uni- versity of Technology Sydney. Contact her at ji...
2013
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.