REVIEW 4 major objections 5 minor 33 references
Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Teaching assistants in an introductory economics course said that AI-generated feedback, offered as suggestions with highlighted evidence and rubric judgments, could speed up grading, make it more consistent, and improve feedback quality.
desk verdict A carefully framed qualitative study whose perceptions-based claims hold up; the stress-test overstates the overreach, but real limitations remain. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a three-step, rubric-wise feedback pipeline that exposes intermediate outputs. For each rubric item, the LLM first extracts the sentences in the essay most relevant to that rubric, then renders a judgment on whether the essay satisfies it, and finally drafts a feedback message informed by established feedback guidelines (specific language, praise, guiding questions). These intermediate outputs — the highlighted sentences and the judgment — are displayed as in-text Word comments alongside the feedback. Showing the evidence lets TAs quickly verify or overturn the AI's judgment, which is what makes the faulty AI output usable as a suggestion.
What would settle it
Run the same course for a full term with the AI suggestion tool embedded in the live grading workflow, randomly assigning essays to TA-grading-with-AI versus TA-grading-without-AI, and measure per-essay grading time, inter-TA agreement on rubric scores, and the number of feedback comments that survive TA editing. If grading time does not drop, inter-TA agreement does not rise, or TAs reject most AI suggestions, the perceived benefits would not translate into practice.
Extended reading notes
Core claim
The central claim is that AI-generated feedback, despite its errors, can serve as effective suggestions in human grading workflows, provided that the generation process is decomposed into visible intermediate steps. The authors built a feedback engine that, for each rubric item, identifies relevant essay sentences, judges whether the rubric is met, and then drafts a feedback message; the sentences and judgment are shown as in-text highlights and comments in Word. In think-aloud sessions, TAs found the AI feedback more rubric-aligned and more personalized than both their own comments and the instructors' historic feedback, while also noting that rubric rigidity can mislead. TAs said they would use such suggestions to save time writing comments, to catch rubric items they missed, and to standardize grading across TAs, as long as they review the essay first and can see the AI's evidence.
Load-bearing premise
The study assumes that what TAs say during think-aloud sessions about AI feedback on already-graded essays matches how they would actually behave when grading live with the tool, and that the five volunteer TAs are representative of all TAs in the course.
Editorial extensions
If this is right
- TAs would adopt AI feedback as a starting point, editing or combining messages rather than writing from scratch, which they expect to cut the time spent composing comments.
- Because AI feedback is rubric-aligned and point-to-point, it can act as a second pair of eyes that catches rubric criteria a TA may have overlooked, supporting within-TA and between-TA consistency.
- Rubric quality becomes a bottleneck: for knowledge-intensive essays, detailed rubrics that spell out domain knowledge, acceptable alternatives, and expected depth are necessary for the AI to generate accurate feedback.
- Presenting intermediate outputs (highlighted sentences and judgments) is a usable form of AI explainability that helps TAs spot and correct hallucinations without reading extra explanatory text.
- TAs plan to read the essay and form their own judgment before consulting AI, meaning the envisioned workflow is human-first with AI as a suggestion layer, not AI-first automation.
Reading between the lines
- The study's positive results rest on self-report; a live deployment where grading time and inter-rater reliability are measured could show whether the perceived speedup and consistency gains materialize in actual grading sessions.
- The double-edged sword of rubric alignment suggests a testable design principle: AI feedback may be more valuable for well-specified, atomistic rubric items than for holistic or open-ended criteria, so future systems might calibrate how strongly to weight the rubric versus the essay's holistic quality.
- The highlight-based explainability finding generalizes beyond grading: any LLM output that a human must verify benefits from showing the source evidence inline, a pattern already used in retrieval-augmented systems but rarely formalized as a UI requirement for feedback tools.
- Because TAs said they would combine or discard AI comments, the marginal value of AI may be highest for TAs with less grading experience; a follow-up study could compare novices' and experts' reliance on and correction of AI suggestions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a qualitative user study in which five teaching assistants (TAs) in an introductory economics course graded essays as usual, then participated in 20 think-aloud sessions reviewing AI-generated feedback on a randomly selected subset of essays they had already graded. The AI feedback was produced by a decomposed LLM pipeline that, for each rubric item, highlights relevant essay sentences, makes a rubric-satisfaction judgment, and generates feedback; the feedback was presented as in-text Word comments. The authors find that TAs perceived AI feedback as more praise-oriented, more personalized, and more rubric-aligned than their own feedback, while also noting risks of over-rigid rubric application and fragmented comments. The paper concludes that TAs see promise in a human-AI collaborative workflow in which AI feedback is offered as suggestions, and it derives design implications about writing clearer rubrics, using highlights for transparency, and exposing intermediate outputs. The central claim is framed in the abstract as TAs' perceptions rather than measured outcomes.
Significance. If taken as a perceptual, exploratory study, the paper is a useful contribution to the human-AI partnership literature in education. Its strengths include the use of real course materials and real graded essays, repeated sessions with the same participants, a concrete deployed artifact (the Word plugin), and a public repository link for the feedback pipeline. The study directly addresses a gap by studying longer, knowledge-intensive essays rather than short answers, and it offers concrete design implications that can guide future systems. However, the evidence base is narrow: five self-selected TAs from one course, retrospective think-aloud protocols, and no behavioral measures of grading speed, consistency, or feedback quality. The paper should therefore be read as generating hypotheses about perceived benefit, not as evidence of realized gains; that distinction needs to be maintained consistently in the conclusions and design recommendations.
major comments (4)
- [§3.3, §4.2, §6] The central claim that AI feedback 'could expedite grading' rests entirely on prospective self-reports collected after the TAs had already completed grading. Section 3.3 describes a think-aloud protocol in which participants were shown AI feedback on essays they had already graded and were asked how they envision using such feedback; Section 4.2 quotes such statements as 'I think that would be faster' and 'it definitely would help save time.' No baseline or experimental measure of grading time, throughput, or comparison condition was collected. The conclusion in Section 6 that participants 'perceived them to accelerate grading' is a fair summary, but the earlier statement in Section 6 that the study 'argues for a future where AI feedback can be provided as suggestions' and the design implication that highlights 'speed up reading and assessments' go beyond the data. Please either add a clearly stated limitation that only perceived benefit was measured, or conduct a small behavioral comparison (e.g., grading time with and without AI suggestions) before making claims about accelerated grading.
- [§5, 'Use highlights to increase transparency'] The design implication that highlighting relevant sentences helps TAs evaluate and correct AI feedback is supported only by self-report from a single condition in which highlights were always present. Section 3.3 does not describe a condition without highlights or with only final feedback, so the specific causal role of the intermediate outputs is not tested. For example, the quote from P4 in Section 4.2 about highlights speeding up reading is a reflection on the exhibited interface, not a comparison against an alternative. Please reframe this as an emergent design hypothesis, or add a within-participant comparison with and without highlights, before presenting it as a validated design recommendation.
- [§3.2, §3.3, Figure 2] The paper introduces two feedback variants in Section 3.2: one in which AI only makes judgments used to retrieve historic feedback, and one in which feedback is generated entirely by AI. However, Section 3.3 does not state how these variants were assigned to essays or participants, whether each participant saw both variants, or whether any conclusions in the findings apply to both variants. Figure 2 appears to show both historic and AI feedback together in the same comment, which further obscures what users actually saw. This is a load-bearing omission because the paper's third design implication ('Provide intermediate outputs on each AI subtask') depends on the distinction between variants. Please clarify the experimental or design procedure and, if both variants were explored, report any differences in participant reactions or explicitly state that the design was not a controlled comparison.
- [§3.3, 'Data analysis methods'] The qualitative analysis is described only as two authors interpreting transcripts and grouping notes via affinity diagramming. There is no inter-rater reliability measure, no saturation discussion, and no member-checking or triangulation with the other TAs. Given that several themes in Section 4 are illustrated with single quotes from one or two participants, the reader cannot gauge how robust the themes are. The paper should either report additional validity measures or include a limitations sentence acknowledging that the thematic analysis was interpretive and that themes may depend on the small participant pool.
minor comments (5)
- [Throughout] There are several typos and minor language issues, including 'quailty' in the Introduction, 'thrid' in a rubric example in Section 4.1, 'wholistic' in Section 5, 'explicity' in Table 1, and a malformed heading in Section 4.2 ('one TASev'). A careful proofreading pass would improve the manuscript.
- [§3.3] The text says essays were 'randomly selected' but does not describe the random selection procedure or report how many essays each TA reviewed per session. Please add this detail, as it affects the generalizability of the comments about the AI feedback.
- [§3.2, Figure 2] Figure 2 would benefit from a caption that explicitly labels which parts correspond to the AI judgment, the historic feedback, and the AI feedback, since the body text relies on these distinctions.
- [§4.1] Some quotes are presented without an explicit count of how many participants expressed each theme (e.g., 'Many participants mentioned' versus 'Some participants considered'). Reporting approximate counts or indicating when a theme was raised by a single participant would help the reader calibrate the strength of the evidence.
- [§5] Table 1's 'Good rubric example' and 'Bad rubric example' columns are useful, but the criteria for choosing these examples are not described. A sentence explaining how the examples were selected from the study data would strengthen the design implication.
Circularity Check
No significant circularity: the paper is an empirical user study whose central claims are grounded in think-aloud data, not in a derivation that assumes its conclusion.
full rationale
This paper is an empirical user study rather than a formal derivation, and none of its central claims reduce to its inputs by construction. The main claim is that TAs perceive AI-generated feedback as potentially useful suggestions, supported by think-aloud sessions (§3.3) and direct quotations from participants (§4.2). This is a measurement of self-reported perception, not a prediction derived from fitted parameters. The feedback engine (§3.2) is described as a pipeline with prompts and rubrics, but the paper does not claim to derive TA perceptions from the engine; rather, it reports how TAs reacted to the engine's output. The one self-citation, ReadingQuizmaker [20], is used to motivate the importance of instructors critically reviewing AI-generated content and is not load-bearing for the main findings. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in via citation. The study is self-contained against its stated evidence base: the conclusions are appropriately framed as perceptions and prospects, with limitations acknowledged in the Discussion. Concerns about demand characteristics or the gap between self-report and actual grading behavior would be validity threats, not circularity. Overall, the derivation chain is not circular; the paper makes no claim that would require its inputs to equal its outputs.
Assumptions & free parameters
free parameters (1)
- LLM temperature =
0.05
assumptions (4)
- domain assumption Think-aloud verbalizations reflect the TAs' actual thought processes and evaluation criteria.
- domain assumption The five participating TAs are representative of the ECON101 TA population and of TAs in similar courses.
- domain assumption The affinity diagram analysis produced reliable themes.
- domain assumption GPT-4o with the described prompts and temperature 0.05 produces feedback of sufficient quality for meaningful critique.
Cite this review
Pith. "Pith review of Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use." pith.science (2026). https://pith.science/paper/QPQM5OI6
@misc{pith2026250515596,
author = {Pith},
title = {Pith review of: Exploring LLM-Generated Feedback for Economics Essays: How Teaching Assistants Evaluate and Envision Its Use},
year = {2026},
howpublished = {\url{https://pith.science/paper/QPQM5OI6}},
note = {Machine review of arXiv:2505.15596}
}
read the original abstract
This project examines the prospect of using AI-generated feedback as suggestions to expedite and enhance human instructors' feedback provision. In particular, we focus on understanding the teaching assistants' perspectives on the quality of AI-generated feedback and how they may or may not utilize AI feedback in their own workflows. We situate our work in a foundational college Economics class, which has frequent short essay assignments. We developed an LLM-powered feedback engine that generates feedback on students' essays based on grading rubrics used by the teaching assistants (TAs). To ensure that TAs can meaningfully critique and engage with the AI feedback, we had them complete their regular grading jobs. For a randomly selected set of essays that they had graded, we used our feedback engine to generate feedback and displayed the feedback as in-text comments in a Word document. We then performed think-aloud studies with 5 TAs over 20 1-hour sessions to have them evaluate the AI feedback, contrast the AI feedback with their handwritten feedback, and share how they envision using the AI feedback if they were offered as suggestions. The study highlights the importance of providing detailed rubrics for AI to generate high-quality feedback for knowledge-intensive essays. TAs considered that using AI feedback as suggestions during their grading could expedite grading, enhance consistency, and improve overall feedback quality. We discuss the importance of decomposing the feedback generation task into steps and presenting intermediate results, in order for TAs to use the AI feedback.
Figures
Reference graph
Works this paper leans on
-
[1]
Education Sciences14(2), 148 (2024)
Almasre, M.: Development and evaluation of a custom gpt for the assessment of students’ designs in a typography course. Education Sciences14(2), 148 (2024)
work page 2024
-
[2]
Innovations in Education and Teaching International pp
Almegren, A., Mahdi, H.S., Hazaea, A.N., Ali, J.K., Almegren, R.M.: Evaluating the quality of ai feedback: A comparative study of ai and human essay grading. Innovations in Education and Teaching International pp. 1–16 (2024)
work page 2024
-
[3]
Advances in neural information processing systems33, 1877–1901 (2020)
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Nee- lakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. Advances in neural information processing systems33, 1877–1901 (2020)
2020
-
[4]
Educational Psychologist49(4), 219–243 (2014)
Chi,M.T.,Wylie,R.:TheICAPframework:Linkingcognitiveengagementtoactive learning outcomes. Educational Psychologist49(4), 219–243 (2014)
work page 2014
-
[5]
In: 2023 IEEE International Conference on Advanced Learning Technologies (ICALT)
Dai, W., Lin, J., Jin, H., Li, T., Tsai, Y.S., Gašević, D., Chen, G.: Can large language models provide feedback to students? a case study on chatgpt. In: 2023 IEEE International Conference on Advanced Learning Technologies (ICALT). pp. 323–325. IEEE (2023)
work page 2023
-
[6]
Eigner, E., Händler, T.: Determinants of llm-assisted decision-making. arxiv 2024. arXiv preprint arXiv:2402.17385
arXiv 2024
-
[7]
International Journal of Educational Technol- ogy in Higher Education (20), 55 (2023)
Escalante, J., Pack, A., Barrett, A.: AI-generated feedback on writing: insights into efficacy and enl student preference. International Journal of Educational Technol- ogy in Higher Education (20), 55 (2023)
work page 2023
-
[8]
International Journal for the Scholarship of Teaching and Learning17(1), 18 (2023)
Finkenstaedt-Quinn, S.A., Watts, F.M., Shultz, G.V., Gere, A.R.: A portrait of mwrite as a research program: a review of research on writing-to-learn in stem through the mwrite program. International Journal for the Scholarship of Teaching and Learning17(1), 18 (2023)
work page 2023
Show all 33 references
-
[9]
Routledge (2018)
Hattie, J., Clarke, S.: Visible learning: Feedback. Routledge (2018)
2018
-
[10]
Review of educational research 77(1), 81–112 (2007)
Hattie, J., Timperley, H.: The power of feedback. Review of educational research 77(1), 81–112 (2007)
2007
-
[11]
arXiv preprint arXiv:2405.00302 (2024)
Heickal, H., Lan, A.: Generating feedback-ladders for logical errors in programming using large language models. arXiv preprint arXiv:2405.00302 (2024)
2024 arXiv
-
[12]
In: Pro- ceedings of the 2016 CHI Conference on Human Factors in Computing Systems
Hicks, C.M., Pandey, V., Fraser, C.A., Klemmer, S.: Framing feedback: Choosing review environment features that support high quality peer assessment. In: Pro- ceedings of the 2016 CHI Conference on Human Factors in Computing Systems. pp. 458–469. ACM (2016)
2016
-
[13]
ACM Computing Surveys55(12), 1–38 (2023)
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A., Fung, P.: Survey of hallucination in natural language generation. ACM Computing Surveys55(12), 1–38 (2023)
2023
-
[14]
In: Proceedings of the 17th International Conference on Educational Data Mining
Jia, Q., Cui, J., Du, H., Rashid, P., Xi, R., Li, R., Gehringer, E.: LLM-generated feedback in real classes and beyond: Perspectives from students and instructors. In: Proceedings of the 17th International Conference on Educational Data Mining. pp. 862–867 (2024)
2024
-
[15]
arXiv preprint arXiv:2501.06658 (2025)
Kakarla, S., Borchers, C., Thomas, D., Bhushan, S., Koedinger, K.R.: Compar- ing few-shot prompting of GPT-4 LLMs with BERT classifiers for open-response assessment in tutor equity training. arXiv preprint arXiv:2501.06658 (2025)
2025 arXiv
-
[16]
Cognitive science36(5), 757–798 (2012)
Koedinger, K.R., Corbett, A.T., Perfetti, C.: The knowledge-learning-instruction framework: Bridging the science-practice chasm to enhance robust student learn- ing. Cognitive science36(5), 757–798 (2012)
2012
-
[17]
Computers and Education: Artificial Intelligence6, 100210 (2024) 14 X
Latif, E., Zhai, X.: Fine-tuning ChatGPT for automatic scoring. Computers and Education: Artificial Intelligence6, 100210 (2024) 14 X. Lu et al
2024
-
[18]
Advances in Neural Information Processing Systems33, 9459–9474 (2020)
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.t., Rocktäschel, T., et al.: Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems33, 9459–9474 (2020)
2020
-
[19]
Advances in Neural Information Processing Systems36(2024)
Li, K., Patel, O., Viégas, F., Pfister, H., Wattenberg, M.: Inference-time inter- vention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems36(2024)
2024
-
[20]
In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Lu, X., Fan, S., Houghton, J., Wang, L., Wang, X.: Readingquizmaker: A human- nlp collaborative system that supports instructors to design high-quality reading quiz questions. In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. pp. 1–18 (2023)
2023
-
[21]
arXiv preprint arXiv:2401.12874 (2024)
Luo, H., Specia, L.: From understanding to utilization: A survey on explainability for large language models. arXiv preprint arXiv:2401.12874 (2024)
2024 arXiv
-
[22]
McNichols, H., Lee, J., Fancsali, S., Ritter, S., Lan, A.: Can large language models replicate its feedback on open-ended math questions? arXiv preprint arXiv:2405.06414 (2024)
2024 arXiv
-
[23]
Moggridge, B., Atkinson, B.: Designing interactions, vol. 17. MIT press Cambridge (2007)
2007
-
[24]
Instructional science37, 375–401 (2009)
Nelson, M.M., Schunn, C.D.: The nature of feedback: How different types of peer feedback affect writing performance. Instructional science37, 375–401 (2009)
2009
-
[25]
Journal of Educational Psychology108(8), 1098 (2016)
Patchan, M.M., Schunn, C.D., Correnti, R.J.: The nature of feedback: How peer feedback features affect students’ implementation rate and quality of revisions. Journal of Educational Psychology108(8), 1098 (2016)
2016
-
[26]
Computational Linguistics49(4), 777–840 (2023)
Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G.S., Turc, I., Reitter, D.: Measuring attribution in natural language generation models. Computational Linguistics49(4), 777–840 (2023)
2023
-
[27]
In: International Conference on Artificial Intelligence in Education
Scarlatos,A.,Smith,D.,Woodhead,S.,Lan,A.:Improvingthevalidityofautomat- ically generated feedback via reinforcement learning. In: International Conference on Artificial Intelligence in Education. pp. 280–294. Springer (2024)
2024
-
[28]
Learning and Instruction91, 101894 (2024)
Steiss, J., Tate, T., Graham, S., Cruz, J., Hebert, M., Wang, J., Moon, Y., Tseng, W., Warschauer, M., Olson, C.B.: Comparing the quality of human and ChatGPT feedback of students’ writing. Learning and Instruction91, 101894 (2024)
2024
-
[29]
In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Wang, R., Zhang, Q., Robinson, C., Loeb, S., Demszky, D.: Bridging the novice- expert gap via models of decision-making: A case study on remediating math mis- takes. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lingu...
2024
-
[30]
Advances in Neural Information Processing Systems35, 24824–24837 (2022)
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems35, 24824–24837 (2022)
2022
-
[31]
arXiv preprint arXiv:2407.18328 (2024)
Wu, X., Saraf, P.P., Lee, G.G., Latif, E., Liu, N., Zhai, X.: Unveiling scoring pro- cesses: Dissecting the differences between llms and human graders in automatic scoring. arXiv preprint arXiv:2407.18328 (2024)
2024 arXiv
-
[32]
arXiv preprint arXiv:2501.09824 (2025)
Xu, C., Lin, J., Wu, T., Aleven, V., Koedinger, K.R.: Improving automated feed- back systems for tutor training in low-resource scenarios through data augmenta- tion. arXiv preprint arXiv:2501.09824 (2025)
2025
-
[33]
In: Proceedings of the 19th ACM Conference on Computer- Supported Cooperative Work & Social Computing
Yuan, A., Luther, K., Krause, M., Vennix, S.I., Dow, S.P., Hartmann, B.: Almost an expert: The effects of rubrics and expertise on perceived value of crowdsourced design critiques. In: Proceedings of the 19th ACM Conference on Computer- Supported Cooperative Work & Social Comp...
2016
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.