{"id":"ddca263d-2f68-4746-a8e4-9fb56ef6f80d","arxiv_id":"2507.20655","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CoGrader combines human-AI metric design, benchmark-driven regrading, and AI-assisted feedback to support instructors in grading project reports.","lead":"CoGrader is a new system that lets instructors use an AI language model to design grading criteria, score reports, compare work against benchmark examples, and draft feedback, while keeping final decisions with the instructor. A study of 12 instructors found high satisfaction and grades similar to existing course grades, but there was no comparison group grading without the tool.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's causal efficiency/consistency claim is not supported: the study has no no-tool control, and the paper's own Table 1 shows wide letter-grade disagreement across participants using CoGrader.","rationale":"The reader's REJECT verdict is sound. My stress-test converges on the same core weakness but frames it slightly differently: the load-bearing gap is causal inference, not just ground-truth validity. Even if we accept the ground truth as a valid standard, the study provides no comparator, so high correlations with ground truth cannot show that CoGrader improved efficiency or consistency; they could simply reflect the shared grading culture of the same institution and course. I add one piece of internal evidence: Table 1's letter-grade distributions show substantial inter-grader disagreement (e.g., R03 spanning four grade levels), so even the consistency claim is not strongly supported by the paper's own quantitative data. The proposed controlled comparison directly tests the causal claim; until that evidence exists, the abstract overstates what was found. I do not question the authors' integrity, and I credit the coherent system design, the formative study, and the detailed engagement logging as useful contributions. The verdict should remain as the reader set it: REJECT, because the central claim is not supported by the current evaluation.","tokens_in":24878,"tokens_out":8790,"duration_ms":102190,"concrete_test":"Run a counterbalanced within-subjects study in which the same participants grade the same five reports twice: once with the full CoGrader workflow and once with their usual manual workflow (or a rubric-only control), recording objective time per report, number of revisions, and inter-rater agreement (e.g., Krippendorff's alpha or exact/adjacent letter-grade agreement). If CoGrader does not significantly reduce grading time and inter-rater variance relative to control, the abstract's improvement claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CoGrader improves grading efficiency and consistency. To establish this, the evaluation must compare against the instructor's normal workflow. Section 5 does not: all 12 participants grade the same five reports with CoGrader, efficiency is measured only by self-report Likert items (Q9-Q20, Figures 7-8), and the Section 5.5 consistency analysis reports only correlations with ground truth (Kendall tau=0.799, Spearman=0.899, Pearson=0.900) without a manual-grading baseline or inferential statistics. Correlations with ground truth are a validity check, not a measure of consistency, and they are confounded by recruitment: ground truth came from the same course's instructor and two senior TAs (Section 5.2), and most participants had taught or TA'd that same data visualization course (Section 5.3), so agreement may reflect shared standards or familiarity rather than an effect of CoGrader. Table 1 independently weakens the consistency claim: letter grades for report R03 span A- to C across 12 graders, and for R05 only 2 of 12 participants assign the exact ground-truth A-. The feedback-reliability claim is likewise not tested with students; only informal discussions are mentioned in Section 6.2. The design and engagement analysis (64.09% override) are valuable, but they do not supply the missing causal comparison.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents CoGrader, a human-AI collaborative grading workflow for project reports. A formative study with six instructors motivates six design requirements, and CoGrader implements collaborative metric design, LLM-based pre-grading, instructor-selected benchmark reports, benchmark-driven regrading, and feedback generation. The evaluation is a 12-participant user study in which instructors grade five reports from a master-level data visualization course; the paper reports Likert-scale ratings of each workflow component, correlations between participant scores and the course's ground truth grades, letter-grade agreement, and interaction logs. The paper concludes that CoGrader improves grading efficiency and consistency and provides reliable peer-comparative feedback to students.","tokens_in":25087,"tokens_out":6269,"duration_ms":71439,"significance":"If the effectiveness claims were adequately supported, the contribution would be useful: project-report grading is a real and underserved problem, and the proposed division of labor between LLM-generated references and instructor-held authority is well motivated. The formative study, the system design, and the interaction-log analysis (e.g., the 64.09% override rate for AI-generated reference scores and comments) are valuable and worth building on. However, the empirical evaluation does not support the causal claims in the abstract, and the paper's own data undermine the consistency claim. The strengths are real but they are strengths of a system-description and formative-study paper, not of an effectiveness evaluation.","major_comments":[{"comment":"The abstract's causal claim that CoGrader 'improves grading efficiency and consistency' is not supported by the study design. All 12 participants used CoGrader; there is no no-tool or manual-grading control arm, no objective measure of time-to-grade, and no measurement of grading drift before versus after using the tool. Efficiency is measured only by self-report Likert items (Q9–Q20, Figures 7–8). A within-subjects comparison against the instructors' normal grading workflow, or at least objective interaction-time logs against a control condition, is needed before any improvement claim can be made.","section":"§5.1–5.5"},{"comment":"The consistency analysis conflates agreement with an external ground truth with grading consistency. The correlations are computed over only five reports, no confidence intervals or inferential statistics are reported, and the label 'internal consistency within individual graders' is misleading because the comparison is to an external standard. More directly, Table 1 undercuts the consistency claim: letter grades for report R03 span A- to C across 12 graders, and for R05 only 2 of 12 participants match the ground-truth A-. These are evidence against inter-grader consistency, not evidence for it.","section":"§5.5, Table 1"},{"comment":"The ground-truth comparison is confounded by recruitment. The ground truth grades come from the same course's instructor and two senior teaching assistants (§5.2), and participants were recruited at the authors' institution, with most having taught or served as teaching assistants for that same data visualization course (§5.3). Agreement with these grades may therefore reflect shared internal standards or prior familiarity with the assignment rather than any effect of CoGrader.","section":"§5.2–5.3"},{"comment":"The claim of 'reliable peer-comparative feedback to students' is not tested with students. Section 6.2 reports only informal discussions with students, and Section 6.5 acknowledges that LLM-generated grading was observed to be inconsistent across attempts (P11). To support the feedback and consistency claims, a student-facing study of feedback usefulness and a reliability analysis (e.g., repeated grading runs or inter-rater agreement statistics such as ICC or Krippendorff's alpha) are required.","section":"§6.2, §6.5"}],"minor_comments":[{"comment":"The term 'reliance' is used where 'reliability' appears to be intended (e.g., RQ2 in §5, Figure 8 heading); please correct this throughout.","section":"Throughout"},{"comment":"The 'Assigned Grades (Ps)' row lists counts such as '9A 2A- 1B+' without explicitly mapping each count to the ordered grade labels; a legend or table heading would make the counts interpretable.","section":"Table 1"},{"comment":"Please clarify how the three correlation values were computed: whether they are averaged across participants, pooled across reports, or computed per participant; the current wording makes the sample size and aggregation ambiguous.","section":"§5.5"},{"comment":"The paper states that ground truth scores were a weighted average of evaluations from one instructor and two senior TAs, but it does not state the weights; since this grade is used as the external standard throughout, the weighting should be specified.","section":"§5.2"}],"recommendation":"reject","confidential_remarks":"The system design and formative study are genuine contributions, and the interaction-log analysis is a strong point. The reason for rejection is the mismatch between the causal effectiveness claims and the evaluation: the missing control condition and the contradictory Table 1 cannot be repaired by re-analysis alone. A future submission with a controlled comparison, objective outcome measures, and more measured claims could warrant reconsideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth reading for the system design, but the central claim—that CoGrader improves grading efficiency and consistency—is not backed by the evidence as presented. The design integrates metric co-design, benchmark-driven regrading, and feedback generation into one instructor-controlled workflow, which looks new relative to the cited literature. The formative study is reasonable, and the engagement analysis (64.09% override rate) is a genuinely useful observation about how instructors interact with AI-generated scores.\n\nThe soft spots are load-bearing. There is no no-tool baseline: all participants used CoGrader, so efficiency is only self-reported Likert ratings. The consistency analysis correlates each participant's grades with the course's ground truth, which is a validity check, not a measure of consistency; Table 1 actually shows wide letter-grade spread across participants (report 03 spans A- to C). The ground truth comes from the same course and most participants taught or TA'd that course, so agreement may reflect shared standards rather than the tool. The feedback-reliability claim rests on informal student discussions, not a formal test.\n\nThese are serious gaps, but not fatal to the paper as a systems contribution. The authors are explicit about limitations, and the interface and workflow are coherent. The paper would benefit from a baseline condition or at least stronger wording in the abstract.\n\nWho this is for: HCI researchers working on AI-assisted assessment and designers of instructor-in-the-loop LLM tools. A serious referee should engage with it; it deserves peer review, with requests for a baseline or tempered claims.","headline":"A well-designed instructor-in-the-loop LLM grading system whose central efficiency/consistency claim is undercut by the absence of a no-tool baseline and reliance on self-report.","tokens_in":25649,"tokens_out":1457,"would_cite":false,"duration_ms":18065,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that CoGrader's workflow — LLM-assisted metric design, instructor-chosen benchmarks, and benchmark-driven regrading — makes project-report grading more efficient and consistent while producing reliable peer-comparative…","keywords":["project report grading","large language models","human-AI collaboration","LLM for education","benchmark-driven grading","AI-assisted feedback","educational assessment"],"falsifier":"A controlled experiment would settle it: recruit two matched groups of instructors with no prior connection to the course, have one group grade the same reports with CoGrader and the other with a conventional rubric only, and compare time per report, inter-grader agreement, and agreement with the official grades; the paper reports no control condition, and its efficiency numbers are self-rated Likert scores. A cheaper check targets consistency directly: run the identical regrade prompt on the same five reports several times and measure score variance, since the paper cites evidence that LLM grading output is not stable across attempts.","tokens_in":24657,"feed_emoji":"🎓","tokens_out":13988,"duration_ms":135340,"temperature":0.7,"pith_summary":"CoGrader aims to show that large language models can help with the part of teaching that resists automation: grading open-ended project reports on subjective criteria such as creativity, originality, and how well a student applied course knowledge. Its proposed answer is a division of labor in which the instructor keeps final authority while the model drafts the routine parts — extracting grading metrics from the assignment sheet, producing initial scores and comments, and regrading every report against the high and low benchmark reports the instructor picks. The central claim is that this workflow improves grading efficiency and consistency and yields reliable peer-comparative feedback for students. A sympathetic reader would care because project-based learning is spreading while the grading workload and the risk of drifting standards are growing; if the claim holds, instructors get a scalable way to grade more fairly and students get feedback anchored to concrete examples of good and weak work. The paper's evidence is a 12-instructor study with positive Likert ratings and strong correlations with the course's original grades.","feed_headline":"Instructor-set benchmarks make LLM grading faster and more consistent","feed_subtitle":"Instructors keep final say while an LLM drafts metrics, scores, and per-student feedback anchored to class benchmarks.","key_machinery":"The load-bearing mechanism is the benchmark-driven regrading loop. In the first pass the LLM drafts scores and comments per metric; the instructor corrects these, then designates one strong submission as Benchmark High and one weak submission as Benchmark Low. The 'Regrade Reports' action then sends every report back through the model with explicit instructions to compare the target report's current evaluation against the two benchmarks and adjust scores and comments accordingly, so each updated comment carries the shape 'this report is stronger than the low benchmark on criterion X because...'. Radar charts comparing the focused report with both benchmarks, plus sortable score-distribution views, let the instructor audit the comparisons at a glance. This loop is what the paper credits for both halves of its claim: consistency, because every report is measured against the same instructor-approved reference points instead of a drifting internal standard, and efficiency, because the most tedious step — revisiting and rescoring earlier reports after benchmarks change — becomes a single automated call. Everything else in CoGrader (metric extraction, feedback synthesis) feeds or consumes this loop.","core_discovery":"The paper's central claim is that the right unit for LLM assistance in grading is the workflow, not the individual grade: a model alone cannot judge design innovation or practical knowledge application, but a model embedded in a loop where the instructor sets the metrics and chooses the reference points can. CoGrader implements this as three connected stages. In metric design, the LLM analyzes the project requirements into objective metrics and suggests extra potential ones, while the instructor selects, edits, adds, and labels each metric as auto-grade or score-reference. In benchmarking-driven grading, the LLM drafts per-metric scores and comments, the instructor reviews and corrects them, selects a Benchmark High and a Benchmark Low, and triggers a regrade in which the model re-evaluates every report explicitly against those two benchmarks, producing comparative comments that cite why a report outperforms or falls short. In feedback generation, the system consolidates the instructor's annotations, edited scores, and AI suggestions into personalized feedback. The evaluation claims that instructors perceived the workflow as efficient and reliable (means of 5.9–6.5 on a 7-point scale across components), that their scores aligned with the original course grades (Kendall's $\\tau = 0.799$, Spearman's $\\rho = 0.899$, Pearson $r = 0.900$), and that they engaged critically rather than deferring — adjusting or overriding 64.09% of AI scores and comments on the subjective reference metrics. The paper itself notes that LLM outputs were not fully stable across attempts and that comments could come out generic, which is why it argues for iterative calibration as future work.","pith_inferences":["The 64.09% override rate on subjective reference metrics cuts against a strong reading of 'reliable': the LLM's draft functions as a scaffold that instructors must correct, so a real deployment should budget human review of most subjective scores rather than expecting to accept AI output wholesale.","The study has no control arm, and its efficiency numbers are self-reported Likert ratings; a direct with/without comparison measuring time per report and inter-grader spread would be the natural next experiment.","A side benefit the paper leaves implicit: the instructor-validated benchmark reports and metric sets become reusable course assets, letting a department accumulate calibrated reference corpora that stabilize grading across semesters and sections.","Scope caution: the evidence covers five reports, one data-visualization course, one model, and participants drawn from the same institution as the ground truth — the workflow may transfer to other subjective assessments, but the effect sizes should not be expected to transfer without re-measurement."],"forward_implications":["When an instructor revises a benchmark mid-way, the regrade function automatically re-scores every report against the updated reference, eliminating the most time-consuming retroactive step in the current manual workflow.","Students receive feedback anchored to concrete exemplars, with comments that state why a report meets, exceeds, or falls short of the nominated high and low benchmarks on each metric.","Because instructors retain authority over metric selection, benchmark choice, score edits, and final feedback, the design is a direct counter to documented failure modes of fully automated grading, such as LLM self-favoritism and generic comments.","If the workflow holds up at larger scale, a high-enrollment project course could give every student comparative, per-metric feedback at a cost that manual grading could not sustain — the drafting and consolidation are automated while the judgment stays human."],"supporting_citations":[{"why":"Supplies the GPT-4o API that powers all three LLM agents (metric design, grading, feedback generation).","marker":"[63]"},{"why":"Supplies Kendall's tau, one of the three correlation statistics used to claim agreement with the ground-truth grades.","marker":"[6]"},{"why":"Supplies Spearman's rank correlation used in the grading-consistency analysis.","marker":"[75]"},{"why":"Supplies Pearson correlation used in the grading-consistency analysis.","marker":"[74]"},{"why":"Documents the LLM self-favoritism bias that motivates the design decision to keep instructors in control of final scores.","marker":"[66]"},{"why":"Provides the retrieval-augmented-generation method behind the file-search tool that lets the LLM quote evidence from reports.","marker":"[16]"},{"why":"Supplies the critique of rubric-based grading (bias and drift) that the benchmark-regrading workflow is designed to address.","marker":"[65]"},{"why":"The LLM-Rubric approach to rubric-driven automated evaluation, the closest baseline that CoGrader extends with human-in-the-loop benchmarking.","marker":"[31]"}],"fun_headline_variants":["LLM grading improves when instructors set the benchmarks","CoGrader: instructor-calibrated LLM grading for project reports","Grading with LLM: instructors set metrics, benchmarks, and final say","Human-LLM loop: instructors tune benchmarks to grade fairly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the premise that the course's official grades — a weighted average of one instructor and two senior teaching assistants' evaluations — are a valid external standard and that agreement with them shows CoGrader improves grading, even though the participants came from the same institution and several had taught or assisted that same course, so the measured agreement could just reflect shared internal standards.","fun_headline_variants_meta":{"raw":{"variants":["LLM grading improves when instructors set the benchmarks","CoGrader: instructor-calibrated LLM grading for project reports","Grading with LLM: instructors set metrics, benchmarks, and final say","Human-LLM loop: instructors tune benchmarks to grade fairly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000604,"raw_usage":{"total_tokens":2878,"prompt_tokens":1066,"completion_tokens":1812,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":1739}},"tokens_in":682,"tokens_out":1812,"duration_ms":16679,"temperature":1.0,"reasoning_tokens":1739,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:22:14.783798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled experiment would settle it: recruit two matched groups of instructors with no prior connection to the course, have one group grade the same reports with CoGrader and the other with a conventional rubric only, and compare time per report, inter-grader agreement, and agreement with the official grades; the paper reports no control condition, and its efficiency numbers are self-rated Likert scores. A cheaper check targets consistency directly: run the identical regrade prompt on the same five reports several times and measure score variance, since the paper cites evidence that LLM grading output is not stable across attempts.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GPT-4o API that powers all three LLM agents (metric design, grading, feedback generation)."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Pearson correlation used in the grading-consistency analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the LLM self-favoritism bias that motivates the design decision to keep instructors in control of final scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the critique of rubric-based grading (bias and drift) that the benchmark-regrading workflow is designed to address."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The LLM-Rubric approach to rubric-driven automated evaluation, the closest baseline that CoGrader extends with human-in-the-loop benchmarking."}],"review_version":1}