{"id":"02d04345-c2a8-482f-8dc8-fb0b98d8aee7","arxiv_id":"2508.11678","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 87 peer-grading studies finds random assignment is the most common strategy and that 3-5 reviews per submission balances grading accuracy with student workload.","lead":"This paper reviews 87 studies on peer grading to find how reviewers are assigned and how many are best. It concludes that random assignment is most common, and that three to five reviews per submission balances accuracy with student workload.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '3–5 reviews' recommendation rests on Figure 3, which double-counts range reports and conflates reviews per submission with reviews per student.","rationale":"The reader correctly identifies the narrow search strategy as a limitation and notes the PRISMA flow inconsistency. My stress-test focuses on a more direct threat to the headline 3–5 recommendation: the counting methodology in §3.3. The range double-counting and the per-submission/per-student conflation are internal to the paper's own analysis and can be checked without new literature. If the re-analysis shifts the mode or weakens the 3–5 majority, the central claim is materially affected. If it does not, the recommendation stands. Because the reader already issued a CONDITIONAL verdict, my concern does not change the verdict; it reinforces the need for the authors to supply a reproducible, unit-consistent histogram before the quantitative claim is accepted.","tokens_in":9140,"tokens_out":2288,"duration_ms":24755,"concrete_test":"Reconstruct the Figure 3 histogram from the 87 included studies, coding each study's reviewer-number datum as a single value: use the midpoint for range reports, and separate 'reviews per submission' from 'reviews per student' into distinct categories. If the modal value remains 3 and the 3–5 grouping still covers a majority of studies when counted this way, the recommendation survives; if the mode shifts or the 3–5 grouping no longer dominates, the §5 conclusion loses its quantitative support and should be softened to a qualitative summary.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion in §5—'Three to five reviews per submission strikes an effective balance'—depends directly on the histogram in Figure 3 and its interpretation in §3.3. That section states two problematic counting rules: (1) when a study reports a range such as '3–5 reviews per submission,' the authors count every value within that range (3, 4, and 5) as separate data points; (2) they explicitly 'treat [reviews per submission and reviews per student] as equivalent' for simplicity. Both rules inflate the counts for mid-range values. A study reporting '3–5 reviews per submission' contributes to all three bars, making '3–5' appear more common than it actually is. More importantly, the recommendation is about the number of reviewers per submission, but many cited studies report the number of reviews required per student (e.g., Mi et al. [45], Fang et al. [46], Lee et al. [47]). These are different quantities—one review per student can yield one review per submission only if the class size and assignment ratio are equal—and conflating them can bias the recommended range. The PRISMA flow inconsistency (text says 9 papers excluded from 112, diagram says 25 excluded, leaving 87) further obscures exactly which 87 studies form the evidence base. Because the 3–5 claim is the paper's headline practical recommendation, its quantitative support should be re-derived without these counting shortcuts.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This systematic literature review (2010–2024, 87 included studies) examines two design choices in peer grading: how reviewers are assigned to submissions and how many reviews are requested per submission. It proposes a four-strategy taxonomy (random, competency-based, social-network-based, bidding), reports that random assignment is the most common strategy (~75% of systems that specify a strategy), and analyzes the effect of the number of reviewers on accuracy, workload, and learning. The paper concludes that 3–5 reviews per submission strikes an effective balance, and that competency-based assignment offers accuracy/fairness advantages over random assignment, while social-network and bidding approaches remain under-evaluated. The review follows the PRISMA framework and reports inter-rater agreement for screening and full-text selection.","tokens_in":9522,"tokens_out":4079,"duration_ms":48103,"significance":"If the conclusions hold, the paper provides a useful synthesis for instructors and platform designers, addressing a gap relative to prior reviews that focused on anonymity or post-hoc grade correction. The four-strategy taxonomy and the concrete examples from 87 studies are valuable. Strengths include a PRISMA-structured protocol, explicit inclusion/exclusion criteria, dual screening with reported inter-rater kappa, and a distinction between accuracy, fairness, workload, and learning outcomes. The main limitations are methodological: the headline quantitative claims rest on two counting conventions in §3.3 that inflate and conflate the evidence, and the search strategy is narrow enough that the prevalence estimates may not be representative. These issues are fixable but require re-analysis rather than copy-editing.","major_comments":[{"comment":"The central conclusion that 'three to five reviews per submission strikes an effective balance' is directly supported by the histogram in Figure 3, but the counting rules stated in §3.3 bias that histogram. The text says that for a reported range such as 3–5, 'we count each value within a range', so a single study contributes to three separate bars. This inflates the counts for 3, 4, and 5 and makes the mid-range appear more common than it actually is. In addition, the text declares 'For simplicity, we treat them as equivalent' when referring to reviews per submission and reviews per student, although these quantities differ whenever the number of submissions and number of students differ, or when calibration reviews are assigned. The recommendation should be re-derived by reporting one value per study (e.g., the midpoint or a consistent choice of fixed value), and by separating per-subm","section":"§3.3, Figure 3"},{"comment":"The study selection numbers are internally inconsistent. The text states that after full-text review 'Nine papers were later excluded', yet Figure 1 reports '25 studies excluded' at the eligibility stage and '87 studies included'. Since 112 papers were assessed, 112−25=87, while 112−9=103. The text and figure cannot both be correct, and this ambiguity obscures exactly which 87 studies form the evidence base. The same inconsistency affects the denominator for the 75% prevalence claim in §3.1: Figure 2 shows 53 random, 15 competency, 2 social-network, and 1 bidding, summing to 71; the text says 'For the other 16 papers, there was no explicit mention'. The 75% is therefore 53/71 among studies that specify a strategy, not 53/87 of all included studies. Please correct the flow diagram and state clearly whether prevalence is computed over all 87 or over the 71 with an identifiable strategy.","section":"§2.2 and Figure 1"},{"comment":"The search was restricted to the exact keywords 'peer grading' and/or 'peer marking', English-language studies, and seven databases, with no snowballing or gray literature. The Limitations paragraph acknowledges that non-English and gray literature were excluded, but the more specific risk is that adjacent terminology—'peer feedback', 'peer evaluation', 'calibrated peer review'—is common in this field and was not searched. Because the 75% random-assignment prevalence and the 3–5 recommendation are aggregate claims over the included corpus, the narrow search is a load-bearing threat to external validity. A concrete test would be to run supplementary searches with the broader terms and check whether the distribution of strategies and the distribution of review counts change materially; if they do not, the claim can stand. This should be reported in the revision.","section":"§2.2, Limitations section"}],"minor_comments":[{"comment":"The PRISMA flow diagram uses '626 studies imported for screening' but the text says 626 publications were returned; the wording is inconsistent. Also, the diagram reports 238 studies irrelevant after abstract screening, while the text says 350 were screened and 112 selected; the arithmetic (350−238=112) should be made explicit.","section":"§2.2"},{"comment":"The sentence 'For the other 16 papers, there was no explicit mention of reviewer assignment strategies' is important but the denominator for the 75% figure is not stated. Please report the denominator explicitly (71 or 87) and consider stating that 16 papers were excluded from the strategy-trend analysis.","section":"§3.1"},{"comment":"The phrase 'the matrices of competency' (near citations [21] and [35]) should be 'the criteria for competency' or 'the indicators of competency'.","section":"§3.2"},{"comment":"Reference [41] is listed as 'A. A. V. Ioannis Caragiannis, George A. Krimpas, ...' which appears malformed; the author list should be checked for consistency with the preceding reference [40]. Also, 'Jingjing et. al' in §3.1 should be 'Jingjing et al.' per the citation style.","section":"References and text"},{"comment":"The statement 'In most cases, these numbers are identical' is presented without a citation or a worked definition. Since the paper later relies on this identity for Figure 3, either add a derivation or qualify the statement explicitly as an assumption.","section":"§3.3"},{"comment":"The x-axis label reads 'Number of Reviewers per Submission' but the text in §3.3 conflates per-submission and per-student counts. The caption should indicate the unit actually plotted after the re-analysis, or separate the two cases.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a CS/education venue and makes a useful contribution despite the methodological concerns. The key issue is that the headline recommendation rests on a histogram whose counting rules are problematic; this can be corrected by re-analyzing the data at the study level. The PRISMA inconsistency and the narrow search should also be fixed. I would not reject because the qualitative taxonomy and the qualitative discussion of trade-offs are sound and independently useful, and the quantitative claims are clearly fixable with a moderate amount of work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: useful but uneven systematic review. The assignment-strategy taxonomy is a real contribution, but the headline 3–5 recommendation doesn't survive close reading of the paper's own quantitative evidence.\n\nWhat's genuinely new: no prior SLR I know of organizes peer-grading research around reviewer-assignment strategy and review count as the two organizing axes. The four categories—random, competency-based, social-network, bidding—are sensible, and the mapping of 87 studies onto them looks careful. The qualitative trade-offs are described fairly: random's simplicity and fairness problems, competency's promise when performance data exist, social and bidding as plausible but under-evaluated. The inter-rater agreement reporting is a plus.\n\nThe soft spots are concentrated in Section 3.3. The central conclusion—'three to five reviews per submission strikes an effective balance'—rests on Figure 3, and Figure 3 is built on two counting rules that inflate the evidence. First, a reported range like 3–5 is counted as three separate data points, so the mid bars are systematically overfilled. Second, reviews per submission and reviews per student are treated as equivalent, even though they're different quantities when class size and assignment configuration vary. Some of the cited studies (Mi, Fang, Lee) are about required reviews per student, not per submission. The recommendation may still be right, but the paper's own histogram doesn't provide the support it claims. This is a load-bearing flaw for the main practical takeaway.\n\nTwo smaller issues. The PRISMA narration says 9 papers were excluded from 112 while the diagram says 25; since 112 minus 25 equals 87, the diagram is probably right, but the text needs fixing. The search (two keywords, English only, no snowballing) almost certainly missed work using terms like 'peer feedback' or 'calibrated peer review.' That's a coverage limitation worth stating more prominently, though the qualitative synthesis probably wouldn't change much. Note that the 75% random-assignment figure is computed over the 71 studies that explicitly name a strategy, which is the correct denominator; that particular number is not an error.\n\nWho's this for? Course designers and MOOC platform teams, not theorists. It gives them a credible map of the literature and a defensible starting range, once the counting is redone. It deserves a serious referee because the taxonomy is genuinely useful and the review is mostly careful. I'd send it out, but with a required revision: re-derive the review-count analysis without the double-counting and per-student/per-submission conflation, and clean up the search protocol description.","headline":"A useful taxonomy of reviewer assignment strategies, but the headline 3–5 review recommendation is not supported by the paper's own counting.","tokens_in":9931,"tokens_out":4462,"would_cite":true,"duration_ms":43685,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic review of 87 studies (2010–2024) argues that three to five reviews per submission balance grading accuracy against student workload, and that competency-based reviewer assignment is a fairer alternative to the dominant rando","keywords":["peer grading","reviewer assignment","competency-based assignment","random assignment","number of reviewers","systematic literature review","MOOC","grading fairness"],"falsifier":"Run a large-scale randomized experiment on a MOOC: assign thousands of submissions to 2, 3, 5, or 8 reviewers, then compare peer grades with instructor grades and track reviewer completion and feedback quality. If accuracy continues to climb materially beyond five reviewers while engagement stays flat, or if two reviewers match expert scores, the 3–5 sweet spot and the accuracy-plateau claim would be refuted.","tokens_in":9095,"feed_emoji":"🎓","tokens_out":10856,"duration_ms":115222,"temperature":0.7,"pith_summary":"Peer grading scales large classes, but its quality depends on who reviews whom and how many reviews each submission receives. This paper synthesizes 87 studies from 2010–2024 to argue that the design choices made before grading—reviewer assignment and review count—do more to prevent noisy grades than post-hoc correction. It finds that random assignment, used by about 75% of systems, is simple but error-prone; competency-based matching, while less common, is associated with fairer and more accurate grades. On quantity, three to five reviews per submission is the recurring recommendation: fewer than three loses reliability, more than five adds workload and disengagement with little accuracy gain. The review positions this range as a practical design target for MOOCs and large classes.","feed_headline":"Three to five reviews per submission is peer grading's sweet spot","feed_subtitle":"A systematic review of 87 studies finds random assignment is common but not best; competency-based matching is fairer.","key_machinery":"The machinery is the review's four-part taxonomy of reviewer-assignment strategies—random, competency-based, social-network-based, and bidding—combined with an accuracy-versus-review-count curve assembled from the included studies. The taxonomy organizes what systems actually do; the curve, which rises steeply from one to about four reviews and flattens after five to six, is what locates the three-to-five recommendation. The paper also uses this structure to identify where evidence is missing, particularly for social and bidding methods.","core_discovery":"The paper's central claim is empirical and synthetic: across 87 included studies, reviewer-assignment strategy and review count are the two preventive design levers with the largest effect on peer-grading quality. It identifies four strategies—random, competency-based, social-network-based, and bidding—and finds random assignment in about 75% of systems despite inconsistent grading and fairness problems. Competency-based assignment, which matches reviewers by prior performance, domain knowledge, calibration accuracy, or reputation, mitigates skewed grading and appears in about 21% of systems. For quantity, three reviews per submission is the most common configuration, and the literature conv","pith_inferences":["Beyond the paper: if accuracy saturates near five reviews, the practical optimum for platforms with calibration data may be lower—three reviews plus calibration could match five uncalibrated reviews, so the 3–5 range is not a hard optimum.","Beyond the paper: the 75% random-assignment figure suggests most deployed peer-grading platforms have adopted the least-optimal strategy; an A/B switch to competency-based matching on an existing course could test this claim directly.","Beyond the paper: re-running the review with broader search terms such as 'peer feedback' or 'peer evaluation' would show whether the 3–5 consensus and the strategy prevalences are artifacts of keyword choice."],"forward_implications":["Platforms that currently assign reviews randomly can expect inconsistent grading; moving to competency-based matching is the paper's recommended preventive fix.","Three reviews is the minimum that appears to match expert grading in several studies; below that, grading validity suffers.","Requiring more than five reviews per student risks rushed, low-effort grading and disengagement without meaningful accuracy gains.","Social-network and bidding assignment methods should be piloted and evaluated before being deployed at scale.","Calibration and weighting techniques can make even three-to-five-review settings accurate, so review count and post-processing should be designed together."],"supporting_citations":[{"why":"Shows that three peer reviewers can produce scores statistically equivalent to expert scores (R² > 0.8), anchoring the lower bound of the 3–5 range.","marker":"[2]"},{"why":"Assigns high-reputation, low-reputation, and random reviewers to each submission, showing how assignment mix affects peer-grading accuracy.","marker":"[6]"},{"why":"Reports RMSE falling sharply from 1 to 4 reviewers, moderately to 6, then marginally beyond, supplying the diminishing-returns shape behind the recommendation.","marker":"[7]"},{"why":"Describes grouping students by cognitive indicators and assigning reviewers from matching clusters, a concrete competency-based assignment approach.","marker":"[21]"},{"why":"Separates reviewers into independent and supervised pools by prior grading accuracy, showing competency-based assignment in practice.","marker":"[23]"},{"why":"Documents accuracy improving from 83% to 93% when reviews increase from 4 to 10, supporting the general claim that more reviews improve accuracy.","marker":"[41]"},{"why":"Recommends a cap of five reviews per student to preserve review quality and avoid student burnout.","marker":"[45]"},{"why":"Finds five reviewers per submission reliable when combined with a hybrid model using some teaching-assistant grades, without overburdening students.","marker":"[46]"},{"why":"Recommends four reviews per submission as the balance point between workload and accuracy.","marker":"[48]"},{"why":"Reports substantial accuracy gains up to five reviewers, moderate gains to eight, and negligible gains after, supporting the plateau claim.","marker":"[50]"}],"fun_headline_variants":["Peer grading's sweet spot: 3-5 reviews per submission","Competency-based reviewers beat random in peer grading, review finds","Random peer grading is common but unfair: 87-study review","For fair peer grading, match reviewers by skill, not luck","How many reviews per peer-graded work? 3-5, says new review"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The whole analysis rests on the assumption that searching seven databases with only the terms 'peer grading' and 'peer marking' in English captured a representative slice of the relevant literature; if studies using other labels like 'peer feedback' or 'peer evaluation' were missed, the prevalence estimates and the 3–5 recommendation could be biased.","fun_headline_variants_meta":{"raw":{"variants":["Peer grading's sweet spot: 3-5 reviews per submission","Competency-based reviewers beat random in peer grading, review finds","Random peer grading is common but unfair: 87-study review","For fair peer grading, match reviewers by skill, not luck","How many reviews per peer-graded work? 3-5, says new review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000568,"raw_usage":{"total_tokens":2521,"prompt_tokens":736,"completion_tokens":1785,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":1693}},"tokens_in":480,"tokens_out":1785,"duration_ms":11494,"temperature":1.0,"reasoning_tokens":1693,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:43:22.058029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a large-scale randomized experiment on a MOOC: assign thousands of submissions to 2, 3, 5, or 8 reviewers, then compare peer grades with instructor grades and track reviewer completion and feedback quality. If accuracy continues to climb materially beyond five reviewers while engagement stays flat, or if two reviewers match expert scores, the 3–5 sweet spot and the accuracy-plateau claim would be refuted.","supporting_citations":[{"cited_title":"Validity of peer grading using calibrated peer review in a guided-inquiry, conceptual physics course,","cited_arxiv_id":null,"evidence_quote":"Shows that three peer reviewers can produce scores statistically equivalent to expert scores (R² > 0.8), anchoring the lower bound of the 3–5 range."},{"cited_title":"Machine and social intelligent peer-assessment systems for assessing large student populations in massive open online education,","cited_arxiv_id":null,"evidence_quote":"Assigns high-reputation, low-reputation, and random reviewers to each submission, showing how assignment mix affects peer-grading accuracy."},{"cited_title":"Improving peer grading reliability with graph mining techniques,","cited_arxiv_id":null,"evidence_quote":"Reports RMSE falling sharply from 1 to 4 reviewers, moderately to 6, then marginally beyond, supplying the diminishing-returns shape behind the recommendation."},{"cited_title":"Peer assessment in moocs based on learners’ profiles clustering,","cited_arxiv_id":null,"evidence_quote":"Describes grouping students by cognitive indicators and assigning reviewers from matching clusters, a concrete competency-based assignment approach."},{"cited_title":"Mechanical ta: Partially auto- mated high-stakes peer grading,","cited_arxiv_id":null,"evidence_quote":"Separates reviewers into independent and supervised pools by prior grading accuracy, showing competency-based assignment in practice."},{"cited_title":"Aggregating partial rankings with applications to peer grading in massive online open courses","cited_arxiv_id":"1411.4619","evidence_quote":"Documents accuracy improving from 83% to 93% when reviews increase from 4 to 10, supporting the general claim that more reviews improve accuracy."},{"cited_title":"Probabilistic graphical models for boosting cardinal and ordinal peer grading in moocs,","cited_arxiv_id":null,"evidence_quote":"Recommends a cap of five reviews per student to preserve review quality and avoid student burnout."},{"cited_title":"Rankwithta: A robust and accurate peer grading mechanism for moocs,","cited_arxiv_id":null,"evidence_quote":"Finds five reviewers per submission reliable when combined with a hybrid model using some teaching-assistant grades, without overburdening students."},{"cited_title":"Vista, E","cited_arxiv_id":null,"evidence_quote":"Recommends four reviews per submission as the balance point between workload and accuracy."},{"cited_title":"Appropriate number of raters for irt based peer assessment evaluation of programming skills,","cited_arxiv_id":null,"evidence_quote":"Reports substantial accuracy gains up to five reviewers, moderate gains to eight, and negligible gains after, supporting the plateau claim."}],"review_version":1}