Pith. sign in

REVIEW 3 major objections 6 minor 53 references

Optimizing Peer Grading: A Systematic Literature Review of Reviewer Assignment Strategies and Quantity of Reviewers

T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This systematic review of 87 studies (2010–2024) argues that three to five reviews per submission balance grading accuracy against student workload, and that competency-based reviewer assignment is a fairer alternative to the dominant rando

desk verdict A useful taxonomy of reviewer assignment strategies, but the headline 3–5 review recommendation is not supported by the paper's own counting. read the letter →

arxiv 2508.11678 v2 pith:3LPW3KIW submitted 2025-08-08 cs.CY

classification cs.CY
keywords peergradingreviewerassignmentcompetency-basedrandomnumberofreviewerssystematicliteraturereviewMOOCfairness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Peer grading scales large classes, but its quality depends on who reviews whom and how many reviews each submission receives. This paper synthesizes 87 studies from 2010–2024 to argue that the design choices made before grading—reviewer assignment and review count—do more to prevent noisy grades than post-hoc correction. It finds that random assignment, used by about 75% of systems, is simple but error-prone; competency-based matching, while less common, is associated with fairer and more accurate grades. On quantity, three to five reviews per submission is the recurring recommendation: fewer than three loses reliability, more than five adds workload and disengagement with little accuracy gain. The review positions this range as a practical design target for MOOCs and large classes.

What carries the argument

The machinery is the review's four-part taxonomy of reviewer-assignment strategies—random, competency-based, social-network-based, and bidding—combined with an accuracy-versus-review-count curve assembled from the included studies. The taxonomy organizes what systems actually do; the curve, which rises steeply from one to about four reviews and flattens after five to six, is what locates the three-to-five recommendation. The paper also uses this structure to identify where evidence is missing, particularly for social and bidding methods.

What would settle it

Run a large-scale randomized experiment on a MOOC: assign thousands of submissions to 2, 3, 5, or 8 reviewers, then compare peer grades with instructor grades and track reviewer completion and feedback quality. If accuracy continues to climb materially beyond five reviewers while engagement stays flat, or if two reviewers match expert scores, the 3–5 sweet spot and the accuracy-plateau claim would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is empirical and synthetic: across 87 included studies, reviewer-assignment strategy and review count are the two preventive design levers with the largest effect on peer-grading quality. It identifies four strategies—random, competency-based, social-network-based, and bidding—and finds random assignment in about 75% of systems despite inconsistent grading and fairness problems. Competency-based assignment, which matches reviewers by prior performance, domain knowledge, calibration accuracy, or reputation, mitigates skewed grading and appears in about 21% of systems. For quantity, three reviews per submission is the most common configuration, and the literature conv

Load-bearing premise

The whole analysis rests on the assumption that searching seven databases with only the terms 'peer grading' and 'peer marking' in English captured a representative slice of the relevant literature; if studies using other labels like 'peer feedback' or 'peer evaluation' were missed, the prevalence estimates and the 3–5 recommendation could be biased.

Editorial extensions

If this is right

  • Platforms that currently assign reviews randomly can expect inconsistent grading; moving to competency-based matching is the paper's recommended preventive fix.
  • Three reviews is the minimum that appears to match expert grading in several studies; below that, grading validity suffers.
  • Requiring more than five reviews per student risks rushed, low-effort grading and disengagement without meaningful accuracy gains.
  • Social-network and bidding assignment methods should be piloted and evaluated before being deployed at scale.
  • Calibration and weighting techniques can make even three-to-five-review settings accurate, so review count and post-processing should be designed together.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if accuracy saturates near five reviews, the practical optimum for platforms with calibration data may be lower—three reviews plus calibration could match five uncalibrated reviews, so the 3–5 range is not a hard optimum.
  • Beyond the paper: the 75% random-assignment figure suggests most deployed peer-grading platforms have adopted the least-optimal strategy; an A/B switch to competency-based matching on an existing course could test this claim directly.
  • Beyond the paper: re-running the review with broader search terms such as 'peer feedback' or 'peer evaluation' would show whether the 3–5 consensus and the strategy prevalences are artifacts of keyword choice.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This systematic literature review (2010–2024, 87 included studies) examines two design choices in peer grading: how reviewers are assigned to submissions and how many reviews are requested per submission. It proposes a four-strategy taxonomy (random, competency-based, social-network-based, bidding), reports that random assignment is the most common strategy (~75% of systems that specify a strategy), and analyzes the effect of the number of reviewers on accuracy, workload, and learning. The paper concludes that 3–5 reviews per submission strikes an effective balance, and that competency-based assignment offers accuracy/fairness advantages over random assignment, while social-network and bidding approaches remain under-evaluated. The review follows the PRISMA framework and reports inter-rater agreement for screening and full-text selection.

Significance. If the conclusions hold, the paper provides a useful synthesis for instructors and platform designers, addressing a gap relative to prior reviews that focused on anonymity or post-hoc grade correction. The four-strategy taxonomy and the concrete examples from 87 studies are valuable. Strengths include a PRISMA-structured protocol, explicit inclusion/exclusion criteria, dual screening with reported inter-rater kappa, and a distinction between accuracy, fairness, workload, and learning outcomes. The main limitations are methodological: the headline quantitative claims rest on two counting conventions in §3.3 that inflate and conflate the evidence, and the search strategy is narrow enough that the prevalence estimates may not be representative. These issues are fixable but require re-analysis rather than copy-editing.

major comments (3)
  1. [§3.3, Figure 3] The central conclusion that 'three to five reviews per submission strikes an effective balance' is directly supported by the histogram in Figure 3, but the counting rules stated in §3.3 bias that histogram. The text says that for a reported range such as 3–5, 'we count each value within a range', so a single study contributes to three separate bars. This inflates the counts for 3, 4, and 5 and makes the mid-range appear more common than it actually is. In addition, the text declares 'For simplicity, we treat them as equivalent' when referring to reviews per submission and reviews per student, although these quantities differ whenever the number of submissions and number of students differ, or when calibration reviews are assigned. The recommendation should be re-derived by reporting one value per study (e.g., the midpoint or a consistent choice of fixed value), and by separating per-subm
  2. [§2.2 and Figure 1] The study selection numbers are internally inconsistent. The text states that after full-text review 'Nine papers were later excluded', yet Figure 1 reports '25 studies excluded' at the eligibility stage and '87 studies included'. Since 112 papers were assessed, 112−25=87, while 112−9=103. The text and figure cannot both be correct, and this ambiguity obscures exactly which 87 studies form the evidence base. The same inconsistency affects the denominator for the 75% prevalence claim in §3.1: Figure 2 shows 53 random, 15 competency, 2 social-network, and 1 bidding, summing to 71; the text says 'For the other 16 papers, there was no explicit mention'. The 75% is therefore 53/71 among studies that specify a strategy, not 53/87 of all included studies. Please correct the flow diagram and state clearly whether prevalence is computed over all 87 or over the 71 with an identifiable strategy.
  3. [§2.2, Limitations section] The search was restricted to the exact keywords 'peer grading' and/or 'peer marking', English-language studies, and seven databases, with no snowballing or gray literature. The Limitations paragraph acknowledges that non-English and gray literature were excluded, but the more specific risk is that adjacent terminology—'peer feedback', 'peer evaluation', 'calibrated peer review'—is common in this field and was not searched. Because the 75% random-assignment prevalence and the 3–5 recommendation are aggregate claims over the included corpus, the narrow search is a load-bearing threat to external validity. A concrete test would be to run supplementary searches with the broader terms and check whether the distribution of strategies and the distribution of review counts change materially; if they do not, the claim can stand. This should be reported in the revision.
minor comments (6)
  1. [§2.2] The PRISMA flow diagram uses '626 studies imported for screening' but the text says 626 publications were returned; the wording is inconsistent. Also, the diagram reports 238 studies irrelevant after abstract screening, while the text says 350 were screened and 112 selected; the arithmetic (350−238=112) should be made explicit.
  2. [§3.1] The sentence 'For the other 16 papers, there was no explicit mention of reviewer assignment strategies' is important but the denominator for the 75% figure is not stated. Please report the denominator explicitly (71 or 87) and consider stating that 16 papers were excluded from the strategy-trend analysis.
  3. [§3.2] The phrase 'the matrices of competency' (near citations [21] and [35]) should be 'the criteria for competency' or 'the indicators of competency'.
  4. [References and text] Reference [41] is listed as 'A. A. V. Ioannis Caragiannis, George A. Krimpas, ...' which appears malformed; the author list should be checked for consistency with the preceding reference [40]. Also, 'Jingjing et. al' in §3.1 should be 'Jingjing et al.' per the citation style.
  5. [§3.3] The statement 'In most cases, these numbers are identical' is presented without a citation or a worked definition. Since the paper later relies on this identity for Figure 3, either add a derivation or qualify the statement explicitly as an assumption.
  6. [Figure 3] The x-axis label reads 'Number of Reviewers per Submission' but the text in §3.3 conflates per-submission and per-student counts. The caption should indicate the unit actually plotted after the re-analysis, or separate the two cases.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the review synthesizes external studies; conclusions are summary statistics, not derived from inputs by construction.

full rationale

This paper is a systematic literature review. Its central claims—the prevalence of random assignment (~75%) and the recommendation of 3–5 reviews per submission—are presented as aggregations of findings from 87 external studies, not as predictions derived from a model or from parameters fitted in this paper. The review does not define reviewer-assignment categories in terms of the outcome, does not fit a parameter and then 'predict' that same parameter, and does not rely on self-citations as load-bearing evidence. The counting choices in Figure 3 (counting each value in a range and treating 'per submission' and 'per student' as equivalent) are methodological limitations that could bias the frequency distribution, but they do not make the conclusion circular: the conclusion is an inductive summary of the literature, not a derivation that assumes the conclusion as an input. The PRISMA flow inconsistency is a reporting error, not circularity. No quoted step exhibits the specific reduction required to establish circularity. Therefore, per the hard rules, I report no significant circularity with score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. It rests on methodological assumptions of the search and synthesis, listed above.

assumptions (3)
  • domain assumption The seven databases and the two keyword families ('peer grading'/'peer marking') capture the relevant literature on reviewer assignment strategies.
    Section 2.2: the entire study set derives from this search; if terminology such as 'peer feedback' or 'peer evaluation' was excluded, prevalence and recommendation claims could be biased.
  • ad hoc to paper The four-strategy taxonomy (random, competency-based, social-network-based, bidding) is an exhaustive and meaningful partition of assignment strategies in the included studies.
    Section 3.1: the authors impose this coding scheme; 16 papers do not fit because they do not mention a strategy, and no inter-coder reliability is reported specifically for this taxonomy.
  • domain assumption Claims about the effect of review count on accuracy generalize across assignment type and context (classroom vs MOOC).
    Section 3.3 aggregates studies with different designs, measures, and contexts to infer a 3-5 review sweet spot; the authors themselves note context-specificity limits generalizability in Section 4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimizing Peer Grading: A Systematic Literature Review of Reviewer Assignment Strategies and Quantity of Reviewers." pith.science (2026). https://pith.science/paper/3LPW3KIW

@misc{pith2026250811678,
  author       = {Pith},
  title        = {Pith review of: Optimizing Peer Grading: A Systematic Literature Review of Reviewer Assignment Strategies and Quantity of Reviewers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3LPW3KIW}},
  note         = {Machine review of arXiv:2508.11678}
}
read the original abstract

Peer assessment has established itself as a critical pedagogical tool in academic settings, offering students timely, high-quality feedback to enhance learning outcomes. However, the efficacy of this approach depends on two factors: (1) the strategic allocation of reviewers and (2) the number of reviews per artifact. This paper presents a systematic literature review of 87 studies (2010--2024) to investigate how reviewer-assignment strategies and the number of reviews per submission impact the accuracy, fairness, and educational value of peer assessment. We identified four common reviewer-assignment strategies: random assignment, competency-based assignment, social-network-based assignment, and bidding. Drawing from both quantitative data and qualitative insights, we explored the trade-offs involved in each approach. Random assignment, while widely used, often results in inconsistent grading and fairness concerns. Competency-based strategies can address these issues. Meanwhile, social and bidding-based methods have the potential to improve fairness and timeliness -- existing empirical evidence is limited. In terms of review count, assigning three reviews per submission emerges as the most common practice. A range of three to five reviews per student or per submission is frequently cited as a recommended spot that balances grading accuracy, student workload, learning outcomes, and engagement.

Figures

Figures reproduced from arXiv: 2508.11678 by the authors.

Figure 1
Figure 1. PRISMA flow diagram of the study selection process. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Distribution of Reviewer Assignment Strategies across literature [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Number of Reviewers per Submission across literature [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 53 canonical work pages

  1. [1]

    Insights about large-scale online peer assessment from an analysis of an astronomy MOOC,

    M. Formanek, M. C. Wenger, S. R. Buxner, C. D. Impey, and T. Sonam, “Insights about large-scale online peer assessment from an analysis of an astronomy MOOC,” Computers & Education, vol. 113, pp. 243–262, 2017

  2. [2]

    Validity of peer grading using calibrated peer review in a guided-inquiry, conceptual physics course,

    E. Price, F. Goldberg, S. Robinson, and M. McKean, “Validity of peer grading using calibrated peer review in a guided-inquiry, conceptual physics course,”Phys. Rev. Phys. Educ. Res., vol. 12, p. 020145, 2016

  3. [3]

    Auto grouping and peer grading system in massive open online course (mooc),

    Y. Chiou and T. K. Shih, “Auto grouping and peer grading system in massive open online course (mooc),” International Journal of Distance Education Technologies (IJDET), vol. 13, no. 3, pp. 25–43, 2015. 10 U. Paul et al

  4. [4]

    Feed-forward assessment, exemplars and peer marking: evidence of efficacy,

    K. Wimshurst and M. Manning, “Feed-forward assessment, exemplars and peer marking: evidence of efficacy,” Assessment & Evaluation in Higher Education, vol. 38, no. 4, pp. 451–465, 2013

  5. [5]

    Improving peer assessment with graph neural networks,

    A. A. Namanloo, J. Thorpe, and A. Salehi-Abari, “Improving peer assessment with graph neural networks,” inProceedings of the 15th International Conference on Ed- ucational Data Mining, A. Mitrovic and N. Bosch, Eds. International Educational Data Mining Society, 2022, pp. 325–332

  6. [6]

    Machine and social intelligent peer-assessment systems for assessing large student populations in massive open online education,

    C. Jimenez Romero, J. Johnson, and R. Castro, “Machine and social intelligent peer-assessment systems for assessing large student populations in massive open online education,” 2013

  7. [7]

    Improving peer grading reliability with graph mining techniques,

    N. Capuano, S. Caballé, and J. Miguel, “Improving peer grading reliability with graph mining techniques,” International Journal of Emerging Technologies in Learning (iJET), vol. 11, no. 07, pp. 24–33, 2016

  8. [8]

    A. A. Russell,The Evolution of Calibrated Peer Review™, ch. 9, pp. 129–143

Show all 53 references
  1. [9]

    Improving es- say peer grading accuracy in massive open online courses using personalized weights from student’s engagement and performance,

    C. García-Martínez, R. Cerezo, M. Bermúdez, and C. Romero, “Improving es- say peer grading accuracy in massive open online courses using personalized weights from student’s engagement and performance,”Journal of Computer As- sisted Learning, vol. 35, no. 2, pp. 201–213, 2019

  2. [10]

    The peerrank method for peer assessment,

    T. Walsh, “The peerrank method for peer assessment,”CoRR, vol. abs/1405.7192, 2014

  3. [11]

    Peer assessment in moocs: Systematic literature review,

    D. Gamage, T. Staubitz, and M. W. and, “Peer assessment in moocs: Systematic literature review,” Distance Education, vol. 42, no. 2, pp. 268–289, 2021

  4. [12]

    Systematic review of approaches to improve peer assessment at scale,

    M. Ravikiran, “Systematic review of approaches to improve peer assessment at scale,” 2020. [Online]. Available: https://arxiv.org/abs/2001.10617

  5. [13]

    An empirical review of anonymity effects in peer assessment, peer feedback, peer review, peer evaluation and peer grading,

    E. Panadero and M. Alqassab, “An empirical review of anonymity effects in peer assessment, peer feedback, peer review, peer evaluation and peer grading,”Assess- ment & Evaluation in Higher Education, vol. 44, no. 8, pp. 1253–1278, 2019

  6. [14]

    Declaración prisma 2020: una guía actualizada para la publicación de revisiones sistemáticas,

    M. J. Page and J. E. McKenzie, “Declaración prisma 2020: una guía actualizada para la publicación de revisiones sistemáticas,”Revista Española de Cardiología, vol. 74, no. 9, pp. 790–799, 2021

  7. [15]

    The measurement of observer agreement for cate- gorical data,

    J. R. Landis and G. G. Koch, “The measurement of observer agreement for cate- gorical data,” Biometrics, vol. 33, no. 1, pp. 159–174, 1977

  8. [16]

    Flipped small group classes and peer marking: incen- tives, student participation and performance in a quasi-experimental approach,

    R. Khatoon and E. Jones, “Flipped small group classes and peer marking: incen- tives, student participation and performance in a quasi-experimental approach,” Assessment & Evaluation in Higher Education, vol. 47, no. 6, pp. 910–927, 2022

  9. [17]

    Simulating massive open on-line courses dy- namics,

    F. Sciarrone and M. Temperini, “Simulating massive open on-line courses dy- namics,” in 2019 18th International Conference on Information Technology Based Higher Education and Training (ITHET), 2019, pp. 1–9

  10. [18]

    Removing bias and incentivizing precision in peer-grading,

    A. Chakraborty, J. Jindal, and S. Nath, “Removing bias and incentivizing precision in peer-grading,” Journal of Artificial Intelligence Research, vol. 79, 2024

  11. [19]

    Grading the graders: Motivating peer graders in a mooc,

    Y. Lu, J. Warren, C. Jermaine, S. Chaudhuri, and S. Rixner, “Grading the graders: Motivating peer graders in a mooc,” ser. WWW ’15. International World Wide Web Conferences Steering Committee, 2015

  12. [20]

    Onthevalidityofpeergradingandacloudteaching assistant system,

    T.VogelsangandL.Ruppertz,“Onthevalidityofpeergradingandacloudteaching assistant system,” ser. LAK ’15. Association for Computing Machinery, 2015, p. 41–50

  13. [21]

    Peer assessment in moocs based on learners’ profiles clustering,

    H. Lynda, B.-D. Farida, B. Tassadit, and L. Samia, “Peer assessment in moocs based on learners’ profiles clustering,” in 2017 8th International Conference on Information Technology (ICIT), 2017, pp. 532–536. Title Suppressed Due to Excessive Length 11

  14. [22]

    Scaling up data science course projects: A case study,

    B. Bhavya, J. Xiao, and C. Zhai, “Scaling up data science course projects: A case study,” in Proceedings of the Eighth ACM Conference on Learning @ Scale, ser. L@S ’21. Association for Computing Machinery, 2021, p. 311–314

  15. [23]

    Mechanical ta: Partially auto- mated high-stakes peer grading,

    J. R. Wright, C. Thornton, and K. Leyton-Brown, “Mechanical ta: Partially auto- mated high-stakes peer grading,” inProceedings of the 46th ACM Technical Sym- posium on Computer Science Education, ser. SIGCSE ’15. Association for Com- puting Machinery, 2015, p. 96–101

  16. [24]

    Improving feedback and discussion in mooc peer assessement using introduced peers,

    D. Gamage, M. E. Whiting, I. Perera, and S. Fernando, “Improving feedback and discussion in mooc peer assessement using introduced peers,” in2018 IEEE In- ternational Conference on Teaching, Assessment, and Learning for Engineering (TALE), 2018, pp. 357–364

  17. [25]

    Collaborative calibrated peer assessment in massive open online courses,

    A. Boudria, Y. Lafifi, and Y. Bordjiba, “Collaborative calibrated peer assessment in massive open online courses,”International Journal of Distance Education Tech- nologies (IJDET), vol. 16, no. 1, pp. 76–102, 2018

  18. [26]

    Non-calibrated peer assessment: An effective assess- ment method for student creative works,

    J. Li, Y. Zhang, and K. Gao, “Non-calibrated peer assessment: An effective assess- ment method for student creative works,” inHCI International 2015 - Posters’ Extended Abstracts, ser. Communications in Computer and Information Science, C. Stephanidis, Ed., 2015, vol. 529, pp. 246–251

  19. [27]

    A blockchain approach to academic assessment,

    S. Alipour, S. Elahimanesh, S. Jahanzad, P. Morassafar, and S. P. Neshaei, “A blockchain approach to academic assessment,” inExtended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, ser. CHI EA ’22. As- sociation for Computing Machinery, 2022

  20. [28]

    Identifying items for moderation in a peer assessment framework,

    S. James, E. Lanham, V. Mak-Hau, L. Pan, T. Wilkin, and G. Wood-Bradley, “Identifying items for moderation in a peer assessment framework,”Knowledge- Based Systems, vol. 162, pp. 211–219, 2018

  21. [29]

    How tool support and peer scoring improved our stu- dents’ attitudes toward peer reviews,

    D. Toll and A. Wingkvist, “How tool support and peer scoring improved our stu- dents’ attitudes toward peer reviews,” inProceedings of the 2017 ACM Conference on Innovation and Technology in Computer Science Education, ser. ITiCSE ’17. Association for Computing Machinery, 2017...

  22. [30]

    Adaptive peer grading and formative assessment,

    G. Albano and N. Capuano, “Adaptive peer grading and formative assessment,” Journal of E-Learning and Knowledge Society, vol. 13, pp. 147–161, 2017

  23. [31]

    Practical methods for semi-automated peer grading in a classroom setting,

    Z. Yuan and D. Downey, “Practical methods for semi-automated peer grading in a classroom setting,” inProceedings of the 28th ACM Conference on User Model- ing, Adaptation and Personalization, ser. UMAP ’20. Association for Computing Machinery, 2020, p. 363–367

  24. [32]

    How learning analytics can help orchestration of formative assessment? data-driven recommen- dations for technology- enhanced learning,

    R. Andriamiseza, F. Silvestre, J.-F. Parmentier, and J. Broisin, “How learning analytics can help orchestration of formative assessment? data-driven recommen- dations for technology- enhanced learning,”IEEE Transactions on Learning Tech- nologies, vol. 16, no. 5, pp. 804–819, 2023

  25. [33]

    Crowdgrader: a tool for crowdsourcing the eval- uation of homework assignments,

    L. de Alfaro and M. Shavlovsky, “Crowdgrader: a tool for crowdsourcing the eval- uation of homework assignments,” inProceedings of the 45th ACM Technical Sym- posium on Computer Science Education, ser. SIGCSE ’14. Association for Com- puting Machinery, 2014, p. 415–420

  26. [34]

    Improving the peer as- sessment experience on mooc platforms,

    T. Staubitz, D. Petrick, M. Bauer, J. Renz, and C. Meinel, “Improving the peer as- sessment experience on mooc platforms,” ser. L@S ’16. Association for Computing Machinery, 2016, p. 389–398

  27. [35]

    Automatedgradingforadvancedtopicscourses,

    E.E.Maicus,“Automatedgradingforadvancedtopicscourses,” Ph.D.dissertation, 2021

  28. [36]

    Learneval peer assessment platform: Iterative devel- opment process and evaluation,

    G. Badea and E. Popescu, “Learneval peer assessment platform: Iterative devel- opment process and evaluation,” IEEE Transactions on Learning Technologies, vol. 15, no. 3, pp. 421–433, 2022. 12 U. Paul et al

  29. [37]

    Trust-aware peer assessment using multi-armed banditalgorithms,

    H. P. Chan, T. Zhao, and I. King, “Trust-aware peer assessment using multi-armed banditalgorithms,” in Proceedings of the 25th International Conference Companion on World Wide Web, ser. WWW ’16 Companion. International World Wide Web Conferences Steering Committee, 2016, p. 899–903

  30. [38]

    Open, collaborative, and ai-augmented peer assessment: Student participation, performance, and perceptions,

    C. Ou, P. Thajchayapong, and D. Joyner, “Open, collaborative, and ai-augmented peer assessment: Student participation, performance, and perceptions,” ser. L@S ’24. Association for Computing Machinery, 2024, p. 496–500

  31. [39]

    Mechanical ta 2: Peer grading with ta and algorithmic support,

    H. Zarkoob and K. Leyton-Brown, “Mechanical ta 2: Peer grading with ta and algorithmic support,” ser. SIGCSE 2024. Association for Computing Machinery, 2024, p. 1470–1476

  32. [40]

    How effective can simple ordinal peer grading be?

    I. Caragiannis, G. A. Krimpas, and A. A. Voudouris, “How effective can simple ordinal peer grading be?”ACM Trans. Econ. Comput., vol. 8, no. 3, 2020

  33. [41]

    Aggregating partial rankings with applications to peer grading in massive online open courses,

    A. A. V. Ioannis Caragiannis, George A. Krimpas, “Aggregating partial rankings with applications to peer grading in massive online open courses,” CoRR, vol. abs/1411.4619, 2014

  34. [42]

    Better peer grading through bayesian inference,

    H. Zarkoob, G. d’Eon, L. Podina, and K. Leyton-Brown, “Better peer grading through bayesian inference,” 2022

  35. [43]

    Peer and self assessment in massive online classes,

    C. Kulkarni, K. P. Wei, H. Le, D. Chia, K. Papadopoulos, J. Cheng, D. Koller, and S. R. Klemmer, “Peer and self assessment in massive online classes,” vol. 20, no. 6, 2013

  36. [44]

    Peer-marking and peer-feedback for coding exercises,

    T. L. Rodgers, “Peer-marking and peer-feedback for coding exercises,”Education for Chemical Engineers, vol. 29, pp. 56–60, 2019

  37. [45]

    Probabilistic graphical models for boosting cardinal and ordinal peer grading in moocs,

    F. Mi and D.-Y. Yeung, “Probabilistic graphical models for boosting cardinal and ordinal peer grading in moocs,” inProceedings of the Twenty-Ninth AAAI Confer- ence on Artificial Intelligence, ser. AAAI’15. AAAI Press, 2015, p. 454–460

  38. [46]

    Rankwithta: A robust and accurate peer grading mechanism for moocs,

    H. Fang, Y. Wang, Q. Jin, and J. Ma, “Rankwithta: A robust and accurate peer grading mechanism for moocs,” in 2017 IEEE 6th International Conference on Teaching, Assessment, and Learning for Engineering (TALE), 2017, pp. 497–502

  39. [47]

    Supporting peer assessment in education with conver- sational agents,

    Y.-C. Lee and W.-T. Fu, “Supporting peer assessment in education with conver- sational agents,” ser. IUI ’19 Companion. Association for Computing Machinery, 2019, p. 7–8

  40. [48]

    Vista, E

    A. Vista, E. Care, and P. Griffin, “A new approach towards marking large-scale complex assessments: Developing a distributed marking system that uses an au- tomatically scaffolding and rubric-targeted interface for guided peer-review,”As- sessing Writing, vol. 24, pp. 1–15, 2015

  41. [49]

    Leveraging social connections to improve peer assessment in moocs,

    H. P. Chan and I. King, “Leveraging social connections to improve peer assessment in moocs,” in Proceedings of the 26th International Conference on World Wide Web Companion, ser. WWW ’17 Companion. International World Wide Web Conferences Steering Committee, 2017, p. 341–349

  42. [50]

    Appropriate number of raters for irt based peer assessment evaluation of programming skills,

    M. Nakayama, M. Uto, M. Temperini, and F. Sciarrone, “Appropriate number of raters for irt based peer assessment evaluation of programming skills,” in2024 21st International Conference on Information Technology Based Higher Education and Training (ITHET), 2024, pp. 1–5

  43. [51]

    Methods for ordinal peer grading,

    K. Raman and T. Joachims, “Methods for ordinal peer grading,” ser. KDD ’14. Association for Computing Machinery, 2014, p. 1037–1046

  44. [52]

    Online peer marking with aggrega- tion functions,

    S. James, L. Pan, T. Wilkin, and L. Yin, “Online peer marking with aggrega- tion functions,” in2017 IEEE International Conference on Fuzzy Systems (FUZZ- IEEE), 2017, pp. 1–6

  45. [53]

    A dynamic review allocation approach for peer assess- ment in technology enhanced learning,

    G. Badea and E. Popescu, “A dynamic review allocation approach for peer assess- ment in technology enhanced learning,”Education and Information Technologies, vol. 27, no. 9, p. 13131–13162, Nov. 2022

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.