Pith. sign in

REVIEW 7 cited by

AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10365 v1 pith:GEHQ4RLQ submitted 2024-08-19 cs.AI

AI-Driven Review Systems: Evaluating LLMs in Scalable and Bias-Aware Academic Reviews

classification cs.AI
keywords reviewshumanreviewpreferencesqualityreviewingevaluateevaluating
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena of human preferences by pairwise comparisons. Gathering human preference may be time-consuming; therefore, we also use an LLM to automatically evaluate reviews to increase sample efficiency while reducing bias. In addition to evaluating human and LLM preferences among LLM reviews, we fine-tune an LLM to predict human preferences, predicting which reviews humans will prefer in a head-to-head battle between LLMs. We artificially introduce errors into papers and analyze the LLM's responses to identify limitations, use adaptive review questions, meta prompting, role-playing, integrate visual and textual analysis, use venue-specific reviewing materials, and predict human preferences, improving upon the limitations of the traditional review processes. We make the reviews of publicly available arXiv and open-access Nature journal papers available online, along with a free service which helps authors review and revise their research papers and improve their quality. This work develops proof-of-concept LLM reviewing systems that quickly deliver consistent, high-quality reviews and evaluate their quality. We mitigate the risks of misuse, inflated review scores, overconfident ratings, and skewed score distributions by augmenting the LLM with multiple documents, including the review form, reviewer guide, code of ethics and conduct, area chair guidelines, and previous year statistics, by finding which errors and shortcomings of the paper may be detected by automated reviews, and evaluating pairwise reviewer preferences. This work identifies and addresses the limitations of using LLMs as reviewers and evaluators and enhances the quality of the reviewing process.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

    cs.CL 2026-05 conditional novelty 6.0

    Official conference reviewer guidelines improve LLM review scores' agreement with human ratings, while LLM-distilled 'reviewer-imitating' guidelines and rubric-style scoring do worse.

  2. ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

    cs.DL 2026-05 unverdicted novelty 6.0

    ARA extracts workflow graphs from papers and scores reproducibility, reaching 61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldStandardDB.

  3. ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

    cs.DL 2026-05 unverdicted novelty 6.0

    ARA uses LLMs to build workflow graphs linking sources, methods, and outputs in papers, then scores reproducibility, reaching ~61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldSt...

  4. BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

    cs.CR 2025-10 conditional novelty 6.0

    An LLM agent generating fabricated papers without experiments gets acceptance-level scores from LLM reviewers up to 82% of the time, and simple integrity-checking mitigations barely beat random.

  5. LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

    cs.CL 2026-06 unverdicted novelty 4.0

    A survey synthesizing LLM methods for peer review critique generation and score prediction, including taxonomies, benchmark limitations, domain biases, and robustness risks such as prompt injection.

  6. When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

    cs.CY 2025-09 conditional novelty 4.0

    GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.

  7. Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

    cs.DL 2025-07 unverdicted novelty 4.0

    The paper proposes a four-role framework for LLMs in scientific innovation and reviews methods, benchmarks, and limitations across Assistant, Collaborator, Scientist, and Evaluator roles.