Pith. sign in

REVIEW 10 cited by

Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.09737 v1 pith:5E6SNJ26 submitted 2025-04-13 cs.AI cs.CLcs.HCcs.LG

Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025

classification cs.AI cs.CLcs.HCcs.LG
keywords feedbackreviewreviewersreviewsagentqualitywereautomated
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Peer review at AI conferences is stressed by rapidly rising submission volumes, leading to deteriorating review quality and increased author dissatisfaction. To address these issues, we developed Review Feedback Agent, a system leveraging multiple large language models (LLMs) to improve review clarity and actionability by providing automated feedback on vague comments, content misunderstandings, and unprofessional remarks to reviewers. Implemented at ICLR 2025 as a large randomized control study, our system provided optional feedback to more than 20,000 randomly selected reviews. To ensure high-quality feedback for reviewers at this scale, we also developed a suite of automated reliability tests powered by LLMs that acted as guardrails to ensure feedback quality, with feedback only being sent to reviewers if it passed all the tests. The results show that 27% of reviewers who received feedback updated their reviews, and over 12,000 feedback suggestions from the agent were incorporated by those reviewers. This suggests that many reviewers found the AI-generated feedback sufficiently helpful to merit updating their reviews. Incorporating AI feedback led to significantly longer reviews (an average increase of 80 words among those who updated after receiving feedback) and more informative reviews, as evaluated by blinded researchers. Moreover, reviewers who were selected to receive AI feedback were also more engaged during paper rebuttals, as seen in longer author-reviewer discussions. This work demonstrates that carefully designed LLM-generated review feedback can enhance peer review quality by making reviews more specific and actionable while increasing engagement between reviewers and authors. The Review Feedback Agent is publicly available at https://github.com/zou-group/review_feedback_agent.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?

    cs.DL 2026-07 conditional novelty 6.0

    ChatGPT-5.4's averaged scores rank journal articles about as reliably as individual expert reviewers, but full-text PDF input does not improve score accuracy over title/abstract input.

  2. SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

    cs.LG 2026-05 conditional novelty 6.0

    SoundnessBench shows frontier LLMs exhibit pervasive optimism bias when rating the soundness of ML research proposals, frequently calling low-soundness ideas sound under standard prompts.

  3. ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

    cs.DL 2026-05 unverdicted novelty 6.0

    ARA extracts workflow graphs from papers and scores reproducibility, reaching 61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldStandardDB.

  4. ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

    cs.DL 2026-05 unverdicted novelty 6.0

    ARA uses LLMs to build workflow graphs linking sources, methods, and outputs in papers, then scores reproducibility, reaching ~61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldSt...

  5. AI Can Learn Scientific Taste

    cs.CL 2026-03 conditional novelty 6.0

    Reinforcement learning on citation-preference pairs teaches a model to predict which papers will be cited more and to propose ideas that LLM judges rate as likely to be cited more—but "taste" here means citation impact.

  6. AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights

    cs.CY 2025-08 conditional novelty 6.0

    LLMs that screen resumes systematically prefer their own generated summaries over human-written ones, with simulated shortlisting advantages of 23 to 60 percent for same-model users.

  7. Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy

    cs.SE 2026-06 unverdicted novelty 5.0

    Agon is a new autonomous research system using prompt economy loops across 444 iterations to demonstrate scalable omnidisciplinary research and a taxonomy separating machine-fixable failures from those needing human judgment.

  8. PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf

    cs.CL 2026-06 unverdicted novelty 4.0

    A multi-agent writing tutor for Overleaf that uses 12 agents and an expert skill library to generate inline comments, with a 14-user study reporting 90.6% actionable and 67.5% valid comments that outperform a GPT-5.2 ...

  9. When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review

    cs.CY 2025-09 conditional novelty 4.0

    GPT-5-mini gives weaker papers systematically higher scores than human reviewers, and hidden field-specific prompts in PDFs can force it to assign perfect scores or suppress weaknesses.

  10. A Backward-Compatible Protocol Upgrade for HotNets

    cs.NI 2026-06 unverdicted novelty 3.0

    HotNets 2026 broadens its scope, adopts distinct review criteria for technical and perspective papers, and experiments with collaborative review and presentation formats.