A fully automated AI-for-AI research system produced 166 papers across 67 topics; human reviews of 140 papers show occasional review-worthy work but mostly low scores and recurring integrity and scope failures.
Has the machine learning review process become more arbitrary as the field has grown? the neurips 2021 consistency experiment
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
support 1representative citing papers
PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.
Intern-Atlas constructs a methodological evolution graph with 9.4 million edges from 1.03 million AI papers to capture how methods emerge, adapt, and transition, enabling better idea evaluation and generation for AI-driven research.
A 120-respondent survey maps ESE community perceptions of review load, quality problems, LLM use in reviewing, and proposed system improvements.
Controlled prompt interventions reveal strong affiliation bias in LLM peer reviews favoring top-ranked institutions, plus effects from seniority and publication history.
A game-theoretic model demonstrates a Nash equilibrium in which authors voluntarily accept random pre-review rejection to reduce reviewer burden and raise evaluation quality.
citing papers explorer
-
FARS: A Fully Automated Research System Deployed at Scale
A fully automated AI-for-AI research system produced 166 papers across 67 topics; human reviews of 140 papers show occasional review-worthy work but mostly low scores and recurring integrity and scope failures.
-
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.
-
Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists
Intern-Atlas constructs a methodological evolution graph with 9.4 million edges from 1.03 million AI papers to capture how methods emerge, adapt, and transition, enabling better idea evaluation and generation for AI-driven research.
-
The State of Peer Review in Empirical Software Engineering: A Community Survey on Review Load, Quality, and GenAI Use
A 120-respondent survey maps ESE community perceptions of review load, quality problems, LLM use in reviewing, and proposed system improvements.
-
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
Controlled prompt interventions reveal strong affiliation bias in LLM peer reviews favoring top-ranked institutions, plus effects from seniority and publication history.
-
Can We Volunteer Out of the Peer Review Crisis?
A game-theoretic model demonstrates a Nash equilibrium in which authors voluntarily accept random pre-review rejection to reduce reviewer burden and raise evaluation quality.