Automatic reviewers fail to detect faulty reasoning in research papers: A new counterfactual evaluation framework.arXiv preprint arXiv:2508.21422, 2025

Nils Dycke, Iryna Gurevych · 2025 · arXiv 2508.21422

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

cs.CL · 2026-05-26 · unverdicted · novelty 6.0

PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.

citing papers explorer

Showing 1 of 1 citing paper.

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers cs.CL · 2026-05-26 · unverdicted · none · ref 29
PRISM benchmark finds LLMs match or exceed humans on isolated review dimensions like novelty verification but none achieve the balanced performance of human reviewers across depth, flaw prioritization, and constructiveness.

Automatic reviewers fail to detect faulty reasoning in research papers: A new counterfactual evaluation framework.arXiv preprint arXiv:2508.21422, 2025

fields

years

verdicts

representative citing papers

citing papers explorer