Pith. sign in

REVIEW 7 cited by

Has the Machine Learning Review Process Become More Arbitrary as the Field Has Grown? The NeurIPS 2021 Consistency Experiment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03262 v1 pith:OEWVM22M submitted 2023-06-05 cs.LG cs.DL

classification cs.LGcs.DL
keywords processexperimentneuripsreviewcommitteesconferenceconsistencyresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the NeurIPS 2021 consistency experiment, a larger-scale variant of the 2014 NeurIPS experiment in which 10% of conference submissions were reviewed by two independent committees to quantify the randomness in the review process. We observe that the two committees disagree on their accept/reject recommendations for 23% of the papers and that, consistent with the results from 2014, approximately half of the list of accepted papers would change if the review process were randomly rerun. Our analysis suggests that making the conference more selective would increase the arbitrariness of the process. Taken together with previous research, our results highlight the inherent difficulty of objectively measuring the quality of research, and suggest that authors should not be excessively discouraged by rejected work.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?

    cs.DL 2026-07 conditional novelty 8.0 of 10

    At ICLR, equally scored borderline papers from outside top-25 institutions are accepted less often, a gap concentrated in preprint-identifiable submissions; outcome tests find no evidence of a higher bar.

  2. FARS: A Fully Automated Research System Deployed at Scale

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    FARS deployed at scale produced 166 AI/ML papers across 67 topics that received 282 structured human reviews indicating some review-worthy outputs alongside recurring failure modes.

  3. ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review

    cs.HC 2026-07 conditional novelty 6.0 of 10

    ReVoicer is a prototype that turns spoken, in-the-moment reactions to a paper into cleaned, tagged annotations and a draft review aligned with the reviewer's own style.

  4. Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review

    cs.DL 2026-04 conditional novelty 6.0 of 10

    An audit of 50,289 ICLR papers shows acceptance odds vary up to 8x across topics at equal reviewer scores, indicating scores are not comparable across research areas.

  5. Publishing Without Journals: An Open, Forkable Archive with Attributed Review

    cs.DL 2026-07 conditional novelty 5.0 of 10

    Unbundle dissemination from certification: open deposit, continuous attributed votable commentary with author replies, and forkable papers that record idea lineage automatically.

  6. How Far Are AI Scientists from Changing the World?

    cs.AI 2025-07 conditional novelty 4.0 of 10

    This survey proposes a four-level capability framework for AI Scientist systems and, using an AI reviewer, finds that current systems produce papers rated well below normal scientific standards.

  7. Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem

    cs.AI 2025-06 conditional novelty 4.0 of 10

    The paper argues that AI-assisted peer review is an urgent priority and that its success depends on collecting richer, structured peer review process data.

Pith tools