Pith. sign in

REVIEW 4 major objections 6 minor 100 references

Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track

T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This position paper argues that machine learning conferences should establish a dedicated "Refutations and Critiques" track so flawed or misleading published work can be officially challenged, corrected, and linked to the original papers.

desk verdict A genuine, well-argued proposal for a refutations track whose main weakness is an undefended assumption that the track's reviews will escape the very constraints used to dismiss review reform. read the letter →

arxiv 2506.19882 v3 pith:TOON3FIP submitted 2025-06-24 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords refutationsandcritiquestrackscientificself-correctionpeerreviewreformpost-publicationcritiquemachinelearningconferencesresearchintegrityreproducibility
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that machine learning conferences, despite their prestige, have no official mechanism for correcting research that peer review wrongly admitted or wrongly highlighted. It proposes a dedicated "Refutations and Critiques" (R&C) track at major ML conferences, where researchers can publish peer-reviewed responses that identify, analyze, and correct misleading, incorrect, or fraudulent claims in impactful prior work. The authors contend that reforming peer review itself is impractical—reviewers are uncompensated, time is short, and qualified reviewers are scarce—so a post-hoc correction track is the more realistic route to a self-correcting field. A sympathetic reader would care because accepted errors currently propagate through citations, while informal critiques on blogs and social media lack verification and permanence.

What carries the argument

The central object is the proposed "Refutations and Critiques" (R&C) track itself—a dedicated submission and review stream inside existing ML conferences. It carries the argument by supplying what the paper says is missing: a high-profile, peer-reviewed destination for critical scholarship. Its design features do the work: bespoke review principles (correctness and rigor; substance; constructiveness; significance), participation by authors of the critiqued work in review, and explicit linking of accepted critiques to the original publication. The paper also leans on the precedent of a dedicated datasets-and-benchmarks review track at a major conference, which carved out specialized reviewing for a previously undervalued contribution type.

What would settle it

A pilot R&C track that attracts mostly trivial or adversarial submissions and produces few citations or corrigenda, or a systematic study showing that paying reviewers and extending review timelines substantially reduces acceptance of flawed papers, would undercut the paper's premise that a separate correction mechanism is the practical route.

Watch

Extended reading notes

Core claim

The paper's central claim is that the health of ML research depends on making critique an official, prestigious category of scholarly contribution. It proposes that major ML conferences add an R&C track whose submissions are judged by four bespoke criteria—correct/rigorous/meticulous, substantive, constructive, and significant—and whose review process gives authors of the criticized paper a role in ensuring fairness. Accepted critiques would be linked directly to the criticized publication, functioning like errata rather than independent papers. The discovery is not a new technical result but a governance proposal: the same peer-review machinery that conferences already operate can, with a dedicated track and tailored standards, convert today's scattered and unvetted refutations into an archival, trustworthy correction stream.

Load-bearing premise

The argument depends on the claim that improving peer review itself—reviewer incentives, review time, or the supply of qualified reviewers—is not a practical way to stop flawed papers, so the field needs a post-hoc correction track instead.

Editorial extensions

If this is right

  • Researchers who find errors in published work would have an archival, peer-reviewed outlet instead of Twitter threads or blog posts.
  • Critiques accepted to the track would be linked to the criticized papers, so readers encountering the original work would also see the correction.
  • Authors of criticized papers would be drawn into the review process, giving them a formal right of response and reducing unfair or frivolous attacks.
  • A successful track would create an incentive for higher-quality original research, since errors would have a visible, official consequence.
  • If one major conference pioneers the track, other leading ML venues may adopt similar mechanisms, standardizing correction across the field.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the same track could eventually host negative results and replication studies, not just refutations of published claims, turning it into a general home for corrective scholarship.
  • The proposal would be testable early: a pilot track that logs whether accepted critiques are cited and whether original authors issue corrigenda would show whether official status actually changes correction behavior.
  • If the track succeeds, evaluation culture may need to shift toward crediting critique work in hiring and promotion, since prestige currently tracks benchmark papers.
  • One unresolved tension the paper leaves implicit: the track's review load falls on the same volunteer reviewers whose incentives the paper says cannot be fixed, so implementation would need to budget for that.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This position paper argues that machine learning conferences (specifically NeurIPS, ICML, and ICLR) should establish a dedicated "Refutations and Critiques" (R&C) track: a peer-reviewed, high-visibility venue for work that identifies, analyzes, and corrects misleading or incorrect claims in prior ML publications. Section 2 builds the case that conferences accept and sometimes highlight flawed research (2.1), lack official rectification mechanisms such as errata or retraction (2.2), cannot practically reform peer review because of reviewer incentives, time, and reviewer supply (2.3), and thereby push disputes into informal channels that harm the field (2.4). Section 3 details the proposal: track design (3.1), naming rationale (3.2), the case for a separate track including the NeurIPS Datasets and Benchmarks Track analogy (3.3), host-venue choice (3.4), and review principles (3.5). Section 4 addresses objections (existing mechanisms, frivolous or adversarial submissions, flawed refutations); Section 5 positions the proposal relative to reproducibility efforts; Section 6 points to an illustrative example submission, a companion paper critiquing an ICLR 2025 Oral; Section 7 concludes with a call to act.

Significance. The paper targets a genuine gap: post-publication critique in ML currently lives in blogs, social media, and OpenReview threads, with no official, visible, peer-reviewed channel. If adopted, an R&C track would give the field a structured correction mechanism, and the paper's concrete design elements are a useful starting point for that discussion. Credit is due for: explicit engagement with counterarguments in Section 4; transparent acknowledgment of the anecdotal evidence base (Introduction and Section 2.4); actionable review principles (Section 3.5); a plausible institutional precedent (the NeurIPS Datasets and Benchmarks Track, Section 3.3); and a falsifiable pilot idea (Section 3.1). The main weaknesses are that the proposal does not show how the track's own reviews escape the fallibility used to dismiss review reform, the 'reform is impractical' premise is asserted rather than demonstrated, and the proposed pilot lacks success criteria. These issues are addressable within the manuscript's scope, so the central claim is defensible.

major comments (4)
  1. [§3.1 vs §2.3] The stress-test concern that the R&C track inherits the very review-process problems it is meant to correct lands on a specific load-bearing tension. Section 2.3 argues that peer review cannot be improved because reviewers are uncompensated volunteers, review timelines are too short for rigorous scrutiny, and qualified reviewers are scarce; Section 3.1 then promises 'a high-profile, reputable, and rigorously peer-reviewed platform' while stating that the track would be 'naturally integrated with existing conference structures, including timelines and review processes.' Sections 3.3 and 4 add procedural elements (bespoke standards, non-double-blind review, author-of-critiqued-paper participation, multiple review rounds), but none of these addresses reviewer time, reviewer incentives, or reviewer supply, and Section 3.5's principles are evaluation criteria rather than structural fixes. The one structural point, in Section 2.3, is that critics themselves have a longer preparation window; that says nothing about the reviewers of the critique. Section 4's concession that published refutations can themselves be flawed (the Zhang et al., 2025 example) makes the gap concrete. The revision should either specify the mechanisms by which R&C reviews achieve reliability despite the Section 2.3 constraints (e.g., small per-cycle volume, targeted recruitment of expert reviewers, a defined author-response and multi-round adjudication protocol), or explicitly reposition the track's contribution as visibility and process rather than added review reliability, and temper the 'rigorously peer-reviewed' claim accordingly.
  2. [§2.3] The argument's pivot—that reforming peer review is impractical, so a post-hoc track is the practical alternative—rests on assertions that are unsupported or in tension with the proposal itself. The 'only two incentives' claim, the reviewer-shortage claim, and the claim that 'we are unaware of more reliable alternatives' are supported respectively by a conference blog post, a Reddit thread, and author assertion. Meanwhile, the NeurIPS 2014 and 2021 consistency experiments cited in Appendix A (Cortes & Lawrence, 2021; Beygelzimer et al., 2021) quantify review noise and are the most direct available evidence for structural fallibility, yet they are not used in Section 2.3. The paper also does not confront the two-edged nature of its own proposal: an R&C track with bespoke standards, non-double-blind review, and author participation is itself a reform of the conference review process, and its success would weaken the claim that review design cannot be improved. Both implications should be addressed explicitly.
  3. [§3.1, §7] The effectiveness claim is not falsifiable as stated. Section 3.1 mentions a pilot 'to evaluate whether the proposed track is effective' but no success criteria are defined anywhere in the paper, while the abstract and Section 7 promise that the track would 'foster a dynamic self-correcting research ecosystem' and 'incentivize more honest, higher quality work.' The Datasets and Benchmarks analogy in Section 3.3 is persuasive precisely because that track had measurable uptake; the R&C proposal should specify analogous pilot metrics, for instance: submission volume and acceptance rate over the pilot cycles; the fraction of accepted R&C papers that lead to documented changes in the scientific record (author-issued corrections, retractions, changed citation or usage practices); structured feedback from authors of critiqued papers; and a protocol for detecting errors in the R&C papers themselves, given the Section 4 concession that refutations can be flawed.
  4. [§2.1, §2.4, §6] The motivating evidence is dominated by the authors' own engagements as critics, which limits its independent force. Several Section 2.1 examples are critiques authored by members of this paper's author list (Schaeffer et al., 2022; 2023a; 2023b; Gerstgrasser et al., 2024; Kazdan et al., 2025; Dey & Donoho, 2024); Section 2.4 asserts that an ICLR 2025 submission 'intentionally suppressed contradictory scientific evidence' (Schaeffer, 2024), an intent claim that a public comment cannot establish and one that sits uneasily with the same section's caveat that 'there are likely two sides to each anecdote'; and Section 6's illustrative example is the authors' companion paper. None of this invalidates the proposal, but the motivating examples do not provide independent adjudication of the disputes they describe. The revision should (a) qualify or replace the 'intentionally suppressed' assertion, (b) add at least one case in which the original authors or the community has publicly acknowledged the error, and (c) state the illustrative example's central claim within the paper so that readers can assess it against the Section 3.5 criteria.
minor comments (6)
  1. [§2.1] The Doc2Vec example states that the paper 'has been cited nearly 14.000 times'; the decimal point is a typo for a thousands separator, and since the non-reproducibility evidence is a StackExchange answer, this example is better framed as an illustration of how difficult disputes are to track than as a documented case of non-reproducibility.
  2. [§2.4] The phrase 'can become popularity contents' appears to be a typo for 'popularity contests.'
  3. [§5] The Journal of Comments and Replications in Economics is discussed without a citation; please add a reference and, ideally, some evidence about its operation and uptake to support the precedent argument.
  4. [§3.5] The criterion that submissions 'solely challenge attribution errors' be rejected appears to conflict with the D-Adaptation example in Section 2.1, where the dispute (Orabona, 2023) is primarily about novelty and prior attribution; please clarify whether prior-art and attribution disputes are within the track's scope.
  5. [§6] Section 6 defers entirely to the companion paper (Schaeffer et al., 2025b) for the illustrative example; since the abstract promises a concrete demonstration, the paper should summarize the example's key claim and sketch how it would be evaluated under the Section 3.5 rubric and the author-response mechanism of Section 4.
  6. [References] The reference 'Pineau, Joelle and others' is malformed and should be reformatted, several entries lack access dates (e.g., Reddit, 2024), and footnote 3's MLRC 2025 URL would be better placed in the reference list.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the proposal is a normative argument whose conclusion does not reduce to its inputs or to self-citations.

full rationale

The paper argues normatively that ML conferences should establish a Refutations and Critiques track, with premises that peer review is fallible, that no official correction mechanism exists, that reforming peer review is impractical, and that a dedicated track would provide a practical post-hoc venue. None of these premises defines the conclusion in terms of itself, and the paper makes no empirical prediction that could be fitted to its own examples. Section 3.1 explicitly proposes a pilot to test effectiveness, so the paper does not even assert that the track will certainly work. Section 4 candidly concedes that refutations themselves can be flawed, citing the Zhang et al. (2025) case, which is a limitation rather than a circular step. The disclaimer in Section 1 that some observations rely on community anecdotes is also explicit and does not disguise those anecdotes as derived results. The many self-citations (e.g., Schaeffer et al. 2022/2023, Gerstgrasser et al. 2024, Kazdan et al. 2025) and the illustrative companion paper (Schaeffer et al. 2025b) are publicly checkable and are used as motivating examples, not as the sole or load-bearing justification; independent examples such as Carlini et al. (2022) and Orabona (2023) carry the argument that flawed work is accepted. Therefore the central claim remains independent of its inputs and no circularity is present.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The proposal relies on several domain assumptions about the ML research ecosystem: peer review is fallible and hard to reform, a dedicated track will attract quality submissions, and original authors will cooperate. These are stated as the paper's operating premises and are not empirically verified. No free parameters or invented entities appear.

assumptions (5)
  • domain assumption Reforming ML conference peer review is impractical due to reviewer incentives, time constraints, and insufficient qualified reviewers.
    Stated in Section 2.3; motivates the choice of a post-hoc track over review overhaul.
  • domain assumption Kuhn's account of scientific progress via iterative correction applies to ML.
    Invoked in Section 1 to justify the need for a correction mechanism.
  • domain assumption A prestigious, peer-reviewed R&C track will attract substantive critiques and improve the scientific record more than existing venues.
    Central to the proposal; asserted in Section 3 with anecdotal support, not pilot data.
  • domain assumption Reviewers can fairly evaluate critiques and distinguish substantive challenges from frivolous or ad hominem attacks.
    Assumed in Sections 3.5 and 4; the paper raises the risk but provides no mechanism to guarantee it.
  • domain assumption Authors of the criticized papers will participate constructively, e.g., by responding to critiques.
    Assumed in Sections 1 and 4; the paper invites this but acknowledges it may not occur.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track." pith.science (2026). https://pith.science/paper/TOON3FIP

@misc{pith2026250619882,
  author       = {Pith},
  title        = {Pith review of: Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TOON3FIP}},
  note         = {Machine review of arXiv:2506.19882}
}
read the original abstract

Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of publications, but have also led to misleading, incorrect, flawed or perhaps even fraudulent studies being accepted and sometimes highlighted at ML conferences due to the fallibility of peer review. While such mistakes are understandable, ML conferences do not offer robust processes to help the field systematically correct when such errors are made. This position paper argues that ML conferences should establish a dedicated "Refutations and Critiques" (R&C) Track. This R&C Track would provide a high-profile, reputable platform to support vital research that critically challenges prior research, thereby fostering a dynamic self-correcting research ecosystem. We discuss key considerations including track design, review principles, potential pitfalls, and provide an illustrative example submission concerning a recent ICLR 2025 Oral. We conclude that ML conferences should create official, reputable mechanisms to help ML research self-correct.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

100 extracted references · 44 canonical work pages

  1. [1]

    URL https://en.wikipedia.org/wiki/Path_integration

    Path integration, n.d. URL https://en.wikipedia.org/wiki/Path_integration. Wikipedia article

  2. [2]

    Deep learning with differential privacy

    Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pp.\ 308--318, 2016

  3. [3]

    Science forum: Consensus-based guidance for conducting and reporting multi-analyst studies

    Balazs Aczel, Barnabas Szaszi, Gustav Nilsonne, Olmo R van den Akker, Casper J Albers, Marcel ALM van Assen, Jojanneke A Bastiaansen, Daniel Benjamin, Udo Boehm, Rotem Botvinik-Nezer, Laura F Bringmann, Niko A Busch, Emmanuel Caruyer, Andrea M Cataldo, Nelson Cowan, Andrew Delios, Noah NN van Dongen, Chris Donkin, Johnny B van Doorn, Anna Dreber, Gilles D...

  4. [4]

    Deep reinforcement learning at the edge of the statistical precipice

    Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare. Deep reinforcement learning at the edge of the statistical precipice. Advances in Neural Information Processing Systems, 34, 2021

  5. [5]

    Baraniuk

    Sina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun, Hossein Babaei, Daniel LeJeune, Ali Siahkoohi, and Richard G. Baraniuk. Self-consuming generative models go MAD . arXiv preprint arXiv:2307.01850, 2023. URL https://arxiv.org/abs/2307.01850

  6. [6]

    Scalable second order optimization for deep learning

    Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. Scalable second order optimization for deep learning. arXiv preprint arXiv:2002.09018, 2020

  7. [7]

    Distributed shampoo was rejected at iclr 2021, later won algoperf, 2024

    Arohan. Distributed shampoo was rejected at iclr 2021, later won algoperf, 2024. URL https://x.com/_arohan_/status/1889357350664020467. Tweet

  8. [8]

    Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples

    Anish Athalye, Nicholas Carlini, and David Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In Jennifer Dy and Andreas Krause (eds.), Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pp.\ 274--283. PMLR, 10--15 Jul 2018. ...

Show all 100 references
  1. [9]

    Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan H...

  2. [10]

    Vector-based navigation using grid-like representations in artificial agents

    Andrea Banino, Caswell Barry, Benigno Uria, Charles Blundell, Timothy Lillicrap, Piotr Mirowski, Alexander Pritzel, Martin J Chadwick, Thomas Degris, Joseph Modayil, et al. Vector-based navigation using grid-like representations in artificial agents. Nature, 557 0 (7705): 0 42...

  3. [11]

    Theoretical guarantees on the best-of-n alignment policy

    Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D'Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh. Theoretical guarantees on the best-of-n alignment policy. arXiv preprint arXiv:2401.01879, 2025. URL https://arxiv.org/abs/2401.01879

  4. [12]

    The NeurIPS 2021 Consistency Experiment

    Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan. The NeurIPS 2021 Consistency Experiment . NeurIPS Blog, December 2021. URL https://blog.neurips.cc/2021/12/08/the-neurips-2021-consistency-experiment/. Accessed: 2025-05-18

  5. [13]

    Instahide disappointingly wins bell labs prize, 2nd place

    Nicholas Carlini. Instahide disappointingly wins bell labs prize, 2nd place. Blog post, dec 2020. URL https://nicholas.carlini.com/writing/2020/instahide-disappointingly-wins-bell-labs-prize.html. Accessed on May 21, 2025

  6. [14]

    efficient defenses against adversarial attacks

    Nicholas Carlini and David Wagner. MagNet and "efficient defenses against adversarial attacks" are not robust to adversarial examples. arXiv preprint arXiv:1711.08478, 2017. URL https://arxiv.org/abs/1711.08478

  7. [15]

    Is private learning possible with instance encoding? arXiv preprint arXiv:2011.05315, 2021

    Nicholas Carlini, Samuel Deng, Sanjam Garg, Somesh Jha, Saeed Mahloujifar, Mohammad Mahmoody, Shuang Song, Abhradeep Thakurta, and Florian Tramer. Is private learning possible with instance encoding? arXiv preprint arXiv:2011.05315, 2021. URL https://arxiv.org/abs/2011.05315

  8. [16]

    privacy for free: How does dataset condensation help privacy

    Nicholas Carlini, Vitaly Feldman, and Milad Nasr. No free lunch in "privacy for free: How does dataset condensation help privacy". arXiv preprint arXiv:2209.14987, 2022. URL https://arxiv.org/abs/2209.14987

  9. [17]

    Incorrect baseline evaluations call into question recent llm-rl claims

    Nikhil Chandak, Shashwat Goel, and Ameya Prabhu. Incorrect baseline evaluations call into question recent llm-rl claims. https://safe-lip-9a8.notion.site/Incorrect-Baseline-Evaluations-Call-into-Question-Recent-LLM-RL-Claims-2012f1fbf0ee8094ab8ded1953c15a37?pvs=4, 2025. Notion Blog

  10. [18]

    A closer look at few-shot classification, 2020

    Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. A closer look at few-shot classification, 2020. URL https://arxiv.org/abs/1904.04232

  11. [19]

    Lawrence

    Corinna Cortes and Neil D. Lawrence. Inconsistency in conference peer review: Revisiting the 2014 neurips experiment. arXiv preprint arXiv:2109.09774, 2021. URL https://arxiv.org/abs/2109.09774

  12. [20]

    Reward model ensembles help mitigate overoptimization

    Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger. Reward model ensembles help mitigate overoptimization. In The Twelfth International Conference on Learning Representations, 2024

  13. [21]

    Ai research summit neurips 2025 receives record-breaking 27,000 paper submissions

    CTOL Digital Solutions . Ai research summit neurips 2025 receives record-breaking 27,000 paper submissions. https://www.ctol.digital/news/ai-research-summit-neurips-2025-receives-record-breaking-27000-paper-submissions/, 2025. Accessed: 2025-06-07

  14. [22]

    Learning-rate-free learning by D -adaptation

    Aaron Defazio and Konstantin Mishchenko. Learning-rate-free learning by D -adaptation. In International Conference on Machine Learning, pp.\ 7449--7479. PMLR, 2023

  15. [23]

    Universality of the ^ 2/6 pathway in avoiding model collapse

    Apratim Dey and David Donoho. Universality of the ^ 2/6 pathway in avoiding model collapse. arXiv preprint arXiv:2410.22812, 2024

  16. [24]

    Distill . Distill. https://distill.pub, 2016. Online journal for machine learning research

  17. [25]

    Model collapse demystified: The case of regression, 2024

    Elvis Dohmatob, Yunzhen Feng, and Julia Kempe. Model collapse demystified: The case of regression, 2024. URL https://arxiv.org/abs/2402.07712

  18. [26]

    Privacy for free: How does dataset condensation help privacy? In International Conference on Machine Learning, pp.\ 5378--5396

    Tian Dong, Bo Zhao, and Lingjuan Lyu. Privacy for free: How does dataset condensation help privacy? In International Conference on Machine Learning, pp.\ 5378--5396. PMLR, 2022

  19. [27]

    Investigating the replicability of preclinical cancer biology

    Timothy M Errington, Maya Mathur, Courtney K Soderberg, Alexandria Denis, Nicole Perfito, Elizabeth Iorns, and Brian A Nosek. Investigating the replicability of preclinical cancer biology. eLife, 10: 0 e71601, dec 2021. ISSN 2050-084X. doi:10.7554/eLife.71601. URL https://doi....

  20. [28]

    Privacy for free?, 2022

    Vitaly Feldman. Privacy for free?, 2022. URL https://x.com/vitalyFM/status/1549599469695512576. Tweet

  21. [30]

    Scaling laws for reward model overoptimization

    Leo Gao, John Schulman, and Jacob Hilton. Scaling laws for reward model overoptimization. In Proceedings of the 40th International Conference on Machine Learning, pp.\ 10835--10866, 2023

  22. [31]

    Roberts, Diyi Yang, David L

    Matthias Gerstgrasser, Rylan Schaeffer, Apratim Dey, Rafael Rafailov, Henry Sleight, John Hughes, Tomasz Korbak, Rajashree Agrawal, Dhruv Pai, Andrey Gromov, Daniel A. Roberts, Diyi Yang, David L. Donoho, and Sanmi Koyejo. Is model collapse inevitable? breaking the curse of re...

  23. [32]

    Intricacies of feature geometry in large language models

    Satvik Golechha, Lucius Bushnaq, Euan Ong, Neeraj Kayal, and Nandi Schoots. Intricacies of feature geometry in large language models. In The Fourth Blogpost Track at ICLR 2025, 2025. URL https://openreview.net/forum?id=Ut3ml7Hdwx

  24. [33]

    Hardwicke and John P

    Tom E. Hardwicke and John P. A. Ioannidis. Populating the data ark: An attempt to retrieve, preserve, and liberate data from the most highly-cited psychology and psychiatry articles. PLOS ONE, 13 0 (8): 0 e0201856, 2018. doi:10.1371/journal.pone.0201856. URL https://journals.p...

  25. [34]

    Hardwicke, Robert T

    Tom E. Hardwicke, Robert T. Thibault, Jessica E. Kosie, Loukia Tzavella, Theiss Bendixen, Sarah A. Handcock, Vivian E. K \"o neke, and John P.A. Ioannidis. Post-publication critique at top-ranked journals across scientific disciplines: a cross-sectional assessment of policies ...

  26. [35]

    Measuring Goodhart 's law

    Jacob Hilton and Leo Gao. Measuring Goodhart 's law. https://openai.com/index/measuring-goodharts-law/, April 2022. Accessed: 2024-05-23

  27. [36]

    Distilling the knowledge in a neural network

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015

  28. [37]

    Reviewing at ICML 2025

    ICML 2025 Program Chairs . Reviewing at ICML 2025 . https://medium.com/@icml2025pc/reviewing-at-icml-2025-a4676d3505db, January 2025

  29. [38]

    Adversarial examples are not bugs, they are features

    Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. Adversarial examples are not bugs, they are features. arXiv preprint arXiv:1905.02175, 2019. URL https://arxiv.org/abs/1905.02175

  30. [39]

    Towards more rigorous evaluations of language models

    Desi R Ivanova, Ilija Ilievski, and Momchil Konstantinov. Towards more rigorous evaluations of language models. In ICLR Blogposts 2025, 2025. URL https://iclr-blogposts.github.io/2025/blog/towards-more-rigorous-llm-evals/. https://iclr-blogposts.github.io/2025/blog/towards-mor...

  31. [40]

    Tweet on sophia paper

    Keller Jordan. Tweet on sophia paper. https://x.com/kellerjordan0/status/1803865273231118677, June 2024. Tweet

  32. [41]

    Tweet on speedrunning

    Keller Jordan. Tweet on speedrunning. https://x.com/kellerjordan0/status/1890183444380172572, February 2025. Tweet

  33. [42]

    Donoho, and Sanmi Koyejo

    Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, and Sanmi Koyejo. Collapse or thrive? perils and promises of synthetic data in a self-generating world. arXiv preprint arXiv:2410.16713, 2025. URL https://arxiv.org/abs/2410.16713

  34. [43]

    Adam: A method for stochastic optimization

    Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014

  35. [44]

    Paper review: Bayesian model selection, the marginal likelihood, and generalization

    Andreas Kirsch. Paper review: Bayesian model selection, the marginal likelihood, and generalization. https://blog.blackhc.net/2022/06/bayesian-model-selection-marginal-likehood-generalization/, June 2022. Accessed on May 21, 2025

  36. [45]

    Does ‘deep learning on a data diet’ reproduce? overall yes, but GraNd at Initialization does not

    Andreas Kirsch. Does ‘deep learning on a data diet’ reproduce? overall yes, but GraNd at Initialization does not. Transactions on Machine Learning Research, 2024 a

  37. [46]

    Important Prior Work Attribution: RhoLoss and Rho-1

    Andreas Kirsch. Important Prior Work Attribution: RhoLoss and Rho-1 . Public comment on OpenReview, dec 2024 b . URL https://openreview.net/forum?id=0NMzBwqaAJ&noteId=dZcOZBIIVe. Accessed: 2025-05-22. Comment on OpenReview forum ID 0NMzBwqaAJ, note ID dZcOZBIIVe

  38. [47]

    Klein, Michelangelo Vianello, Fred Hasselman, Byron G

    Richard A. Klein, Michelangelo Vianello, Fred Hasselman, Byron G. Adams, Reginald B. Adams, Sinan Alper, Mark Aveyard, Jordan R. Axt, Mayowa T. Babalola, Štěpán Bahník, et al. Many labs 2: Investigating variation in replicability across samples and settings. Advances in Method...

  39. [48]

    Responsible reviewing initiative for NeurIPS 2025, 5 2025

    Piotr Koniusz, Nancy Chen, Marzyeh Ghassemi, Razvan Pascanu, Hsuan-Tien Lin, Lora Aroyo, Francesco Locatello, and Konstantina Palla. Responsible reviewing initiative for NeurIPS 2025, 5 2025. URL https://blog.neurips.cc/2025/05/02/responsible-reviewing-initiative-for-neurips-2025/

  40. [49]

    Thomas S. Kuhn. The Structure of Scientific Revolutions. University of Chicago Press, 1962

  41. [50]

    Distributed representations of sentences and documents

    Quoc Le and Tomas Mikolov. Distributed representations of sentences and documents. arXiv preprint arXiv:1405.4053, May 2014. Published in Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 2014. JMLR: W&CP volume 32

  42. [51]

    Complaints about icml desk rejects, 2024

    Sam Lee. Complaints about icml desk rejects, 2024. URL https://x.com/Samlee1124117/status/1917535739266392260. Tweet

  43. [52]

    Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2024

    Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma. Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2024. URL https://arxiv.org/abs/2305.14342

  44. [53]

    e lle Loosli and St \'e phane Canu. Comments on the

    Ga \"e lle Loosli and St \'e phane Canu. Comments on the "core vector machines: Fast SVM training on very large data sets". Journal of Machine Learning Research, 8 0 (2), 2007

  45. [54]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019

  46. [55]

    Reassessing emnlp 2024’s best paper: Does divergence-based calibration for membership inference attacks hold up? In ICLR Blogposts 2025, 2025

    Pratyush Maini and Anshuman Suri. Reassessing emnlp 2024’s best paper: Does divergence-based calibration for membership inference attacks hold up? In ICLR Blogposts 2025, 2025. URL https://iclr-blogposts.github.io/2025/blog/calibrated-mia/. https://iclr-blogposts.github.io/202...

  47. [56]

    The curse of low task diversity: On the failure of transfer learning to outperform maml and their empirical equivalence, 2022

    Brando Miranda, Patrick Yu, Yu-Xiong Wang, and Sanmi Koyejo. The curse of low task diversity: On the failure of transfer learning to outperform maml and their empirical equivalence, 2022. URL https://arxiv.org/abs/2208.01545

  48. [57]

    Is pre-training truly better than meta-learning? arXiv preprint arXiv:2306.13841, 2023

    Brando Miranda, Patrick Yu, Saumya Goyal, Yu-Xiong Wang, and Sanmi Koyejo. Is pre-training truly better than meta-learning? arXiv preprint arXiv:2306.13841, 2023. URL https://arxiv.org/abs/2306.13841

  49. [58]

    Iclr 2025 best paper was rejected at neurips 2024 with no significant changes, 2024

    Prateek Mittal. Iclr 2025 best paper was rejected at neurips 2024 with no significant changes, 2024. URL https://x.com/prateekmittal_/status/1918350357144420859. Tweet

  50. [59]

    ML Retrospectives: A Venue for Self-Reflection in ML Research

    ML Retrospectives Community . ML Retrospectives: A Venue for Self-Reflection in ML Research . https://ml-retrospectives.github.io/, 2019. Accessed: 2025-05-22

  51. [60]

    Information theoretic guarantees for policy alignment in large language models

    Youssef Mroueh. Information theoretic guarantees for policy alignment in large language models. arXiv preprint arXiv:2406.05883, 2024

  52. [61]

    A discussion of 'adversarial examples are not bugs, they are features': Adversarial examples are just bugs, too

    Preetum Nakkiran. A discussion of 'adversarial examples are not bugs, they are features': Adversarial examples are just bugs, too. Distill, 2019. doi:10.23915/distill.00019.5. https://distill.pub/2019/advex-bugs-discussion/response-5

  53. [62]

    Explaining heterogeneity in medial entorhinal cortex with task-driven neural networks

    Aran Nayebi, Alexander Attinger, Malcolm Campbell, Kiah Hardcastle, Isabel Low, Caitlin S Mallory, Gabriel Mel, Ben Sorscher, Alex H Williams, Surya Ganguli, et al. Explaining heterogeneity in medial entorhinal cortex with task-driven neural networks. Advances in Neural Inform...

  54. [63]

    NeurIPS 2025 Area Chair (AC) Guidelines

    NeurIPS Foundation . NeurIPS 2025 Area Chair (AC) Guidelines . https://neurips.cc/Conferences/2025/AC-Guidelines, 2025 a . Accessed: 2025-05-22

  55. [64]

    NeurIPS 2025 Reviewer Guidelines

    NeurIPS Foundation . NeurIPS 2025 Reviewer Guidelines . https://neurips.cc/Conferences/2025/ReviewerGuidelines, 2025 b . Accessed: 2025-05-22

  56. [65]

    NeurIPS 2025 Senior Area Chair (SAC) Guidelines

    NeurIPS Foundation . NeurIPS 2025 Senior Area Chair (SAC) Guidelines . https://neurips.cc/Conferences/2025/SAC-Guidelines, 2025 c . Accessed: 2025-05-22

  57. [66]

    Turning up the heat: Min-p sampling for creative and coherent LLM outputs

    Minh Nguyen, Andrew Baker, Clement Neo, Allen Roush, Andreas Kirsch, and Ravid Shwartz-Ziv. Turning up the heat: Min-p sampling for creative and coherent LLM outputs. arXiv preprint arXiv:2407.01082, 2024

  58. [67]

    Yet another ICML award fiasco

    Francesco Orabona. Yet another ICML award fiasco. https://parameterfree.com/yet-another-icml-award-fiasco/, August 2023. Blog post on *Parameter-free Learning and Optimization Algorithms*

  59. [68]

    ICLR 2018 Reproducibility Challenge

    Joelle Pineau, Genevieve Fried, Rosemary Nan Ke, and Hugo Larochelle. ICLR 2018 Reproducibility Challenge . https://www.cs.mcgill.ca/ jpineau/ICLR2018-ReproducibilityChallenge.html, 2017. Accessed on May 20, 2025

  60. [69]

    ICLR reproducibility challenge second edition, 2019

    Pineau, Joelle and others . ICLR reproducibility challenge second edition, 2019. https://www.cs.mcgill.ca/ jpineau/ICLR2019-ReproducibilityChallenge.html, 2019. Accessed: 2024-05-22

  61. [70]

    Safety alignment should be made more than just a few tokens deep

    Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. Safety alignment should be made more than just a few tokens deep. arXiv preprint arXiv:2406.05946, 2024. URL https://arxiv.org/abs/2406.05946

  62. [71]

    A siren song of open source reproducibility, examples from machine learning

    Edward Raff and Andrew L Farris. A siren song of open source reproducibility, examples from machine learning. In Proceedings of the 2023 ACM Conference on Reproducibility and Replicability, pp.\ 115--120, 2023

  63. [72]

    Rapid learning or feature reuse? towards understanding the effectiveness of maml

    Aniruddh Raghu, Maithra Raghu, Samy Bengio, and Oriol Vinyals. Rapid learning or feature reuse? towards understanding the effectiveness of maml. arXiv preprint arXiv:1909.09157, 2020. URL https://arxiv.org/abs/1909.09157

  64. [73]

    Do not write that jailbreak paper

    Javier Rando. Do not write that jailbreak paper. In ICLR Blogposts 2025, 2025. URL https://iclr-blogposts.github.io/2025/blog/do-not-write-jailbreak-papers/. https://iclr-blogposts.github.io/2025/blog/do-not-write-jailbreak-papers/

  65. [74]

    Are you a reviewer for neurips24? please read

    Reddit. Are you a reviewer for neurips24? please read. https://www.reddit.com/r/MachineLearning/comments/1d9o8tn/r_are_you_a_reviewer_for_neurips24_please_read/, [Month when posted] 2024. Reddit post

  66. [75]

    How does batch normalization help optimization? Advances in neural information processing systems, 31, 2018

    Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. How does batch normalization help optimization? Advances in neural information processing systems, 31, 2018

  67. [76]

    Strong Model Collapse

    Rylan Schaeffer. Scientific claims require discussion & citation of intentionally-omitted contradictory prior work (part 2). Public Comment on OpenReview forum for "Strong Model Collapse", Dec 2024. URL https://openreview.net/forum?id=et5l9qPUhm&noteId=3cph8WmKrC

  68. [77]

    No free lunch from deep learning in neuroscience: A case study through models of the entorhinal-hippocampal circuit

    Rylan Schaeffer, Mikail Khona, and Ila Fiete. No free lunch from deep learning in neuroscience: A case study through models of the entorhinal-hippocampal circuit. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural Information Proces...

  69. [78]

    Testing assumptions underlying a unified theory for the origin of grid cells

    Rylan Schaeffer, Mikail Khona, Adrian Bertagnoli, Sanmi Koyejo, and Ila Rani Fiete. Testing assumptions underlying a unified theory for the origin of grid cells. arXiv preprint arXiv:2311.16295, 2023 a . URL https://arxiv.org/abs/2311.16295

  70. [79]

    Disentangling fact from grid cell fiction in trained deep path integrators, 2023 b

    Rylan Schaeffer, Mikail Khona, Sanmi Koyejo, and Ila Rani Fiete. Disentangling fact from grid cell fiction in trained deep path integrators, 2023 b . URL https://arxiv.org/abs/2312.03954

  71. [80]

    Self-supervised learning of representations for space generates multi-modular grid cells

    Rylan Schaeffer, Mikail Khona, Tzuhsuan Ma, Cristobal Eyzaguirre, Sanmi Koyejo, and Ila Fiete. Self-supervised learning of representations for space generates multi-modular grid cells. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), Advances in N...

  72. [81]

    Are emergent abilities of large language models a mirage? In A

    Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage? In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (eds.), Advances in Neural Information Processing Systems, volume 36, pp.\ 55565--55581. Curran A...

  73. [82]

    Position: Model collapse does not mean what you think

    Rylan Schaeffer, Joshua Kazdan, Alvan Caleb Arulandu, and Sanmi Koyejo. Position: Model collapse does not mean what you think. arXiv preprint arXiv:2503.03150, 2025 a . URL https://arxiv.org/abs/2503.03150

  74. [83]

    Min-p, max exaggeration: A critical analysis of min-p sampling in language models, 2025 b

    Rylan Schaeffer, Joshua Kazdan, and Yegor Denisov-Blanch. Min-p, max exaggeration: A critical analysis of min-p sampling in language models, 2025 b . URL https://arxiv.org/abs/2506.13681

  75. [84]

    How do large language monkeys get their power (laws)? arXiv preprint arXiv:2502.17578, 2025 c

    Rylan Schaeffer, Joshua Kazdan, John Hughes, Jordan Juravsky, Sara Price, Aengus Lynch, Erik Jones, Robert Kirk, Azalia Mirhoseini, and Sanmi Koyejo. How do large language monkeys get their power (laws)? arXiv preprint arXiv:2502.17578, 2025 c . URL https://arxiv.org/abs/2502.17578

  76. [85]

    Why has predicting downstream capabilities of frontier AI models with scale remained elusive? arXiv preprint arXiv:2406.04391, 2025 d

    Rylan Schaeffer, Hailey Schoelkopf, Brando Miranda, Gabriel Mukobi, Varun Madan, Adam Ibrahim, Herbie Bradley, Stella Biderman, and Sanmi Koyejo. Why has predicting downstream capabilities of frontier AI models with scale remained elusive? arXiv preprint arXiv:2406.04391, 2025...

  77. [86]

    Ai models collapse when trained on recursively generated data

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data. Nature, 631 0 (8022): 0 755--759, 2024. ISSN 1476-4687. doi:10.1038/s41586-024-07566-y

  78. [87]

    A unified theory for the origin of grid cells through the lens of pattern formation

    Ben Sorscher, Gabriel Mel, Surya Ganguli, and Samuel Ocko. A unified theory for the origin of grid cells through the lens of pattern formation. Advances in neural information processing systems, 32, 2019

  79. [88]

    Deconstructing self-supervised monocular reconstruction: The design decisions that matter

    Jaime Spencer, Chris Russell, Simon Hadfield, and Richard Bowden. Deconstructing self-supervised monocular reconstruction: The design decisions that matter. Transactions on Machine Learning Research, 2022

  80. [89]

    Learning to summarize with human feedback

    Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in neural information processing systems, 33: 0 3008--3021, 2020

  81. [90]

    Tenenbaum, and Phillip Isola

    Yonglong Tian, Yue Wang, Dilip Krishnan, Joshua B. Tenenbaum, and Phillip Isola. Rethinking few-shot image classification: a good embedding is all you need? arXiv preprint arXiv:2003.11539, 2020. URL https://arxiv.org/abs/2003.11539

  82. [91]

    Position: Considerations for differentially private learning with large-scale public pretraining

    Florian Tram\` e r, Gautam Kamath, and Nicholas Carlini. Position: Considerations for differentially private learning with large-scale public pretraining. In Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenka...

  83. [92]

    Discussion on doc2vec issues, 2024

    Hacker News User. Discussion on doc2vec issues, 2024. URL https://news.ycombinator.com/item?id=38689976. Hacker News thread

  84. [93]

    Reproducibility issues with doc2vec, 2014

    StackExchange User. Reproducibility issues with doc2vec, 2014. URL https://stats.stackexchange.com/questions/123562/. StackExchange question

  85. [94]

    Comments on doc2vec reproducibility, 2015

    StackExchange User. Comments on doc2vec reproducibility, 2015. URL https://stats.stackexchange.com/a/222501/62060. StackExchange answer

  86. [95]

    Announcing the NeurIPS 2021 datasets and benchmarks track, April 2021

    Joaquin Vanschoren and Serena Yeung. Announcing the NeurIPS 2021 datasets and benchmarks track, April 2021. URL https://blog.neurips.cc/2021/04/07/announcing-the-neurips-2021-datasets-and-benchmarks-track/. NeurIPS Blog

  87. [96]

    Neurips 2014 rejected knowledge distillation, 2019

    Oriol Vinyals. Neurips 2014 rejected knowledge distillation, 2019. URL https://x.com/OriolVinyalsML/status/1129420305246629899. Tweet

  88. [97]

    ensemble everything everywhere

    Jie Zhang, Christian Schlarmann, Kristina Nikolić, Nicholas Carlini, Francesco Croce, Matthias Hein, and Florian Tramèr. Evaluating the robustness of the "ensemble everything everywhere" defense, 2025. URL https://arxiv.org/abs/2411.14834

  89. [98]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  90. [99]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  91. [100]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  92. [101]

    Ensemble Everything Everywhere

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.