REVIEW 2 major objections 4 minor 122 references
Modern AI advances mainly by operational rigor—benchmarks and deployment reliability—while conceptual clarity and scientific understanding lag, explaining both its speed and its uncertainties.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 16:22 UTC pith:TBZXODEC
load-bearing objection Clean conceptual synthesis that names why deep learning runs on operational rigor; useful organizing lens, not a tested theory. the 2 major comments →
The Role of Rigor in Artificial Intelligence
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The distinctive trajectory of AI arises from how conceptual, epistemic, and operational rigor interact across paradigms, resulting in the primacy of operational rigor in modern deep learning; that primacy explains both the field’s rapid capability gains and its lasting uncertainties, and it clarifies what is required to turn AI into a mature science and reliable technology.
What carries the argument
A three-part framework of rigor—conceptual (clear foundational concepts and paradigms), epistemic (reproducibility, predictability, explainability), and operational (benchmarks, reliability, and safety practices)—used as the diagnostic lens for AI’s history, methods, and future bottlenecks.
Load-bearing premise
The analysis rests on the premise that this three-way split of rigor is the right and sufficiently complete way to diagnose AI’s scientific status, rather than some other taxonomy of standards.
What would settle it
If a future AI paradigm (or a careful historical re-analysis of current deep learning) showed that lasting capability gains required simultaneous advances in conceptual and epistemic rigor rather than operational metric-chasing, or if an alternative rigor taxonomy better predicted progress and failure modes, the central claim would be undercut.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that AI's distinctive trajectory—rapid capability gains with persistent conceptual and scientific uncertainty—arises from the interaction of three forms of rigor: conceptual (clarity of foundational terms and paradigms), epistemic (reproducibility, predictability, explainability), and operational (benchmarks, reliability, and safety). It applies this framework to contested notions of intelligence and understanding, the empirical character of deep learning, the strengths and pathologies of benchmarks, and the historical succession of symbolic, classical statistical-learning, and connectionist/deep-learning paradigms. The central claim is that modern deep learning has elevated operational rigor above the other two, which both explains progress and clarifies the obstacles to maturing AI as a science and reliable technology.
Significance. If accepted as a useful analytic lens, the paper supplies a coherent organizing vocabulary for a multidisciplinary field whose progress is often discussed in fragmented or polemical terms. The historical sketch of paradigms and the treatment of benchmarks, reproducibility distinctions, and alignment are standard but well-integrated; the explicit contrast between AGI (favoring operational rigor) and alignment (requiring conceptual and epistemic rigor) is a clear contribution. The work is philosophical and taxonomic rather than theorematic or empirical; its value lies in clarifying structure and priorities rather than in new derivations or falsifiable predictions. Strengths include careful citation of the literature and a measured tone that avoids both hype and pure critique.
major comments (2)
- The three-part taxonomy is introduced by stipulation in the Introduction and then applied throughout; its exhaustiveness and superiority relative to alternatives (e.g., Olteanu et al. [1] or classical philosophy-of-science categories) are not independently argued or tested. Because the central claim—that the distinctive trajectory of AI arises from how these forms interact, with operational rigor primary under deep learning—depends on this partition, the manuscript should either (a) defend the partition more explicitly against nearby alternatives or (b) state more clearly that the framework is provisional and heuristic rather than uniquely privileged.
- Section 4.2's historical narrative (symbolic → classical statistical learning → connectionism/deep learning) is standard and well-cited, but the claim that the current primacy of operational rigor is "historically contingent rather than inevitable" remains under-supported. A brief discussion of what would count as evidence that a future paradigm rebalanced the three forms (or of counter-examples already present) would strengthen the load-bearing historical claim.
minor comments (4)
- The footnote distinguishing the present three-part scheme from Olteanu et al. [1] is useful but brief; a short paragraph in the Introduction or Conclusion comparing the two taxonomies would help readers locate the contribution.
- Section 2.3's discussion of explainability vs. interpretability is careful, yet the claim that deep learning "defies effective hierarchical abstraction" could be sharpened with one or two concrete examples of failed localization (e.g., attribution of a particular failure mode to data vs. architecture vs. optimization).
- Occasional informal phrasing ("alchemy," "jagged intelligence") is already hedged, but ensuring each such term is immediately tied to a citation or definition would further reduce ambiguity.
- References to scaling laws and infinite-width theories are accurate; a brief note on known caveats (already alluded to via Hooker [48]) would keep the predictability discussion balanced.
Circularity Check
No significant circularity: the three-part rigor taxonomy is a stipulated analytic lens applied to known history, not a derivation that reduces to its inputs by construction.
full rationale
This is a philosophical/analytic position paper, not a quantitative derivation. The central claim—that AI's distinctive trajectory (rapid capability growth with lagging conceptual and scientific understanding) arises from the interaction of conceptual, epistemic, and operational rigor, with operational rigor becoming primary under deep learning—is an organizing thesis introduced by stipulation in the Introduction and then used to re-describe well-known historical paradigms (symbolic AI, classical statistical learning, connectionism/deep learning), the empirical character of modern deep learning, and the roles of benchmarks and alignment. There are no equations, fitted parameters, or 'predictions' that reduce by construction to inputs. The only mild definitional element is the author's introduction of the three categories themselves; once accepted as a provisional lens, the subsequent historical and diagnostic claims do not loop back to force those categories. Citations are overwhelmingly to independent sources (Turing, McCarthy, Legg & Hutter, Kaplan et al. scaling laws, ImageNet, Goodhart's law literature, alignment papers, etc.); the single self-citation is the author's own prior NeurIPS paper on n-gram statistics of transformers, which is used only as an illustrative example of training-data regurgitation and is not load-bearing for the framework or the primacy-of-operational-rigor thesis. No uniqueness theorem is imported from the author's prior work, no ansatz is smuggled via self-citation, and no known empirical pattern is merely renamed as a new result. Score 1 reflects only the ordinary philosophical practice of defining terms and then applying them; the paper is self-contained as an interpretive essay and exhibits no circular reduction of the kind the analyzer is charged to detect.
Axiom & Free-Parameter Ledger
axioms (4)
- ad hoc to paper Rigor in AI is usefully partitioned into conceptual, epistemic, and operational forms that interact across paradigms.
- domain assumption Modern deep learning is characterized by a tight feedback loop in which the same metrics used for evaluation are also optimization targets.
- domain assumption AI produces the artifacts it studies, so conceptual and epistemic inquiry always chase a moving target.
- domain assumption Earlier AI paradigms (symbolic, classical statistical learning) embodied different balances of the three rigor forms than deep learning.
invented entities (1)
-
Three-part rigor framework (conceptual / epistemic / operational)
no independent evidence
read the original abstract
Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations associated with mature disciplines. Unlike traditional sciences, where reliable technology typically emerges from theoretical understanding, modern AI has progressed largely through performance-driven iteration and "alchemical" experimentation. This tension motivates a systematic analysis of AI through the lens of rigor. We introduce a three-part framework consisting of conceptual rigor (clarifying foundational concepts), epistemic rigor (establishing scientific understanding), and operational rigor (ensuring reliable performance and deployment). Using this framework, we analyze competing conceptions of intelligence and understanding, the strengths and limitations of the empirical approach to deep learning, the power and pitfalls of benchmarks, and the obstacles to theory development posed by modern AI systems. We argue that the distinctive trajectory of AI arises from how forms of rigor interact across paradigms, resulting in the primacy of operational rigor in modern deep learning. This perspective helps explain both AI's rapid advances and its persistent uncertainties, while clarifying the challenges involved in transforming AI into a mature science and reliable technology.
Reference graph
Works this paper leans on
-
[1]
Rigor in AI: Doing rigorous AI work requires a broader, responsible AI-informed conception of rigor
Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz, Flavio Calmon, Margaret Mitchell, Michael Ekstrand, Reuben Binns, and Solon Barocas. Rigor in AI: Doing rigorous AI work requires a broader, responsible AI-informed conception of rigor. In The Thirty-Ninth Annual Conference on Neural Information Processing Systems Position Pa...
2025
-
[2]
Alan M. Turing. Computing machinery and intelligence.Mind, 59(236):433–460, 1950
1950
-
[3]
A proposal for the Dartmouth summer research project on Artificial Intelligence, 1955
John McCarthy, Marvin Minsky, Nathaniel Rochester, and Claude Shannon. A proposal for the Dartmouth summer research project on Artificial Intelligence, 1955. Proposal for the 1956 Dart- mouth workshop. URL:http://jmc.stanford.edu/articles/dartmouth/dartmouth.pdf
1955
-
[4]
Luciano Floridi and Anna C. Nobre. Anthropomorphising machines and computerising minds: The crosswiring of languages between artificial intelligence and brain & cognitive sciences.Cen- tre for Digital Ethics (CEDE) Research Paper, 2024. Available at SSRN:https://ssrn.com/ abstract=4738331
2024
-
[5]
Pei Wang.Non-Axiomatic Reasoning System: Exploring the Essence of Intelligence. Ph.d. thesis, Indiana University, Bloomington, IN, USA, 1995
1995
-
[6]
A collection of definitions of intelligence
Shane Legg and Marcus Hutter. A collection of definitions of intelligence. InAdvances in Artificial General Intelligence: Concepts, Architectures and Algorithms, pages 17–24, Amsterdam, The Netherlands, 2007. IOS Press. 13
2007
-
[7]
On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019
Fran¸ cois Chollet. On the measure of intelligence.arXiv preprint arXiv:1911.01547, 2019
Pith/arXiv arXiv 1911
-
[8]
Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
S´ ebastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. Sparks of artificial general intelligence: Early experiments with GPT-4, 2023
2023
-
[9]
Godfather of AI
60 Minutes. “Godfather of AI”: Geoffrey Hinton – the 60 minutes interview, Oct 2023. YouTube video, 13:12. URL:https://www.youtube.com/watch?v=qrvK_KuIeJk
2023
-
[10]
Godfather of AI
The Royal Institution. Will AI outsmart human intelligence? - with “Godfather of AI” Ge- offrey Hinton, Jul 2025. YouTube video, 47:15. URL:https://www.youtube.com/watch?v= IkdziSLYzHw
2025
-
[11]
A conversation with Yann LeCun AI: Lifeline or landmine?, Feb
World Governments Summit. A conversation with Yann LeCun AI: Lifeline or landmine?, Feb
-
[12]
URL:https://www.youtube.com/watch?v=rf9jgZYAni8
YouTube video, 24:51. URL:https://www.youtube.com/watch?v=rf9jgZYAni8
-
[13]
Krakauer, John W
David C. Krakauer, John W. Krakauer, and Melanie Mitchell. Large language models and emergence: A complex systems perspective, 2025
2025
-
[14]
Krakauer
Melanie Mitchell and David C. Krakauer. The Debate over Understanding in AI’s Large Language Models.Proceedings of the National Academy of Sciences, 120(13):e2215907120, March 2023
2023
-
[15]
Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, Fran¸ cois Candelon, and Karim R
Fabrizio Dell’Acqua, Edward McFowland III, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran, Lisa Krayer, Fran¸ cois Candelon, and Karim R. Lakhani. Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowl- edge worker productivity and quality. Working Paper 24-013, Harvard Business S...
2023
-
[16]
Andrej Karpathy. Jagged Intelligence, 2024. Tweet on X (formerly Twitter), July 2024. URL: https://x.com/karpathy/status/1816531576228053133
arXiv 2024
-
[17]
Why 9.11 is larger than 9.9
OpenAI Community. Why 9.11 is larger than 9.9. . . . . . incredible. Online forum post, Jul 2024. Discussion on comparing numeric values; community.openai.com thread 869824. URL:https: //community.openai.com/t/why-9-11-is-larger-than-9-9-incredible/869824
2024
-
[18]
a is b” fail to learn “b is a
Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. The reversal curse: LLMs trained on “a is b” fail to learn “b is a”. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[19]
Hutchinson, London, 1949
Gilbert Ryle.The Concept of Mind. Hutchinson, London, 1949
1949
-
[20]
Fodor.The Language of Thought
Jerry A. Fodor.The Language of Thought. Harvard University Press, 1975
1975
-
[21]
Psychological predicates
Hilary Putnam. Psychological predicates. In W. H. Capitan and D. D. Merrill, editors,Art, Mind, and Religion, pages 37–48. University of Pittsburgh Press, 1967
1967
-
[22]
The symbol grounding problem.Physica D: Nonlinear Phenomena, 42(1–3):335– 346, 1990
Stevan Harnad. The symbol grounding problem.Physica D: Nonlinear Phenomena, 42(1–3):335– 346, 1990
1990
-
[23]
From System 1 Deep Learning to System 2 Deep Learning
Yoshua Bengio. From System 1 Deep Learning to System 2 Deep Learning. NeurIPS 2019 keynote talk on YouTube, 2019. Accessed 2026-01-29. URL:https://www.youtube.com/watch? v=FtUbMG3rlFs
2019
-
[24]
Othello-GPT
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Vi´ egas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. InInternational Conference on Learning Representations (ICLR) 2023, 2023. Also known as “Othello-GPT” study on transformer emergent representations
2023
-
[25]
Genie 3: A new frontier for world mod- els
Jack Parker-Holder and Shlomi Fruchter. Genie 3: A new frontier for world mod- els. DeepMind Blog, Aug 2025. Available at:https://deepmind.google/blog/ genie-3-a-new-frontier-for-world-models/. 14
2025
-
[26]
A Path Towards Autonomous Machine Intelligence
Yann LeCun. A Path Towards Autonomous Machine Intelligence. Working paper, OpenReview,
-
[27]
Version 0.9.2, June 27, 2022
2022
-
[28]
Meta’s AI chief Yann LeCun on AGI, open-source, and AI risk.TIME, February
Charlotte Alter. Meta’s AI chief Yann LeCun on AGI, open-source, and AI risk.TIME, February
-
[29]
Interview with Yann LeCun
-
[30]
Meaning without reference in large language models
Steven Piantadosi and Felix Hill. Meaning without reference in large language models. In NeurIPS 2022 Workshop on Neuro Causal and Symbolic AI (nCSI), 2022
2022
-
[31]
Nastase, Martin Chodorow, Mengru Wu, and Ping Li
Qianwen Xu, Yixuan Peng, Samuel A. Nastase, Martin Chodorow, Mengru Wu, and Ping Li. Large language models without grounding recover non-sensorimotor but not sensorimotor fea- tures of human concepts.Nature Human Behaviour, 9(9):1871–1886, 2025
2025
-
[32]
Bender and Alexander Koller
Emily M. Bender and Alexander Koller. Climbing towards NLU: On meaning, form, and under- standing in the age of data. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5185–5198, Online, July 2020. Association for Computational Linguistics
2020
-
[33]
Universal intelligence: A definition of machine intelligence
Shane Legg and Marcus Hutter. Universal intelligence: A definition of machine intelligence. Minds and Machines, 17(4):391–444, 2007
2007
-
[34]
Trends in the dollar training cost of machine learning systems
Ben Cottier. Trends in the dollar training cost of machine learning systems. Epoch AI Blog, 2023. Published January 31, 2023; accessed 2026-01-29. URL:https://epoch.ai/blog/ trends-in-the-dollar-training-cost-of-machine-learning-systems
2023
-
[35]
How much power will frontier AI training demand in 2030? Epoch AI Blog, 2025
Josh You and David Owen. How much power will frontier AI training demand in 2030? Epoch AI Blog, 2025. Published August 11, 2025; accessed 2026-01-29. URL:https://epoch.ai/blog/ power-demands-of-frontier-ai-training
2030
-
[36]
The AI index 2023 annual report
Nestor Maslej, Loredana Fattorini, Erik Brynjolfsson, John Etchemendy, Katrina Ligett, Terah Lyons, James Manyika, Helen Ngo, Juan Carlos Niebles, Vanessa Parli, Yoav Shoham, Rus- sell Wald, Jack Clark, and Raymond Perrault. The AI index 2023 annual report. Re- port, Institute for Human-Centered Artificial Intelligence (HAI), Stanford University, 2023. In...
2023
-
[37]
Unreproducible research is reproducible
Xavier Bouthillier, C´ esar Laurent, and Pascal Vincent. Unreproducible research is reproducible. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors,Proceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 725–734. PMLR, 09–15 Jun 2019
2019
-
[38]
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. Deep reinforcement learning that matters. InProceedings of the 32nd AAAI Conference on Artificial Intelligence, AAAI’18, pages 3207–3214. AAAI Press, 2018
2018
-
[39]
Are GANs created equal? A large-scale study.arXiv preprint arXiv:1711.10337, 2017
Mario Lucic, Karol Kurach, Marcin Michalski, Sylvain Gelly, and Olivier Bousquet. Are GANs created equal? A large-scale study.arXiv preprint arXiv:1711.10337, 2017
Pith/arXiv arXiv 2017
-
[40]
Random search and reproducibility for neural architecture search
Liam Li and Ameet Talwalkar. Random search and reproducibility for neural architecture search. In Ryan P. Adams and Vibhav Gogate, editors,Proceedings of The 35th Uncertainty in Artificial Intelligence Conference, volume 115 ofProceedings of Machine Learning Research, pages 367–377. PMLR, 22–25 Jul 2020
2020
-
[41]
Julian D
Moritz Herrmann, F. Julian D. Lange, Katharina Eggensperger, Giuseppe Casalicchio, Marcel Wever, Matthias Feurer, David R¨ ugamer, Eyke H¨ ullermeier, Anne-Laure Boulesteix, and Bernd Bischl. Position: why we must rethink empirical research in machine learning. InProceedings of the 41st International Conference on Machine Learning, ICML’24. JMLR.org, 2024
2024
-
[42]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models, 2020. 15
2020
-
[43]
Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, An- drei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell. Overcoming catastrophic for- getting in neural networks.Proceedings of the National Academy of Sciences, 114(13):3...
2017
-
[44]
Extracting training data from large language models
Nicholas Carlini, Florian Tram` er, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, ´Ulfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650. USENIX Association, August 2021
2021
-
[45]
Understanding transformers via n-gram statistics
Timothy Nguyen. Understanding transformers via n-gram statistics. InAdvances in Neural Information Processing Systems 37 (NeurIPS 2024). Curran Associates, Inc., 2024
2024
-
[46]
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Cl´ ement Hongler. Neural tangent kernel: Convergence and generalization in neural networks. InAdvances in Neural Information Processing Systems, pages 8580–8589, Red Hook, NY, USA, 2018. Curran Associates, Inc
2018
-
[47]
Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl- Dickstein, and Jeffrey Pennington
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl- Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. InAdvances in Neural Information Processing Systems, volume 32, 2019. NeurIPS 2019; accessed 2026-01-29
2019
-
[48]
Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari. Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit. In Kamalika Chaudhuri and Ruslan Salakhut- dinov, editors,Proceedings of Machine Learning Research, volume 99 ofProceedings of Machine Learning Research, pages 2388–2464. PMLR, 2019
2019
-
[49]
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer.arXiv preprint arXiv:2203.03466, 2022
Pith/arXiv arXiv 2022
-
[50]
Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianine- jad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically, 2017
2017
-
[51]
On the slow death of scaling
Sara Hooker. On the slow death of scaling. SSRN Scholarly Paper ID 5877662. Available at SSRN:https://ssrn.com/abstract=5877662, December 2025
2025
-
[52]
Michaud, Berkan Ottlik, and Joseph Turnbull
Jamie Simon, Daniel Kunin, Alexander Atanasov, Enric Boix-Adser` a, Blake Bordelon, Jeremy Cohen, Nikhil Ghosh, Florentin Guth, Arthur Jacot, Mason Kamb, Dhruva Karkada, Eric J. Michaud, Berkan Ottlik, and Joseph Turnbull. There will be a scientific theory of deep learning. arXiv preprint arXiv:2604.21691, 2026
Pith/arXiv arXiv 2026
-
[53]
Goodfellow, Jonathon Shlens, and Christian Szegedy
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adver- sarial examples. InInternational Conference on Learning Representations (ICLR), 2015
2015
-
[54]
Accounting for variance in machine learning benchmarks
Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Nazanin Mohammadi Sepahvand, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Tal Arbel, Chris Pal, Gael Varoquaux, and Pascal Vincent. Accounting for variance in machine learning benchmarks. In A. Smola, A. Dimakis, and ...
2021
-
[55]
AI researchers allege that machine learning is alchemy
Matthew Hutson. AI researchers allege that machine learning is alchemy. News article,Sci- ence, May 3 2018. Accessed: 2025-01-05. URL:https://www.science.org/content/article/ ai-researchers-allege-machine-learning-alchemy
2018
-
[56]
Brief introduction to deep learning and the “alchemy” controversy
Sanjeev Arora. Brief introduction to deep learning and the “alchemy” controversy. Video, YouTube, 2019. Presented at Deep Learning: Alchemy or Science?, Institute for Advanced Study. URL:https://www.youtube.com/watch?v=kqhg-o-KEns. 16
2019
-
[57]
Zico Kolter
J. Zico Kolter. Is this really science? A lukewarm defense of alchemy. InWorkshop on Scientific Methods for Understanding Neural Networks, NeurIPS 2024, 2024
2024
-
[58]
Vapnik.The Nature of Statistical Learning Theory
Vladimir N. Vapnik.The Nature of Statistical Learning Theory. Springer, New York, 2 edition, 2000
2000
-
[59]
Cambridge University Press, 2014
Shai Shalev-Shwartz and Shai Ben-David.Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press, 2014
2014
-
[60]
Approximation theory of the MLP model in neural networks.Acta Numerica, 8:143–195, 1999
Allan Pinkus. Approximation theory of the MLP model in neural networks.Acta Numerica, 8:143–195, 1999
1999
-
[61]
Springer Science & Business Media, New York, second edition, 2006
Jorge Nocedal and Stephen Wright.Numerical Optimization. Springer Science & Business Media, New York, second edition, 2006
2006
-
[62]
Introduction to online convex optimization, 2023
Elad Hazan. Introduction to online convex optimization, 2023
2023
-
[63]
Reconciling modern machine- learning practice and the bias–variance trade-off.Proc
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine- learning practice and the bias–variance trade-off.Proc. Natl. Acad. Sci. U.S.A., 116(32):15849– 15854, 2019. PMCID: PMC6689936, PMID: 31341078
2019
-
[64]
The implicit bias of gradient descent on separable data, 2017
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data, 2017
2017
-
[65]
No spurious local minima in nonconvex low rank problems: A unified geometric analysis, 2017
Rong Ge, Chi Jin, and Yi Zheng. No spurious local minima in nonconvex low rank problems: A unified geometric analysis, 2017
2017
-
[66]
Du, Jason D
Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai. Gradient descent finds global minima of deep neural networks, 2018
2018
-
[67]
Knowledge editing for large language models: A survey.ACM Computing Surveys, 57(3):59:1–59:35, 2024
Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. Knowledge editing for large language models: A survey.ACM Computing Surveys, 57(3):59:1–59:35, 2024
2024
-
[68]
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers. InAdvances in Neural Information Processing Systems 37, pages 83345–83373, Vancouver, Canada, 2024. Neural Information Processing Systems Foundation, Inc
2024
-
[69]
Toward faithful retrieval-augmented generation with sparse autoencoders, 2025
Guangzhi Xiong, Zhenghao He, Bohan Liu, Sanchit Sinha, and Aidong Zhang. Toward faithful retrieval-augmented generation with sparse autoencoders, 2025
2025
-
[70]
Cor- rectness assessment of code generated by large language models using internal representations
Tuan-Dung Bui, Thanh Trong Vu, Thu-Trang Nguyen, Son Nguyen, and Hieu Dinh Vo. Cor- rectness assessment of code generated by large language models using internal representations. Journal of Systems and Software, 230:112570, December 2025
2025
-
[71]
Zachary C. Lipton. The Mythos of Model Interpretability: In Machine Learning, the Concept of Interpretability Is Both Important and Poorly Defined.Commun. ACM, 61(10):36–43, September 2018
2018
-
[72]
Sarthak Jain and Byron C. Wallace. Attention Is Not Explanation. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3543–3556, 2019
2019
-
[73]
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. Attention is not not explanation. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 11–20, Hong Kong, China, November 2019. ...
2019
-
[74]
Everything, everywhere, all at once: Is mechanistic interpretability identifiable? InThe Thirteenth International Confer- ence on Learning Representations, 2025
Maxime M´ eloux, Silviu Maniu, Fran¸ cois Portet, and Maxime Peyrard. Everything, everywhere, all at once: Is mechanistic interpretability identifiable? InThe Thirteenth International Confer- ence on Learning Representations, 2025. 17
2025
-
[75]
Schiller, Filippos Stamatiou, and Anders Søgaard
Iwan Williams, Ninell Oldenburg, Ruchira Dhar, Joshua Hatherley, Constanza Fierro, Nina Ra- jcic, Sandrine R. Schiller, Filippos Stamatiou, and Anders Søgaard. Mechanistic interpretability needs philosophy.arXiv:2506.18852 [cs.CL], 2025. Preprint; accessed 2026-01-29
Pith/arXiv arXiv 2025
-
[76]
Cambridge University Press, 2 edition, 2009
Judea Pearl.Causality: Models, Reasoning, and Inference. Cambridge University Press, 2 edition, 2009
2009
-
[77]
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 248–255, 2009
2009
-
[78]
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona. Recognition in terra incognita. InProceedings of the European Conference on Computer Vision (ECCV), pages 456–473, Munich, Germany,
-
[79]
Zemel, Wieland Brendel, Matthias Bethge, and Felix A
Robert Geirhos, J¨ orn-Henrik Jacobsen, Claudio Michaelis, Richard S. Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
2020
-
[80]
Benchmarking neural network robustness to common corruptions and perturbations
Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. InInternational Conference on Learning Representations, 2019
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.