REVIEW 4 major objections 5 minor 59 references
A not-too-simple solution to Goodman's new riddle of induction in the age of AI
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that Goodman's new riddle of induction dissolves once direct measurements are required to yield convex error-boxes and model complexity is measured by the shortest empirically equivalent formulation.
desk verdict A transparent synthesis of Gärdenfors and the author's own 2013 model that makes a real proposal, but the green/grue asymmetry is largely stipulated by Postulate 1, so 'solution' overstates what is demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the pair consisting of Postulate 1 and Definition 3. Postulate 1 requires the outcome of a single direct measurement of a property $Q$ to be a central value $Q_0$ together with a convex error-box containing $Q_0$; it rules out split error regions as legitimate direct outcomes. Definition 3 defines the epistemic complexity $C(M)$ of a model $M$ as the minimum, over all logically and empirically equivalent formulations $M'\equiv M$, of the length of the assumptions $P(M')$; here empirical equivalence (Definition 2) requires the translation to preserve the measurement outcomes and precisions of every measurable property. Together these make the $\Xi=0$ reformulation inadmissible, because that rewriting breaks empirical equivalence and demands a directly measurable $\Xi$ with a connected error-box, so the minimum is non-trivial and reformulation-independent. The paper's one-sentence summary of the mechanism: 'it is the constraint of convexity that enables a non-trivial notion of complexity.'
What would settle it
A single direct measurement whose reported uncertainty is a genuinely disconnected or non-convex region—a multimodal expected distribution for one measurement—would falsify Postulate 1 and collapse the complexity gap, because grue would then be as directly measurable as green. Alternatively, any case of model selection that the broad scientific community rejects but Definition 4 admits, or accepts but Definition 4 rules out, would falsify the paper's descriptive claim.
Extended reading notes
Core claim
The central claim is that the asymmetry between green and grue is not a matter of language or convention but of measurability: although any model can be expressed as $\Xi=0$, no empirically equivalent formulation can make $\Xi$ a directly measurable quantity with a central value and a connected error-box, because the error-box of a grue measurement splits at the critical time $t_0$ and, for emeralds first seen before $t_0$, remains undetermined later. The paper formalizes this through Postulate 1 (convexity of single direct measurement outcomes) and Definition 3 (epistemic complexity as minimum length over logically and empirically equivalent formulations), and shows that the grue model is then strictly more complex than the green model with no empirical advantage, so Definition 4 rules it out. The paper stresses that this solves the new riddle—identifying the hidden assumptions behind scientists' actual model selection—and not the old riddle of justifying induction by future success.
Load-bearing premise
The argument hangs on Postulate 1, that a single direct measurement always reports a central value inside one convex error-box and never a split or disconnected region; if scientists could legitimately report a split region as a direct measurement, grue would be directly measurable and the complexity advantage of green disappears.
Editorial extensions
If this is right
- The selection rule of Definition 4 gives a precise meaning to 'explaining more with less': any model that is more complex and no more accurate than a rival is ruled out, with no trade-off involved.
- Because epistemic complexity is invariant under logically and empirically equivalent reformulations, simplicity comparisons remain meaningful across very different theories, including across scientific revolutions.
- The same framework dismisses conspiracy theories: extra ad-hoc assumptions like '$\Xi$-people' buy conciseness only by sacrificing empirical accuracy, because the corresponding measurements are not available.
- Bayesian confirmation requires prior probabilities; the paper argues the prior choice can be anchored only by this reformulation-independent complexity measure, otherwise the $\Xi$ trick makes priors arbitrary.
- This is a solution to the new riddle, not the old one: it describes the hidden assumptions behind scientists' actual choices rather than promising to justify the future success of science.
Reading between the lines
- Extending the paper's approach: any predicate that forces split error-boxes in a shared direct-measurement basis should be empirically disfavored exactly like grue; this is a testable prediction for language design and machine-learning feature engineering.
- The paper asserts rather than proves that the minimum in Definition 3 is non-trivial and attained; a formal proof, or a realistic model class where the minimum fails to be attained, would either complete or stress the foundation.
- A broader research program implied here is grounding non-empirical epistemic values (simplicity, naturalness, projectibility) in measurability constraints rather than in metaphysical natural kinds or entrenchment.
- For AI, the framework suggests a concrete auditing rule: a learned model that gains apparent simplicity by redefining its inputs so that measurement error-boxes split is a 'grue model' and should be downgraded; this could be operationalized as a test for shortcut learning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a solution to Goodman's new riddle of induction by combining two ideas: (i) a postulate that the outcome of a single direct measurement is always a central value with a convex error region, and (ii) an "epistemic complexity" defined as the minimum length of the assumptions over all logically and empirically equivalent reformulations. The author argues that a grue-like formulation cannot be empirically equivalent to the standard green formulation because measuring grue requires a non-convex, split error-box, so any grue reformulation is more complex without empirical advantage. The paper also argues that this combination solves the riddle in the sense of identifying the hidden assumptions behind scientific model selection, and it offers historical and conspiracy-theory examples as illustrations.
Significance. If the central claims hold, the paper would provide a principled, reformulation-invariant criterion that rules out grue-like predicates without appealing to psychological entrenchment or subjective simplicity, and it would connect a classic philosophical riddle to practical questions in AI model selection. The paper is clearly written, engages seriously with the literature on conceptual spaces and complexity, and explicitly identifies the adequacy conditions for a solution. The main contribution is conditional, however: the two load-bearing assumptions—convexity as a constitutive feature of direct measurements and the non-triviality of the minimum in Definition 3—are asserted rather than demonstrated. The paper's descriptive examples are suggestive but do not yet provide independent evidence for those assumptions.
major comments (4)
- [Section 2.1, Postulate 1] The central asymmetry between green and grue rests on excluding split error regions from the outcome of a single direct measurement. The author explicitly acknowledges that Post. 1, together with Def. 1, offers an "implicit (partial) definition of direct measurements" and that "it is up to the model to decide which are the direct measurements." Consequently, the claim that grue is not directly measurable is not an empirical result but a consequence of the postulate; a proponent of grue could simply stipulate that a grue-meter's single-readout distribution is bimodal and insist that it is direct. To carry the argument, the paper needs an independent characterization of directness, or at least a systematic empirical survey showing that no accepted scientific direct measurement has a non-convex error-box. The two examples given (room temperature and Fig. 2) are not enough to support the load placed on this postulate.
- [Section 2.2, Definition 3] The epistemic complexity C(M) is defined as a minimum over "all possible equivalent formulations (in any language)" of length[P(M')]. The paper does not specify the length function, does not prove that the minimum exists, and does not prove that it is non-trivial for the relevant cases. The statement that restricting to logically and empirically equivalent formulations ensures that the Xi = 0 formulation is no longer legitimate, and that the shortest formulation is "in general, not trivial anymore," is an assertion. Without a precise definition of the language class and a proof that no short, convex, directly measurable reformulation can encode the same empirical content, the complexity gap between green and grue could collapse, and the definition might be ill-defined or yield a constant for all models.
- [Section 2.2, Definition 2] The empirical equivalence relation used in Definition 3 requires "same precision and same outcome" for each measurable property. Precision is not an external fact; it is part of the model assumptions, since Definition 1 includes Delta(b) for every directly measurable quantity b in B. This creates a circularity: whether two formulations are equivalent depends on the very precision values whose effect on complexity is being assessed. The paper should specify how equivalence is to be judged independently of the model's own stipulations, or explain why this dependence does not undermine the claimed reformulation independence.
- [Sections 3.2 and 5] The paper's justification strategy is descriptive accuracy: the model is accepted if no counterexample against scientific consensus exists. But the paper never operationalizes "broad scientific consensus" nor explains how to identify a counterexample independently of the framework. The Bielefeld conspiracy example is essentially a restatement of the same unavailability-of-records argument used for grue, so it does not provide independent support. The "no counterexample" claim is therefore a conjecture rather than a test of the model; the author should either provide a falsification protocol or soften the claim accordingly.
minor comments (5)
- [Section 2.1, Postulate 1'] The paper says that all important conclusions are maintained under the more general Postulate 1', but it does not demonstrate this; a short verification would be helpful, especially given that the probability-distribution formulation is what makes the convexity claim checkable.
- [Section 2.2, Definition 3] The term length[P(M')] is never defined. If it is intended as a string length in some formal language, that language must be specified; if it is an informal notion of amount of assumptions, the claim of precision is strained.
- [Throughout] There are several typographical errors, including "Goodnam" in Section 1.2, "constrints" in Section 2, "explicitely" and "alghough" in Section 2.1, "Gardenfor's" in footnote 9, and "discipleines" in Section 5. These should be corrected.
- [Figure 3] The right panel of Fig. 3 would be more informative if the axes were labeled explicitly (e.g., wavelength and time), so that the reader can see exactly how the error-box splits in the grue/bleen representation.
- [Section 4.2] The discussion of knowledge-what is interesting but seems only loosely connected to the formal definitions in Section 2; the author should state explicitly whether the irreducibility of knowledge-what is supposed to justify Postulate 1 or is merely a philosophical aside.
Circularity Check
The green/grue asymmetry is stipulated by Postulate 1's definition of direct measurements, making the central 'prediction' true by construction.
-
self definitional
[Section 2.1, Postulate 1 and following paragraphs (around Fig. 3)]
"Postulate 1 The result of a valid single direct measurement of (k-dimensional) property Q is always expressed as a (k-dimensional) central value Q0 and a (k-dimensional) convex set (error-box) that contains Q0. ... Note that I have not defined direct measurement explicitely. Post. 1—together with Def. 1 below—offers an implicit (partial) definition of direct measurements. ... Split error-boxes are incompatible with a direct measurement, although they are fully acceptable for indirect measurement."
The paper's central claim is that grue is not directly measurable and therefore more complex than green. But 'direct measurement' is not an independent concept; it is stipulated by Post.1 to be convex. The grue error-box (Fig. 3) is non-convex, so 'grue is not directly measurable' is a direct substitution into the definition. The paper even concedes that Post.1 is an implicit (partial) definition. Consequently the asymmetry the solution relies on is built into the premise: a hypothetical grue-meter whose single-readout distribution is bimodal is rejected for no reason other than the postulate. The later complexity argument (Def.3) then inherits this stipulated asymmetry rather than discovering it.
full rationale
Most of the paper is an explicit axiomatic construction rather than a hidden circularity: Post.1 (convexity of direct measurements) and Def.3 (minimal length over empirically equivalent formulations) are stated as postulates/definitions, and the applications (grue, Bielefeld conspiracy) are worked out from them. However, the central asymmetry is partly stipulated. Post.1 is acknowledged to be an implicit (partial) definition of 'direct measurement,' and the paper then uses the non-convexity of the grue error-box to conclude that grue is not directly measurable. That conclusion is a direct consequence of the definition: any proposed bimodal single-readout 'grue-meter' is inadmissible by fiat. Hence the claimed reduction of Goodman's riddle to complexity is conditional on a premise that already encodes the desired distinction. A second load-bearing assertion is Def.3's non-triviality: the claim that restricting to empirically equivalent formulations makes the minimum length 'in general, not trivial anymore' is asserted without proof; if the minimum collapsed to a constant for all models, the complexity gap would disappear. This is an omitted proof rather than a circular step per se. The paper's appeals to Scorzato (2013) for the framework and for the absence of counterexamples are self-citations, but they support applications rather than the core derivation. Overall, the solution has independent, testable content against scientific-consensus examples, so the circularity is partial, not total.
Assumptions & free parameters
free parameters (3)
- Length function in Definition 3 =
unspecified
- Precision Delta(b) for the direct measurement basis B =
unspecified
- Metric d for discrete measurements =
unspecified
assumptions (5)
- domain assumption A valid single direct measurement always yields a convex error region (Postulate 1/1').
- ad hoc to paper The minimum in Definition 3 over all logically and empirically equivalent formulations is non-trivial.
- domain assumption For any two models on the same topic, a common sub-model containing all directly measurable quantities exists.
- domain assumption Broad scientific consensus is a reliable benchmark for model selection and for testing the proposed model.
- domain assumption The set of measurable quantities can be captured by a finite basis B of directly measurable concepts.
invented entities (1)
-
Epistemic complexity C(M)
Cite this review
Pith. "Pith review of A not-too-simple solution to Goodman's new riddle of induction in the age of AI." pith.science (2026). https://pith.science/paper/JCNZOEZM
@misc{pith2026250717212,
author = {Pith},
title = {Pith review of: A not-too-simple solution to Goodman's new riddle of induction in the age of AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/JCNZOEZM}},
note = {Machine review of arXiv:2507.17212}
}
read the original abstract
I review the works of G\"ardenfors (1990) and Scorzato (2013) and show that their combination provides an elegant solution of Goodman's new riddle of induction. The solution is based on two main ideas: (1) clarifying what is expected from a solution: understanding that philosophy of science is a science itself, with the same limitations and strengths as other scientific disciplines; (2) understanding that the concept of complexity of a model's assumptions and the concept of direct measurements must be characterized together. Although both measurements and complexity have been the subject of a vast literature, within the philosophy of science, essentially no other attempt has been made to combine them. The widespread expectation, among modern philosophers, that Goodman's new riddle cannot be solved is not defensible without a serious exploration of such a natural approach. A clarification of this riddle has always been very important, but it has become even more crucial in the age of AI.
Reference graph
Works this paper leans on
-
[1]
Akaike, H. (1973). Information Theory as an Extension of the Maximum Likelihood Principle . In B. Petrov and F. Csaki (Eds.), Second International Symposium on Information Theory , pp.\ 267--281. Budapest: Akademiai Kiado
work page 1973
-
[2]
Barnett, L. (1950, Jan). The Meaning of Einstein's New Theory -- Interview of A. Einstein . Life Magazine\/ 28 , 22
work page 1950
-
[3]
Behroozi, M. (2022). Largest inscribed rectangles in geometric convex sets
work page 2022
-
[4]
Beisbart, C. (2021). Opacity thought through: on the intransparency of computer simulations . Synthese\/ 199\/ (3-4), 11643--11666
work page 2021
-
[5]
Bird, A. and E. Tobin (2008). Natural Kinds . In E. N. Zalta (Ed.), Stanford Encyclopedia of Philosophy\/ (Spring 2024 Edition ed.). Stanford University
work page 2008
-
[6]
Brzovi\' c , Z. (2014). Natural Kinds . The Internet Encyclopedia of Philosophy\/ ISSN-2161-0002 , 1
work page 2014
-
[7]
Carnap, R. (1950). The Logical Foundations of Probability . Chicago: University of Chicago Press
work page 1950
-
[8]
Carnap, R. (1966). Der Logische Aufbau der Welt \/ (3rd ed.). Hamburg, Germany: Felix Meiner
work page 1966
Show all 59 references
-
[9]
Chaitin, G. J. (1975). Randomness and mathematical proof . Scientific American\/ 232\/ (5), 47--53
1975
-
[10]
Choi, S. and M. Fara (2021). Dispositions . In E. N. Zalta (Ed.), The Stanford Encyclopedia of Philosophy\/ ( S pring 2021 ed.). Metaphysics Research Lab, Stanford University
2021
-
[11]
Cohnitz, D. and M. Rossberg (2024). Nelson Goodman . In E. N. Zalta and U. Nodelman (Eds.), The Stanford Encyclopedia of Philosophy\/ ( S pring 2024 ed.). Metaphysics Research Lab, Stanford University
2024
-
[12]
Crupi, V. (2021). Confirmation . In E. N. Zalta (Ed.), The Stanford Encyclopedia of Philosophy\/ ( S pring 2021 ed.). Metaphysics Research Lab, Stanford University
2021
-
[13]
Duhem, P. M. M. (1954). The Aim and Structure of Physical Theory . Princeton: Princeton University Press
1954
-
[14]
Elgin, C. Z. (Ed.) (1997). The Philosophy of Nelson Goodman: Selected Essays . New York: Garland
1997
-
[15]
Feigl, H. (1970). The ``Orthodox'' View of Theories: Remarks in Defense as well as Critique . In M. Radner and S. Winokur (Eds.), Minnesota Studies in the Philosophy of Science , Volume 4, pp.\ 3--16. University of Minnesota Press
1970
-
[16]
Feynman, R. P., R. B. Leighton, and M. L. Sands (1963). The Feynman lectures on physics; New millennium ed. New York, NY: Addison-Wesley Pub. Co
1963
-
[17]
Fletcher, S. C. (2016). Similarity, topology, and physical significance in relativity theory . The British Journal for the Philosophy of Science\/ 67\/ (2), 365--389
2016
-
[18]
Fletcher, S. C. (2024). On the Alleged Incommensurability of Newtonian and Relativistic Mass . Erkenntnis\/ , 1--22
2024
-
[19]
G\" a rdenfors, P. (1990). Induction, Conceptual Spaces and AI . Philosophy of Science\/ 57\/ (1), 78–95
1990
-
[20]
G \"a rdenfors, P. (2000). Conceptual spaces: the geometry of thought . A Bradford book. MIT Press
2000
-
[21]
G \"a rdenfors, P. (2019). Convexity Is an Empirical Law in the Theory of Conceptual Spaces: Reply to Hern \'a ndez-Conde . In M. Kaipainen, F. Zenker, A. Hautamäki, and P. Gärdenfors (Eds.), Conceptual Spaces: Elaborations and Applications , Volume 405, pp.\ 77. Springer
2019
-
[22]
G\" a rdenfors, P. and A. Stephens (2017). Induction and Knowledge-What . European Journal for Philosophy of Science\/ 8\/ (3), 1--21
2017
-
[23]
Goodman, N. (1946). A Query on Confirmation . Journal of Philosophy\/ 43 , 383--385
1946
-
[24]
Goodman, N. (1955). Fact, Fiction, and Forecast \/ (2nd ed.). Cambridge, MA: Harvard University Press
1955
-
[25]
Goodman, N. (1983). Fact, Fiction, and Forecast \/ (4th ed.). Cambridge, MA: Harvard University Press
1983
-
[26]
Goodman, S. (2008). A Dirty Dozen: Twelve P-Value Misconceptions . Seminars in Hematology\/ 45\/ (3), 135--140. Interpretation of Quantitative Research
2008
-
[27]
Hern\' a ndez - Conde, J. V. (2017). A Case Against Convexity in Conceptual Spaces . Synthese\/ 194\/ (10), 4011--4037
2017
-
[28]
Kelly, K. T. (2007). Ockham’s Razor, Empirical Complexity, and Truth-finding Efficiency . Theoretical Computer Science\/ 383 , 270--289
2007
-
[29]
Khamsi, M. and W. Kirk (2001). An Introduction to Metric Spaces and Fixed Point Theory . John Wiley & Sons, Ltd
2001
-
[30]
Kolmogorov, A. N. (1965). Three Approaches to the Quantitative Definition of Information . Problems Inform. Transmission\/ 1 , 1--7
1965
-
[31]
Kraj cek, J. (2004). Proof complexity . In European congress of mathematics (ECM), Stockholm, Sweden , pp.\ 221--231
2004
-
[32]
Leitgeb, H. (2007). A new analysis of quasianalysis . Journal of Philosophical Logic\/ 36 , 181--226
2007
-
[33]
Leitgeb, H. (2024). Vindicating the Verifiability Criterion . Philosophical Studies\/ 181\/ (1), 223--245
2024
-
[34]
Mormann, T. (1995). Incompatible empirically equivalent theories: A structural explication . Synthese\/ 103 , 203--249
1995
-
[35]
SI definition of Meter
NIST (2019). SI definition of Meter . https://www.nist.gov/si-redefinition/meter
2019
-
[36]
Oberheim, E. and P. Hoyningen-Huene (2025). The Incommensurability of Scientific Theories . In E. N. Zalta and U. Nodelman (Eds.), The Stanford Encyclopedia of Philosophy\/ (Spring 2025 ed.). Stanford University
2025
-
[37]
Piattelli - Palmarini, M. (1980). Language and Learning: The Debate Between Jean Piaget and Noam Chomsky . Harvard University Press
1980
-
[38]
Quine, W. v. O. (1950). Two Dogmas of Empiricism . The Philosophical Review\/ 60 , 20--43
1950
-
[39]
Quine, W. v. O. (1969). Ontological Relativity and Other Essays . New York: Columbia University Press
1969
-
[40]
Quine, W. v. O. (1975). On Empirically Equivalent Systems of the World . Erkenntnis\/ 9 , 313
1975
-
[41]
Quine, W. v. O. (1991). Two Dogmas in Retrospect . Canadian Journal of Philosophy\/ 21\/ (3), 265--274
1991
-
[42]
Scholz, S. (2024). Conceptual Spaces: A Solution to Goodman's New Riddle of Induction? Philosophia\/ 52\/ (4), 915--934
2024
-
[43]
Schwarz, G. (1978). Estimating the Dimension of a Model . Annals of Statistics\/ 4 , 461--464
1978
-
[44]
Scorzato, L. (2013). On the role of simplicity in science . Synthese\/ 190 , 2867--2895
2013
-
[45]
Scorzato, L. (2015). Science and Illusions . preprint: philsci-archive.pitt.edu/15570
2015
-
[46]
Scorzato, L. (2016). A simple model of scientific progress . In L. Felline, F. Paoli, and E. Rossanese (Eds.), New Developments in Logic and Philosophy of Science , Volume 3 of SILFS . College Publications
2016
-
[47]
Scorzato, L. (2024). Reliability and Interpretability in Science and Deep Learning . Minds and Machines\/ 34\/ (3), 27
2024
-
[48]
Sprenger, J. and S. Hartmann (2019). Bayesian Philosophy of Science: Variations on a Theme by the Reverend Thomas Bayes . Oxford and New York: Oxford University Press
2019
-
[49]
Stalker, D. F. (Ed.) (1994). Grue!: The New Riddle of Induction . Chicago and La Salle, IL: Open Court
1994
-
[50]
Stanford, K. (2021). Underdetermination of Scientific Theory . In E. N. Zalta (Ed.), The Stanford Encyclopedia of Philosophy\/ ( W inter 2021 ed.). Metaphysics Research Lab, Stanford University
2021
-
[51]
Starr, W. (2022). Counterfactuals . In E. N. Zalta and U. Nodelman (Eds.), The Stanford Encyclopedia of Philosophy\/ ( W inter 2022 ed.). Metaphysics Research Lab, Stanford University
2022
-
[52]
Str\" o s s ner, C. (2022). Criteria for Naturalness in Conceptual Spaces . Synthese\/ 200\/ (2), 1--36
2022
-
[53]
Teller, P. (1969). Goodman's Theory of Projection . British Journal for the Philosophy of Science\/ 20\/ (3), 219--238
1969
-
[54]
Visser, A. (1991). The Formalization of Interpretability . Studia Logica\/ 50\/ (1), 81--105
1991
-
[55]
Visser, A. (2004). Categories of theories and interpretations . Logic Group Preprint Series\/ 228 , 1--64
2004
-
[56]
Votsis, I. (2016). Philosophy of Science and Information . In L. Floridi (Ed.), The Routledge Handbook of Philosophy of Information . Routledge
2016
-
[57]
Bielefeld conspiracy --- W ikipedia , the free encyclopedia
Wikipedia (2025). Bielefeld conspiracy --- W ikipedia , the free encyclopedia. https://en.wikipedia.org/wiki/Bielefeld_conspiracy. [Online; accessed 05-May-2025]
2025
-
[58]
Williamson, T. and J. Stanley (2001). Knowing how. Journal of Philosophy\/ 98\/ (8), 411--444
2001
-
[59]
Zenil, H. (2020). A Review of Methods for Estimating Algorithmic Complexity: Options, Challenges, and New Directions . Entropy\/ 22\/ (6), 1--28
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.