REVIEW 3 major objections 5 minor 1 cited by
Reply to "Comment on 'Experimentally adjudicating between different causal accounts of Bell-inequality violations via statistical model selection'"
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Model inclusion does not cause overfitting: the reply argues overfitting requires mistaking statistical fluctuations for real features, so train-and-test verdicts against superdeterministic models stand.
desk verdict A logically sound reply whose central point about model inclusion and overfitting holds, but its empirical illustration is under-specified and narrower than the authors' framing suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is the train-and-test methodology together with a distinction between equality constraints and inequality constraints on the set of correlations a model can produce. No-signalling is an equality constraint: classically and quantumly common-cause models satisfy it for all parameter values, whereas structurally radical models include parameter values that violate it, leaving them vulnerable to mistaking finite-run fluctuations away from no-signalling for real features. Bell inequalities are inequality constraints, and the reply notes that a more expressive model can also overfit when the data lie within statistical error of the boundary between the two models' achievable sets. The dephased Bell experiment is the load-bearing example: qCC is strictly more expressive than cCC, both satisfy the same equality constraints, and the data are far from the Bell-inequality boundary, so the extra expressive power is innocuous and no overfitting occurs.
What would settle it
Recompute the train-and-test errors for qCC versus cCC on the raw data of the dephased experiment: if qCC shows a strictly larger test error than cCC while its training error is at least as small, the paper's central example of model inclusion without overfitting fails; if the errors agree within statistical uncertainty, the example stands.
Extended reading notes
Core claim
The central claim is that overfitting in a train-and-test comparison is not a generic consequence of model inclusion. A model $M$ that reproduces all the operational statistics of $M'$ and more besides will overfit $M'$ only when its extra expressive power gives it a way to mistake finite-run statistical fluctuations for real features. The paper exhibits a concrete case: in a dephased Bell experiment, qCC and cCC stand in a model-inclusion relation, yet qCC's training and test errors are essentially the same as cCC's, so no overfitting appears. In contrast, the structurally radical models cCE0 and cSD0, which allow parameter values that violate the no-signalling condition, do overfit because they can interpret fluctuations away from no-signalling as real effects. The reply concludes that the original finding—that certain superdeterministic models overfit relative to qCC—is a valid reason to disfavour them.
Load-bearing premise
The reply's decisive counterexample rests on the correctness of the original train-and-test analysis reported as Fig. 6 and Appendix C.3 of the original article: if that analysis contains an error, or if its model classes do not represent the superdeterministic models the comment defends, the reply's refutation loses its footing.
Editorial extensions
If this is right
- Train-and-test model selection remains a legitimate arbiter between structurally radical and structurally conservative causal explanations of Bell violations; being more expressive does not by itself protect a superdeterministic model from an overfitting verdict.
- The original experiment's disfavouring of parameter-unrestricted superdeterministic models (cSD0) stands unless proponents articulate a specific parameter restriction and show that it removes the overfitting vulnerability.
- Overfitting can appear in a model-inclusion relation when the data lie within statistical error of an equality constraint (like no-signalling) or an inequality constraint (like a Bell inequality) that the more expressive model can violate.
- Feeding additional details of the experimental procedure into the fitting model can prevent a more expressive model from mistaking fluctuations for real features, so the possibility of overfitting is not a reductio against train-and-test methodology.
- The same methodology applies to other beyond-quantum theories, such as those permitting violations of the Tsirelson bound, with the same caveat that a detailed prior can rescue the model.
Reading between the lines
- If this argument is right, the debate over superdeterminism should shift from whether overfitting is a fair criterion to which parameter restrictions are physically motivated: the reply explicitly leaves open that a parameter-restricted, quantum-extending superdeterministic model could outperform the quantum common-cause model.
- The equality-versus-inequality constraint distinction suggests a general diagnostic: when comparing nested model classes, overfitting risk is concentrated where the larger class relaxes an equality constraint such as no-signalling, rather than merely extending an inequality bound.
- A direct testable extension is to apply the same train-and-test analysis to other radical causal structures—retrocausal or signalling models—and check whether they exhibit the same overfitting signature on finite-run Bell data.
- The reply's logic also implies that a proponent of any beyond-quantum theory can deflect an overfitting verdict by supplying a parameter prior informed by experimental details; the burden is on the proponent to supply that prior.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This reply to Hance and Hossenfelder's comment defends the original paper's terminology and its use of overfitting as an adjudication criterion in causal model selection. The authors dispute HH's claim that they misrepresented superdeterministic models, clarifying that 'classical' referred specifically to the framework of causal modelling, not to classical mechanics, and that HH's own preferred superdeterministic models fall into the cSD class considered in the original work. They then address HH's argument that model inclusion necessarily leads to overfitting. Their central logical point is that strict model inclusion does not imply overfitting: overfitting occurs when a model mistakes statistical fluctuations for real features. As a concrete example, they cite a dephased Bell experiment from their prior paper in which the more expressive qCC model did not overfit the less expressive cCC model despite standing in a model inclusion relation. The reply also discusses mechanisms that can cause overfitting under model inclusion, distinguishing equality constraints such as no-signalling from inequality constraints such as Bell inequalities, and argues that incorporating additional experimental details can prevent a reductio against the train-and-test methodology.
Significance. If the argument holds, the paper makes a useful conceptual clarification: model inclusion is not sufficient for overfitting, and a train-and-test diagnosis can legitimately count against a class of superdeterministic models when the overfitting arises from a specific statistical mechanism. The reply's logical point is sound and does not depend on the empirical example. It also provides a valuable distinction between equality and inequality constraints as sources of overfitting, and it is candid that parameter-restricted superdeterministic models remain untested. The main weakness is that the decisive empirical example of model inclusion without overfitting is not self-contained and is reported without uncertainty quantification. The paper does not provide new code or data, and relies on the authors' own prior work for its central counterexample, which should be stated more transparently.
major comments (3)
- [Section III and Appendix A1] The central empirical claim that qCC does not overfit cCC in the dephased experiment is supported only by the statement that the qCC model had training and test errors 'essentially the same' as those of the cCC model, citing Fig. 6 and Appendix C.3 of Ref. [1]. No numerical values, uncertainties, or statistical test are reproduced in this reply. Because this example is the paper's only concrete demonstration of model inclusion without overfitting, the absence of a detected difference could simply reflect insufficient statistical power, which would not establish that the extra expressive power is innocuous. The authors should either reproduce the relevant error values with uncertainties and a significance test, or explicitly present the example as an illustration rather than a demonstrated fact.
- [Appendix A1] The conclusion that qCC's extra expressive power is 'of the innocuous variety' is based on the assertion that 'we do not see any other avenues for mistaking statistical fluctuations for real features' in the qCC-versus-cCC comparison. This is a non-exhaustive argument. The authors identify the no-signalling equality constraint as one mechanism and argue that Bell-inequality violations are unlikely in the dephased experiment, but they do not provide a complete model-based enumeration of possible overfitting mechanisms. A stronger statement would be either to give a principled argument that no other mechanism exists for this pair of models, or to weaken the conclusion to 'no overfitting was observed in the realized experiment.'
- [Appendix A2] The reply argues that including more experimental detail can prevent a more expressive model from overfitting, and that this avoids a reductio of the train-and-test methodology. However, this makes the overfitting verdict dependent on the modeller's choice of how much experimental detail to include. The authors do not provide a criterion for selecting the appropriate level of detail. If a proponent of a more expressive model can always avoid overfitting by coarse-graining the description of the preparation, then overfitting is not a property of the theory alone but of the model-experiment pairing. The reply should state explicitly that overfitting is defined relative to a chosen level of experimental description and should discuss how comparisons should be standardized.
minor comments (5)
- [Section II heading] The heading contains a typo: 'SUPDETERMINISTIC' should be 'SUPERDETERMINISTIC'.
- [Section II] The word 'stiuplate' should be 'stipulate'.
- [Appendix A1] There is a duplicated word: 'strictly more expressive than than the cCC model' should read 'strictly more expressive than the cCC model'.
- [Appendix A2] The phrase 'parameteric downconversion' should be 'parametric downconversion'.
- [References] Reference [2] is listed as 'to be published'; an updated citation with volume or DOI should be provided if available.
Circularity Check
No circular derivation: the reply's logical point stands independently, and its use of the authors' prior experiment is self-citation for empirical illustration, not circular reasoning.
full rationale
The central claim of the reply is a general logical thesis: if model M has strictly more expressive power than model M', model inclusion does not by itself imply that M will overfit M' in a train-and-test analysis. This thesis is argued directly from the mechanism of overfitting ('mistaking statistical fluctuations for real features') and from the structural observation that qCC and cCC satisfy the same equality constraints, differing only by inequality constraints. No equation of the reply defines the conclusion in terms of its inputs, and no fitted parameter from the prior analysis is relabeled as a prediction here. The empirical counterexample from the authors' previous experiment is cited rather than re-derived, and the numerical details are not reproduced, but that prior experiment is an independently published, externally checkable data analysis, not a definitional or logical component of the present argument. The absence of uncertainty quantification for the 'essentially the same' training and test errors is a legitimate concern about evidential strength, but it is not circularity. The self-citations to Refs. [1] and [6] support background facts and an illustrative example, while the logical point about model inclusion and overfitting remains self-contained; hence the low score.
Assumptions & free parameters
assumptions (4)
- domain assumption The train-test results and model definitions from the authors' prior article (Ref. [1], especially Fig. 6 and Appendix C.3) are correct and representative.
- domain assumption Overfitting is caused by a model mistaking statistical fluctuations for real features.
- standard math Structurally radical causal models (cSD0, cCE0) can realize violations of the no-signalling condition, while structurally conservative models (cCC, qCC) cannot.
- domain assumption The classical framework for causal modelling is the appropriate setting for characterizing the superdeterministic models discussed in the original paper.
Cite this review
Pith. "Pith review of Reply to "Comment on 'Experimentally adjudicating between different causal accounts of Bell-inequality violations via statistical model selection'"." pith.science (2026). https://pith.science/paper/ERAF7ZPG
@misc{pith2026241202829,
author = {Pith},
title = {Pith review of: Reply to "Comment on 'Experimentally adjudicating between different causal accounts of Bell-inequality violations via statistical model selection'"},
year = {2026},
howpublished = {\url{https://pith.science/paper/ERAF7ZPG}},
note = {Machine review of arXiv:2412.02829}
}
read the original abstract
Our article described an experiment that adjudicates between different causal accounts of Bell inequality violations by a comparison of their predictive power, finding that certain types of models that are structurally radical but parametrically conservative, of which a class of superdeterministic models are an example, overfit the data relative to models that are structurally conservative but parametrically radical in the sense of endorsing an intrinsically quantum generalization of the framework of causal modelling. In their comment (arXiv:2206.10619), Hance and Hossenfelder argue that we have misrepresented the purpose of superdeterministic models. We here dispute this claim by recalling the different classes of superdeterministic models we defined in our article and our conclusions regarding which of these are disfavoured by our experimental results. Their confusion on this point seems to have arisen in part from the fact that we characterized superdeterministic models within a causal modelling framework and from the fact that we referred to this framework as "classical" in order to contrast it with an intrinsically quantum alternative. In this reply, therefore, we take the opportunity to clarify these points. They also claim that if one is adjudicating between a pair of models, where one model can account for strictly more operational statistics than the other, the first model will tend to overfit the data relative to the second. Because this model inclusion relation can arise for pairs of models in a reductionist heirarchy, they conclude that overfitting should not be taken as evidence against the first model. We point out here that, contrary to this claim, one does not expect overfitting to arise generically in cases of model inclusion, so that it is indeed sometimes appropriate to consider overfitting as a criterion for adjudicating between such models.
Forward citations
Cited by 1 Pith paper
-
The Observational Partial Order of Causal Structures with Latent Variables
For causal structures with latent variables and a fixed node order, the paper fully determines the observational dominance order at three visible nodes and bounds it at four nodes, showing that most equivalence classe...
Reference graph
Works this paper leans on
-
[1]
P. J. Daley, K. J. Resch, and R. W. Spekkens, Physical Review A 105, 042220 (2022)
work page 2022
-
[2]
J. R. Hance and S. Hossenfelder, to be published in Phys- ical Review A (2024)
work page 2024
-
[3]
C. A. Hooker, The Logico-Algebraic Approach to Quan- tum Mechanics: Volume I: Historical Evolution , vol. 5 (Springer Science & Business Media, 2012)
work page 2012
- [4]
-
[5]
M. S. Leifer and R. W. Spekkens, Physical Review A 88, 052130 (2013)
work page 2013
-
[6]
C. J. Wood and R. W. Spekkens, New Journal of Physics 5 17, 033002 (2015)
work page 2015
-
[7]
M. D. Mazurek, M. F. Pusey, K. J. Resch, and R. W. Spekkens, PRX Quantum 2, 020302 (2021)
work page 2021
- [8]
Show all 14 references
-
[9]
Valentini, Physics Letters A 156, 5 (1991)
A. Valentini, Physics Letters A 156, 5 (1991)
1991
-
[10]
Valentini, Physics Letters A 158, 1 (1991)
A. Valentini, Physics Letters A 158, 1 (1991)
1991
-
[11]
Valentini, arXiv preprint astro-ph/0412503 (2004)
A. Valentini, arXiv preprint astro-ph/0412503 (2004)
2004 arXiv
-
[12]
Navascu´ es, Y
M. Navascu´ es, Y. Guryanova, M. J. Hoban, and A. Ac ´ ın, Nature communications 6, 6288 (2015). Appendix A: Further comments on the relationship between model inclusion and overfitting
2015
-
[13]
We then fit the data from this experiment to four distinct causal models and ana- lyzed how they each performed relative to the train-and- test methodology
An example of model inclusion without overfitting In Appendix C.3 of our article [1], we described a de- phased version of our experiment, that is, a Bell exper- iment where the entangled state is subject to dephasing and thus becomes a separable state, so that no Bell in- equa...
-
[14]
A novel opportunity for overfitting in a model inclusion setting Nonetheless, it is worth considering the question that this raises: what sorts of circumstances could lead to the situation that a more expressive model overfits the data relative to a less expressive model? To ans...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.