{"id":"54e9b2db-a16b-4a33-a1ea-77968e179957","arxiv_id":"2412.02829","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A reply arguing that model inclusion alone does not imply overfitting, so overfitting can remain a valid criterion against structurally radical superdeterministic models.","lead":"This reply defends an earlier experiment comparing causal explanations of Bell inequality violations. It argues that a more flexible model does not automatically overfit the data, so overfitting can still count as evidence against some superdeterministic models.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reply's empirical counterexample rests on a null result with no reported uncertainty: 'essentially the same' errors cannot establish absence of overfitting.","rationale":"The reply makes a sound conceptual distinction: strict model inclusion does not force overfitting, because overfitting occurs when statistical fluctuations are mistaken for real features. I am not objecting to that logical point. The crucial empirical demonstration, however, is the qCC versus cCC comparison in the dephased experiment. Reporting that the test errors are 'essentially the same' is a null result; without confidence intervals or a pre-specified equivalence margin, it does not establish the absence of overfitting. The same concern applies, though less centrally, to the cSD0 and cCE0 overfitting claims, which are cited only from the prior article. Since this empirical counterexample is the load-bearing support for applying the overfitting criterion to the class of superdeterministic models at issue, and the reply does not include the relevant numerical evidence, the conditional verdict is appropriate. I do not see an internal inconsistency in the logical argument, and the reply's explicit caveat about parameter-restricted models honestly limits the scope of its conclusions. The proposed check—recomputing the paired test-error difference with uncertainty—would settle whether the example genuinely demonstrates absence of overfitting or merely non-detection at the achieved statistical precision.","tokens_in":632,"tokens_out":2926,"duration_ms":92723,"concrete_test":"Obtain the full numerical output from Appendix C.3 and Fig. 6 of Ref. [1], including per-fold training and test errors for all four models. Compute the paired difference in test error between qCC and cCC, with bootstrap or per-fold confidence intervals. If the 95% confidence interval excludes zero with qCC worse, the counterexample fails. If the interval is wide enough to include a practically relevant overfitting effect (for example, a relative difference larger than 10%), the example is inconclusive; a new experiment with more data, or an analytic construction of a model-inclusion pair with provably equal generalization error, would be required to support the reply's central point. Independently, re-run the original analysis code to confirm that cSD0 and cCE0 overfit qCC in the Bell experiment.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing assertion is that in the dephased experiment the qCC model does not overfit the cCC model despite qCC being strictly more expressive, thereby demonstrating model inclusion without overfitting. But this is an empirical null result, and the reply only reports that the training and test errors are 'essentially the same', citing Fig. 6 and Appendix C.3 of Ref. [1] without reproducing the numbers or any uncertainty quantification. If the qCC test error is merely not significantly worse than that of cCC, this could be due to insufficient data rather than to the extra expressive power being innocuous. The logical point—that model inclusion does not logically entail overfitting—is well made, but the empirical example is the only support offered for the specific mechanism, and it is under-specified. Moreover, the reply's own standard says that overfitting is diagnosed by mistaking statistical fluctuations for real features; here the absence of a detected difference could itself be a statistical fluctuation. Thus the counterexample does not yet establish that the extra expressive power of qCC over cCC is innocuous in the realized experiment; it only establishes that any overfitting, if present, was not detected at the achieved precision. The companion claims that cSD0 and cCE0 overfit qCC in the Bell experiment are likewise not independently reproducible from this reply alone. Finally, the reply itself concedes that parameter-restricted superdeterministic models might not overfit, so the scope of the conclusion is narrower than a general refutation of Hance and Hossenfelder's position.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This reply to Hance and Hossenfelder's comment defends the original paper's terminology and its use of overfitting as an adjudication criterion in causal model selection. The authors dispute HH's claim that they misrepresented superdeterministic models, clarifying that 'classical' referred specifically to the framework of causal modelling, not to classical mechanics, and that HH's own preferred superdeterministic models fall into the cSD class considered in the original work. They then address HH's argument that model inclusion necessarily leads to overfitting. Their central logical point is that strict model inclusion does not imply overfitting: overfitting occurs when a model mistakes statistical fluctuations for real features. As a concrete example, they cite a dephased Bell experiment from their prior paper in which the more expressive qCC model did not overfit the less expressive cCC model despite standing in a model inclusion relation. The reply also discusses mechanisms that can cause overfitting under model inclusion, distinguishing equality constraints such as no-signalling from inequality constraints such as Bell inequalities, and argues that incorporating additional experimental details can prevent a reductio against the train-and-test methodology.","tokens_in":10719,"tokens_out":6951,"duration_ms":73637,"significance":"If the argument holds, the paper makes a useful conceptual clarification: model inclusion is not sufficient for overfitting, and a train-and-test diagnosis can legitimately count against a class of superdeterministic models when the overfitting arises from a specific statistical mechanism. The reply's logical point is sound and does not depend on the empirical example. It also provides a valuable distinction between equality and inequality constraints as sources of overfitting, and it is candid that parameter-restricted superdeterministic models remain untested. The main weakness is that the decisive empirical example of model inclusion without overfitting is not self-contained and is reported without uncertainty quantification. The paper does not provide new code or data, and relies on the authors' own prior work for its central counterexample, which should be stated more transparently.","major_comments":[{"comment":"The central empirical claim that qCC does not overfit cCC in the dephased experiment is supported only by the statement that the qCC model had training and test errors 'essentially the same' as those of the cCC model, citing Fig. 6 and Appendix C.3 of Ref. [1]. No numerical values, uncertainties, or statistical test are reproduced in this reply. Because this example is the paper's only concrete demonstration of model inclusion without overfitting, the absence of a detected difference could simply reflect insufficient statistical power, which would not establish that the extra expressive power is innocuous. The authors should either reproduce the relevant error values with uncertainties and a significance test, or explicitly present the example as an illustration rather than a demonstrated fact.","section":"Section III and Appendix A1"},{"comment":"The conclusion that qCC's extra expressive power is 'of the innocuous variety' is based on the assertion that 'we do not see any other avenues for mistaking statistical fluctuations for real features' in the qCC-versus-cCC comparison. This is a non-exhaustive argument. The authors identify the no-signalling equality constraint as one mechanism and argue that Bell-inequality violations are unlikely in the dephased experiment, but they do not provide a complete model-based enumeration of possible overfitting mechanisms. A stronger statement would be either to give a principled argument that no other mechanism exists for this pair of models, or to weaken the conclusion to 'no overfitting was observed in the realized experiment.'","section":"Appendix A1"},{"comment":"The reply argues that including more experimental detail can prevent a more expressive model from overfitting, and that this avoids a reductio of the train-and-test methodology. However, this makes the overfitting verdict dependent on the modeller's choice of how much experimental detail to include. The authors do not provide a criterion for selecting the appropriate level of detail. If a proponent of a more expressive model can always avoid overfitting by coarse-graining the description of the preparation, then overfitting is not a property of the theory alone but of the model-experiment pairing. The reply should state explicitly that overfitting is defined relative to a chosen level of experimental description and should discuss how comparisons should be standardized.","section":"Appendix A2"}],"minor_comments":[{"comment":"The heading contains a typo: 'SUPDETERMINISTIC' should be 'SUPERDETERMINISTIC'.","section":"Section II heading"},{"comment":"The word 'stiuplate' should be 'stipulate'.","section":"Section II"},{"comment":"There is a duplicated word: 'strictly more expressive than than the cCC model' should read 'strictly more expressive than the cCC model'.","section":"Appendix A1"},{"comment":"The phrase 'parameteric downconversion' should be 'parametric downconversion'.","section":"Appendix A2"},{"comment":"Reference [2] is listed as 'to be published'; an updated citation with volume or DOI should be provided if available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reply to a comment, and its central logical point is defensible. The main obstacle is the lack of quantitative support for the decisive empirical example. If the authors can supply the exact error values, uncertainties, and a statistical comparison from Ref. [1], or clearly downgrade the status of the example, the paper would be suitable for publication. The reliance on the authors' own prior work is acceptable in a reply, but the current level of detail makes the central claim difficult to verify."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jon,\n\nQuick take: the reply's core logical point is right. HH's claim that more expressive power automatically means overfitting is simply false; overfitting is about mistaking statistical fluctuations for real effects. The paper does a good job explaining that, and the clarification of what 'classical' means in their causal-model framework is fair and addresses a real misunderstanding.\n\nThe genuinely new bit is the second appendix section, where they point out that equality constraints (like no-signalling) and inequality constraints (like Bell inequalities) offer different overfitting avenues. That distinction is useful and not something I recall being explicit before. The hypothetical example where qCC overfits cCC near a Bell-inequality boundary is a nice thought experiment.\n\nThe soft spot is the empirical example meant to make the point concrete. They say that in the dephased experiment qCC and cCC had 'essentially the same' training and test errors, so model inclusion did not cause overfitting. But there is no uncertainty quantification, no numbers, no statistical test for equivalence. If the test error of qCC is slightly worse but within noise, that doesn't demonstrate absence of overfitting; it just means they didn't detect any. That is a real gap, and the stress-test note is right about it. The data and code from the original paper should have been referenced more explicitly, or the error bars reproduced.\n\nAlso, the reply itself concedes that parameter-restricted superdeterministic models might not overfit, so the refutation only targets the parameter-unrestricted class. That is honest, but it means they are not claiming to have settled the whole dispute.\n\nOverall: the logical argument holds, the clarifications are useful, and the novel distinction about equality vs inequality constraints is a genuine contribution. The empirical support is thinner than the prose suggests, but that is fixable in a revised version. I would send this to peer review. The authors should be asked to supply error bars or a statistical comparison for the qCC vs cCC errors, and ideally make the prior analysis more easily accessible. It is a serious reply to a serious comment, and it deserves referee time.\n\nBest,\n[Your name]","headline":"A logically sound reply whose central point about model inclusion and overfitting holds, but its empirical illustration is under-specified and narrower than the authors' framing suggests.","tokens_in":11255,"tokens_out":1906,"would_cite":false,"duration_ms":19656,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Model inclusion does not cause overfitting: the reply argues overfitting requires mistaking statistical fluctuations for real features, so train-and-test verdicts against superdeterministic models stand.","keywords":["superdeterminism","Bell inequalities","overfitting","model selection","causal models","train-and-test","no-signalling","measurement independence"],"falsifier":"Recompute the train-and-test errors for qCC versus cCC on the raw data of the dephased experiment: if qCC shows a strictly larger test error than cCC while its training error is at least as small, the paper's central example of model inclusion without overfitting fails; if the errors agree within statistical uncertainty, the example stands.","tokens_in":10291,"feed_emoji":"⚛️","tokens_out":7626,"duration_ms":65127,"temperature":0.7,"pith_summary":"This reply defends train-and-test model selection as a fair way to adjudicate between causal explanations of Bell-inequality violations. Against the objection that overfitting merely reflects model inclusion, the paper argues that a strictly more expressive model $M$ need not overfit a less expressive model $M'$; overfitting occurs specifically when $M$ mistakes statistical fluctuations for real features. The decisive evidence is a dephased Bell experiment in which the quantum common-cause model (qCC) is strictly more expressive than the classical common-cause model (cCC) yet shows essentially identical training and test errors, while structurally radical models that can violate no-signalling do overfit. If the reply is correct, the original experimental verdict that disfavours a class of superdeterministic models remains intact, and the same diagnostic can be turned on other beyond-quantum proposals.","feed_headline":"Model inclusion does not cause overfitting","feed_subtitle":"Train-and-test comparisons remain a fair test of superdeterministic models of Bell violations.","key_machinery":"The mechanism that carries the argument is the train-and-test methodology together with a distinction between equality constraints and inequality constraints on the set of correlations a model can produce. No-signalling is an equality constraint: classically and quantumly common-cause models satisfy it for all parameter values, whereas structurally radical models include parameter values that violate it, leaving them vulnerable to mistaking finite-run fluctuations away from no-signalling for real features. Bell inequalities are inequality constraints, and the reply notes that a more expressive model can also overfit when the data lie within statistical error of the boundary between the two models' achievable sets. The dephased Bell experiment is the load-bearing example: qCC is strictly more expressive than cCC, both satisfy the same equality constraints, and the data are far from the Bell-inequality boundary, so the extra expressive power is innocuous and no overfitting occurs.","core_discovery":"The central claim is that overfitting in a train-and-test comparison is not a generic consequence of model inclusion. A model $M$ that reproduces all the operational statistics of $M'$ and more besides will overfit $M'$ only when its extra expressive power gives it a way to mistake finite-run statistical fluctuations for real features. The paper exhibits a concrete case: in a dephased Bell experiment, qCC and cCC stand in a model-inclusion relation, yet qCC's training and test errors are essentially the same as cCC's, so no overfitting appears. In contrast, the structurally radical models cCE0 and cSD0, which allow parameter values that violate the no-signalling condition, do overfit because they can interpret fluctuations away from no-signalling as real effects. The reply concludes that the original finding—that certain superdeterministic models overfit relative to qCC—is a valid reason to disfavour them.","pith_inferences":["If this argument is right, the debate over superdeterminism should shift from whether overfitting is a fair criterion to which parameter restrictions are physically motivated: the reply explicitly leaves open that a parameter-restricted, quantum-extending superdeterministic model could outperform the quantum common-cause model.","The equality-versus-inequality constraint distinction suggests a general diagnostic: when comparing nested model classes, overfitting risk is concentrated where the larger class relaxes an equality constraint such as no-signalling, rather than merely extending an inequality bound.","A direct testable extension is to apply the same train-and-test analysis to other radical causal structures—retrocausal or signalling models—and check whether they exhibit the same overfitting signature on finite-run Bell data.","The reply's logic also implies that a proponent of any beyond-quantum theory can deflect an overfitting verdict by supplying a parameter prior informed by experimental details; the burden is on the proponent to supply that prior."],"forward_implications":["Train-and-test model selection remains a legitimate arbiter between structurally radical and structurally conservative causal explanations of Bell violations; being more expressive does not by itself protect a superdeterministic model from an overfitting verdict.","The original experiment's disfavouring of parameter-unrestricted superdeterministic models (cSD0) stands unless proponents articulate a specific parameter restriction and show that it removes the overfitting vulnerability.","Overfitting can appear in a model-inclusion relation when the data lie within statistical error of an equality constraint (like no-signalling) or an inequality constraint (like a Bell inequality) that the more expressive model can violate.","Feeding additional details of the experimental procedure into the fitting model can prevent a more expressive model from mistaking fluctuations for real features, so the possibility of overfitting is not a reductio against train-and-test methodology.","The same methodology applies to other beyond-quantum theories, such as those permitting violations of the Tsirelson bound, with the same caveat that a detailed prior can rescue the model."],"supporting_citations":[{"why":"Supplies the original experiment, the train-and-test data analysis, Fig. 6, and the dephased-experiment results that the reply's central counterexample relies on.","marker":"[1]"},{"why":"The comment being rebutted; its claim that model inclusion necessarily implies overfitting is the target of the reply.","marker":"[2]"},{"why":"Defines conditional density operators, providing the framework for the intrinsically quantum causal model qCC whose comparison with cCC is the key example.","marker":"[5]"},{"why":"Shows classical causal models are hidden variable models and introduces faithfulness, which the reply connects to the overfitting mechanism.","marker":"[6]"},{"why":"Supports the argument that precision-frontier experiments can detect small deviations from operational quantum theory, countering the comment's claim about necessary limits.","marker":"[7]"},{"why":"Supplies the framework of equality and inequality constraints on causal-model correlations, used to distinguish no-signalling from Bell-inequality constraints.","marker":"[8]"},{"why":"Provides the example of a post-quantum theory violating the Tsirelson bound, used to illustrate how overfitting can arise near an inequality-constraint boundary.","marker":"[12]"}],"fun_headline_variants":["Nested models need not overfit: Bell test stands","Overfitting not generic when one model includes another","Superdeterministic overfitting hinges on misreading noise","Train-test still rules out some superdeterministic models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reply's decisive counterexample rests on the correctness of the original train-and-test analysis reported as Fig. 6 and Appendix C.3 of the original article: if that analysis contains an error, or if its model classes do not represent the superdeterministic models the comment defends, the reply's refutation loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["Nested models need not overfit: Bell test stands","Overfitting not generic when one model includes another","Superdeterministic overfitting hinges on misreading noise","Train-test still rules out some superdeterministic models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00045,"raw_usage":{"total_tokens":2325,"prompt_tokens":1060,"completion_tokens":1265,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":1200}},"tokens_in":676,"tokens_out":1265,"duration_ms":10625,"temperature":1.0,"reasoning_tokens":1200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:03:33.884086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the train-and-test errors for qCC versus cCC on the raw data of the dephased experiment: if qCC shows a strictly larger test error than cCC while its training error is at least as small, the paper's central example of model inclusion without overfitting fails; if the errors agree within statistical uncertainty, the example stands.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original experiment, the train-and-test data analysis, Fig. 6, and the dephased-experiment results that the reply's central counterexample relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The comment being rebutted; its claim that model inclusion necessarily implies overfitting is the target of the reply."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines conditional density operators, providing the framework for the intrinsically quantum causal model qCC whose comparison with cCC is the key example."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows classical causal models are hidden variable models and introduces faithfulness, which the reply connects to the overfitting mechanism."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the argument that precision-frontier experiments can detect small deviations from operational quantum theory, countering the comment's claim about necessary limits."},{"cited_title":"Wolfe, R","cited_arxiv_id":null,"evidence_quote":"Supplies the framework of equality and inequality constraints on causal-model correlations, used to distinguish no-signalling from Bell-inequality constraints."},{"cited_title":"Navascu´ es, Y","cited_arxiv_id":null,"evidence_quote":"Provides the example of a post-quantum theory violating the Tsirelson bound, used to illustrate how overfitting can arise near an inequality-constraint boundary."}],"review_version":1}