{"id":"9d240360-580b-43b5-8ac5-aa46e3e1eb01","arxiv_id":"2507.14159","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Siamese network trained only on non-critical percolation configurations predicts 3D site and bond percolation thresholds, but its learned representation is essentially the normalized largest-cluster size.","lead":"Using a Siamese neural network trained only on very low and very high occupancy examples, this paper locates 3D percolation thresholds with percent-level accuracy. The approach uses few labels near the transition, but its learned signal coincides with the standard largest-cluster order parameter.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's ν claim is unsupported: §4.3 fixes ν≈0.88 a priori for the data collapse and Table 4 reports ν=0.88 with no uncertainty or fitting, so the method never estimates ν; a free-ν collapse fit is required to substantiate or retract the claim.","rationale":"The reader's weakest_assumption is the p_c intersection criterion, which is a legitimate concern but is partially mitigated by the reported anchor-dependent thresholds: even the most discrepant entries in Tables 1 and 2 are within roughly 1–2% of the literature values, and the finite-size trends point in the correct direction. The ν claim, by contrast, is contradicted by the manuscript's own protocol. Section 4.3 explicitly fixes ν≈0.88, and Table 4 reports identical values with no uncertainty and no fitting procedure, so the paper provides no evidence that the method can estimate ν. This is a precise internal mismatch rather than a disagreement with prior consensus, and it directly affects a headline quantitative claim. The missing FCC transfer experiment is also notable, but it is a missing result rather than a direct contradiction of a described procedure; the ν issue is more decisive because the described procedure shows that the claimed estimate was not actually produced. The reader's rationale does mention the hand-fixed exponent, so there is partial agreement, though the reader's stated weakest assumption is the p_c extraction. Since the reader already assigned a CONDITIONAL verdict, which is the appropriate disposition given this concern, the verdict should remain unchanged: the paper needs revision to either perform a genuine free-ν estimation or visibly soften the abstract's ν claim.","tokens_in":24328,"tokens_out":7849,"duration_ms":89525,"concrete_test":"Re-analyze the data behind Figs. 7 and 8 with ν as a free parameter: for each anchor and L, collapse the SNN output curves onto a common master curve y=f((p−p_c)L^{1/ν}) and minimize a quantitative collapse metric (e.g., bootstrap-averaged squared residuals over ν and p_c together). Report the best-fit ν with its confidence interval for both site and bond percolation. If the optimum lies far from the literature value 0.8765, or if the collapse metric is flat over a range wider than about 10% of ν, the abstract's ν-estimation claim is unsupported and should be replaced by a fixed-ν consistency check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The label-efficiency claim has two quantitative components: p_c and ν. The p_c extraction in §4.2/4.3 is under-formalized, but Tables 1–2 show plausible anchor-dependent values. The ν component is not under-formalized; it is absent. The abstract states that the method 'yields estimates of the critical exponent ν consistent with literature values within statistical uncertainty,' but §4.3 performs the data collapse by setting ν≈0.88: 'when the critical exponent is set to ν≈0.88, the data curves ... collapse.' Table 4 then reports ν=0.88 for every anchor, before and after iteration, with no uncertainty, no goodness-of-fit measure, and no screening over ν. Therefore ν is an input chosen by the authors, not a quantity inferred from the SNN outputs. A single visual collapse at a hand-picked exponent cannot establish agreement 'within statistical uncertainty,' because no statistical procedure connects the curves to ν. If the collapse was intended only as a consistency check at the known literature value, the abstract overstates the result; if it was intended as an estimate, the estimation step is missing. This is load-bearing because the central claim explicitly advertises critical-exponent estimation as part of the label-efficient payoff, and the only evidence offered is a fixed-parameter replot.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a Siamese neural network (SNN) on pairwise similarity labels derived from 3D site and bond percolation configurations, using as input the largest connected cluster extracted by depth-first search. Labels are generated only from configurations in the non-critical intervals [0,0.1] and [0.9,1], giving 22 labeled probability points across five system sizes. The authors propose that the intersection of the SNN's positive and negative similarity curves marks the percolation threshold, and they extract p_c via finite-size scaling. They also present data-collapse plots for the SNN output and report the critical exponent ν=0.88. The central advertised results are percent-level estimates of p_c, ν consistent with literature within statistical uncertainty, and transfer from simple cubic to face-centered cubic lattices without retraining.","tokens_in":24590,"tokens_out":3667,"duration_ms":37819,"significance":"If fully substantiated, a label-efficient SNN that locates 3D percolation thresholds and estimates critical exponents from only 22 labeled probability points would be a useful addition to the ML-for-critical-phenomena toolkit. The paper has several strengths: it uses a large set of independent Monte Carlo configurations per probability point, shares weights between the two branches, explicitly studies the effect of extending the labeling interval via iteration, and reports anchor-based robustness tables. However, the current manuscript does not support the critical-exponent claim or the face-centered-cubic transfer claim, and the p_c extraction procedure is insufficiently formalized. The significance is therefore conditional on substantial additional analysis.","major_comments":[{"comment":"The claim that the method “yields estimates of the critical exponent ν consistent with literature values within statistical uncertainty” is not supported by the reported analysis. In §4.3 the data collapse is performed by setting ν≈0.88 beforehand (“when the critical exponent is set to ν≈0.88, the data curves … collapse”), and Table 4 lists ν=0.88 and ν_it=0.88 for every anchor, with no uncertainty, no goodness-of-fit measure, and no search over ν. The exponent is therefore an input chosen by the authors, not an estimate produced by the SNN. To substantiate or retract the abstract claim, the authors should perform a data-collapse fit in which ν is a free parameter, report the fitted value with uncertainty for each anchor, and state a quantitative collapse criterion.","section":"Abstract; §4.3; Figs. 7-8; Table 4"},{"comment":"The abstract's headline transfer claim — that a network trained solely on simple cubic lattices identifies the phase transition in face-centered cubic lattices without retraining — has no corresponding experimental section, figure, or table anywhere in the main text. No fcc model, fcc simulation, or fcc result is described in Sections 2-5. Either add a complete fcc subsection with simulation details and quantitative results, or remove the claim from the abstract and introduction, since it is currently unsupported.","section":"Abstract; Sections 2-5"},{"comment":"The procedure for extracting p_c is not formally defined. The text states that p_c is identified as the intersection of the average similarity curves of positive and negative sample pairs, but no equation or algorithm specifies how the curves are averaged, how the intersection is computed, or how uncertainties are propagated from individual similarity outputs to the finite-size estimates. The anchor-dependent thermodynamic-limit results in Tables 1-2 spread from 0.3090 to 0.3149 (site) and from 0.2484 to 0.2530 (bond); these spreads are larger than the reported uncertainties and are comparable to the claimed percent-level accuracy. The monotone-crossing premise also lacks a derivation. The authors should formalize the estimator, provide an error budget, and either demonstrate that anchor dependence is within statistical error or introduce an anchor-independent estimator.","section":"§4.2; §4.3; Tables 1-2"},{"comment":"The Monte Carlo calibration itself deviates from the accepted literature values by more than the reported statistical errors. Figure 4(f) gives p_c=0.3146(14) for 3D site percolation, while the text quotes the standard value 0.3116; Figure 5(f) gives p_c=0.2513(7) for bond percolation, versus 0.2488. These discrepancies are roughly 1% and are many times the stated uncertainties, yet the text says the results “align well with theoretical predictions.” Because these same configurations feed the SNN, the source of the bias should be identified (for example, the sigmoid fitting form, the choice of 1/L extrapolation without correction-to-scaling terms, or insufficient system sizes) and the calibration repeated or the discrepancy discussed explicitly.","section":"§4.1; Figs. 4-5"},{"comment":"There is a circularity risk in the labeling scheme that should be addressed with a control experiment. Positive labels are assigned when two configurations come from the same interval in {[0,0.1], [0.9,1]} and negative labels when they come from different intervals. This target already places the two labeled classes on opposite sides of any plausible p_c, so the network is effectively trained to output a monotone function of p; reading off a threshold from that monotone function may reflect the label construction rather than the true transition. The paper should test this by training the identical SNN on labels derived from an arbitrary split of the same probabilities (for example, positive for p in [0,0.1] vs. negative for p in [0.45,0.55] on a model with no transition at 0.3), or by shuffling the p-values attached to configurations, and showing that the extracted intersection no longer tracks the true p_c. Such a control would directly address whether the representation and the intersection carry information beyond the prescribed label intervals.","section":"§4.2; Labeling strategy; §4.3"}],"minor_comments":[{"comment":"There is a duplicated text fragment: “Unlike traditional supervised learning that assigns Unlike traditional supervised learning, which assigns labels to individual samples…” should be corrected.","section":"§4.2"},{"comment":"The phrase “22 labeled probability points” is ambiguous: each probability point contains 1000 configurations, so the labeled data volume is 22×1000 configurations per system size. Please state both the number of probability values and the number of configurations per value.","section":"Abstract; §4.2"},{"comment":"The statement that all sigmoid fits achieve “a goodness of fit exceeding 99.9%” is not defined; please report the specific metric (e.g., R², χ² per degree of freedom) and its value for each fit.","section":"§4.1; Figs. 4-5"},{"comment":"The architecture description is inconsistent: Eq. (4) says the similarity evaluator takes the concatenated embeddings, while Eq. (6) defines a distance D_W between embeddings and the text says that distance “or alternatively, the concatenated embeddings” is passed to the evaluator. Please specify which input the evaluator actually uses.","section":"§3; Eqs. (4)-(6)"},{"comment":"The captions read “FFS extrapolation”; this should be “FSS extrapolation” in both places.","section":"Captions of Figs. 6(b) and 6(e)"},{"comment":"The phrase “first successful application of the SNN method for predicting critical thresholds in three-dimensional percolation models” is stronger than the evidence presented; earlier work cited in the paper (e.g., Ref. 24) already applies SNNs to phase diagrams, and no comparison with prior 3D SNN results is given. Please soften or support this claim.","section":"Introduction; Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The paper's core idea is interesting and the Monte Carlo dataset is a useful resource, but the advertised ν estimate and fcc transfer are not present in the actual analysis. The former can likely be fixed by a free-ν collapse with reported uncertainties; the latter requires new experiments. The p_c extraction also needs a precise estimator definition and a control for the labeling circularity. I would not reject, but the revision must go beyond cosmetic changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: the core idea is plausible and the p_c results are in the right ballpark, but the abstract sells more than the paper delivers. The nu \"estimate\" is a fixed input, the FCC transfer and representation analyses appear only in the abstract, and the MC validation is several sigma off from literature.\n\nWhat's actually new: applying a pairwise Siamese network trained only on extreme-interval labels (p<0.1 or p>0.9) to 3D site and bond percolation, then reading p_c from the crossing of positive/negative similarity curves. That combination is not in the cited prior work, and it works at the percent level. They also check anchor dependence and show that extending the labeling interval gives diminishing returns, which is a useful observation. The paper is honest about the main method being a label-efficient wrapper rather than a new physical observable.\n\nWhere it gets soft. The abstract claims the method \"yields estimates of the critical exponent nu consistent with literature values within statistical uncertainty.\" That is not supported. Section 4.3 sets nu approx 0.88 to make the collapse, and Table 4 just reports 0.88 with no uncertainty or fitting. No free-nu fit is done, so nu is an input, not an estimate. That needs fixing before the claim can stand.\n\nTwo other claims in the abstract have no body: the FCC transfer result (\"identifies the phase transition in face-centered cubic lattices without retraining\") appears nowhere in the text, and the representation analysis (learned statistic coincides with S_max/L^3, r>0.99) is also missing. Both are advertised as headline results. That is a serious gap, not a minor omission.\n\nThe MC validation also has issues. Your sigmoid-FSS extrapolation gives site p_c=0.3146(14) and bond p_c=0.2513(7), but the literature values 0.3116 and 0.2488 are 2-3 sigma away. Calling that \"align well\" is generous. And the SNN thresholds vary by anchor: site 0.3090-0.3149, bond 0.2484-0.2530. That spread is larger than the per-anchor error bars, so the reported uncertainties understate the systematic error.\n\nBottom line: the label-efficient p_c detection is a legitimate incremental result, but the paper as written is overclaimed. It deserves a serious referee, but only with major revision: add the FCC experiments, show the representation analysis, replace the fixed-nu collapse with a free-nu fit (or retract the nu claim), and either provide code/data or at least report training details and the curve-intersection procedure formally.","headline":"Plausible label-efficient p_c detection for 3D percolation, but the abstract overclaims: nu is fixed not estimated, FCC transfer and representation results are missing from the body, and the MC baseline is off by several sigma.","tokens_in":25150,"tokens_out":3716,"would_cite":false,"duration_ms":37457,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Siamese neural network trained on just 22 labeled probability points from non-critical regions locates 3D percolation thresholds and the critical exponent ν.","keywords":["Siamese neural network","percolation threshold","critical exponent","3D site and bond percolation","label-efficient learning","largest-cluster size","finite-size scaling"],"falsifier":"Train the SNN with the same 22 labeled probability points but with anchors placed inside and immediately below the critical region, and compare the crossing points to the high-precision thresholds 0.3116 and 0.2488; if the crossing shifts systematically with anchor by more than the reported error bars, the extracted threshold is an artifact of the labeling protocol rather than the true critical point.","tokens_in":1618,"feed_emoji":"🔬","tokens_out":1804,"duration_ms":89519,"temperature":0.7,"pith_summary":"The paper tries to establish that phase transitions in three-dimensional percolation can be located with almost no labeled data: a Siamese neural network trained only on binary similarity labels for pairs drawn from the far non-critical regions [0,0.1] and [0.9,1] estimates the percolation threshold with percent-level accuracy and yields the critical exponent ν consistent with literature values. The network was never shown a configuration near the transition, yet its learned representation quantitatively matches the normalized largest-cluster size S_max/$L^{3}$, the standard finite-size order parameter. If true, this offers a practical route to criticality detection in settings where no order parameter is explicitly defined and labeled data are scarce.","feed_headline":"22 labels locate 3D percolation thresholds to within 1 percent","feed_subtitle":"A similarity-learning network finds p_c and the exponent ν without any labels near the transition.","key_machinery":"The central machinery is a Siamese neural network with two shared-weight fully connected branches that embed DFS-extracted largest-cluster configurations into a latent space, followed by a similarity evaluator trained with binary cross-entropy on positive and negative pairs. The load-bearing identity is that the learned embedding collapses onto the normalized largest-cluster size S_max/$L^{3}$, so the network's similarity score effectively measures the finite-size order parameter of percolation; the threshold is read as the crossing of average similarity curves for same-region versus cross-region pairs, and finite-size scaling extrapolates that crossing to the thermodynamic limit.","core_discovery":"The paper reports that, for site and bond percolation on a three-dimensional simple cubic lattice, a Siamese network trained solely on configuration pairs labeled as same-region or different-region, using 22 labeled probability points taken entirely from [0,0.1]∪[0.9,1], recovers the percolation threshold from the crossing of the positive and negative similarity curves. Finite-size scaling extrapolation gives p_c ≈ 0.309–0.315 for site percolation and ≈ 0.248–0.253 for bond percolation, compared with literature values of 0.3116 and 0.2488, and data collapse of the network outputs gives ν ≈ 0.88, consistent with the known value 0.8765. The paper further claims that the learned embedding coincides with the normalized largest-cluster size S_max/$L^{3}$ at correlation r > 0.99, and that a network trained solely on simple cubic lattices identifies the phase transition in face-centered cubic lattices without retraining.","pith_inferences":["If the learned embedding really is S_max/L^3, the same architecture should work for any model whose transition is governed by a spanning or largest-cluster observable, including continuum percolation and correlated percolation variants; that is a testable prediction beyond the paper.","The method could serve as an order-parameter discovery tool: when no quantitative order parameter is known, the SNN embedding itself may be used as the scaling variable in a finite-size collapse.","The anchor-to-anchor spread in the reported thermodynamic-limit thresholds (about 0.006 for site and 0.005 for bond, larger than the quoted extrapolation errors) invites a systematic study of how the crossing-point estimator depends on anchor location; a monotone trend would indicate a systematic component.","Applying the identical 22-point protocol to two-dimensional percolation, where p_c and ν are known to high precision, would provide a cheap external calibration of whether the crossing rule is unbiased."],"forward_implications":["Three-dimensional site and bond percolation thresholds on cubic lattices can be extracted to within about one percent using only 22 labeled probability points, none of them near the critical region.","The same network output, after data collapse, yields the correlation-length exponent ν ≈ 0.88, matching the literature value 0.8765 within statistical uncertainty.","The learned representation is, up to a correlation exceeding 0.99, the normalized largest-cluster size S_max/L^3, indicating that the network discovers the finite-size order parameter from similarity labels alone.","A network trained on simple cubic lattices transfers to face-centered cubic lattices without retraining, suggesting the learned statistic is not tied to one lattice geometry.","Extending the labeling interval toward the critical region does not significantly improve the threshold or exponent estimates, supporting the claim of label efficiency."],"supporting_citations":[{"why":"Supplies the Siamese-neural-network approach to unsupervised phase-transition detection that this work extends to 3D percolation.","marker":"[24]"},{"why":"Provides the Siamese architecture and one-shot similarity-learning basis adopted for configuration pairs.","marker":"[30]"},{"why":"Establishes the semi-supervised transfer-learning baseline for percolation that this method builds on.","marker":"[27]"},{"why":"Demonstrates domain-adversarial transfer for critical point prediction, which motivates the cross-lattice transfer claim.","marker":"[28]"},{"why":"Supplies the literature values of the percolation thresholds (0.3116 site, 0.2488 bond) used as the comparison standard.","marker":"[31]"},{"why":"Provides high-precision Monte Carlo results for three-dimensional site percolation used to validate the simulation data.","marker":"[43]"},{"why":"Supplies the efficient Monte Carlo algorithm used to generate the percolation configurations.","marker":"[13]"},{"why":"Provides the finite-size scaling relations used for threshold extrapolation and data collapse.","marker":"[36]"}],"fun_headline_variants":["22 labels locate 3D percolation threshold to 1%","Siamese net learns criticality from 22 far-off labels","Label-efficient AI finds percolation threshold and exponent","Train on simple cubic, identify FCC phase transition","No labels near transition? 22 far ones suffice for 3D"],"cache_read_input_tokens":27264,"weakest_assumption_plain":"The paper assumes that the point where the positive and negative similarity curves cross marks the true percolation threshold, independent of the chosen anchor and labeling interval, even though the network was trained only on labels from outside the critical region.","fun_headline_variants_meta":{"raw":{"variants":["22 labels locate 3D percolation threshold to 1%","Siamese net learns criticality from 22 far-off labels","Label-efficient AI finds percolation threshold and exponent","Train on simple cubic, identify FCC phase transition","No labels near transition? 22 far ones suffice for 3D"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000172,"raw_usage":{"total_tokens":1294,"prompt_tokens":984,"completion_tokens":310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":600,"completion_tokens_details":{"reasoning_tokens":225}},"tokens_in":600,"tokens_out":310,"duration_ms":4078,"temperature":1.0,"reasoning_tokens":225,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:58:44.326995+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the SNN with the same 22 labeled probability points but with anchors placed inside and immediately below the critical region, and compare the crossing points to the high-precision thresholds 0.3116 and 0.2488; if the crossing shifts systematically with anchor by more than the reported error bars, the extracted threshold is an artifact of the labeling protocol rather than the true critical point.","supporting_citations":[{"cited_title":"& Wetzel, S","cited_arxiv_id":null,"evidence_quote":"Supplies the Siamese-neural-network approach to unsupervised phase-transition detection that this work extends to 3D percolation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Siamese architecture and one-shot similarity-learning basis adopted for configuration pairs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the semi-supervised transfer-learning baseline for percolation that this method builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates domain-adversarial transfer for critical point prediction, which motivates the cross-lattice transfer claim."},{"cited_title":"& Aharony, A","cited_arxiv_id":null,"evidence_quote":"Supplies the literature values of the percolation thresholds (0.3116 site, 0.2488 bond) used as the comparison standard."},{"cited_title":"& Blöte, H","cited_arxiv_id":null,"evidence_quote":"Provides high-precision Monte Carlo results for three-dimensional site percolation used to validate the simulation data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the efficient Monte Carlo algorithm used to generate the percolation configurations."},{"cited_title":"& Moloney, N","cited_arxiv_id":null,"evidence_quote":"Provides the finite-size scaling relations used for threshold extrapolation and data collapse."}],"review_version":1}