{"id":"30ae55bd-8453-42f6-a3c0-f02a7e94383c","arxiv_id":"2412.06768","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"A vision transformer predicts topological visibility and LDOS-based Majorana indicators from simulated nanowire conductance traces with high held-out accuracy, but with no experimental validation.","lead":"This paper trains a vision transformer on simulated conductance measurements to predict whether a disordered Majorana nanowire is in a topological phase, and to map the full phase diagram. The method also infers LDOS-based Majorana indicators from conductance traces, but the entire validation is on simulated data, so its practical value depends on how well the simulations match real devices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The P>0.9998 false-positive claim is not statistically supported: zero observed failures on a small test set gives a much wider confidence bound, and cutoffs are selected on the same test data.","rationale":"I focused on the statistical support for the headline confidence claim because it is the paper's most quantitative assertion and it is used to motivate experimental applicability. The supervised-learning core is credible: the network is trained on KWANT simulations, and the reported held-out RMSE values indicate that it learns the simulator's conductance-to-indicator mapping. The weakness is not the architecture but the evaluation protocol for 'P>0.9998'. A finite-sample zero-false-positive frequency, with the cutoff selected on the same test data, cannot support a probability below 0.0002; the exact binomial bound is orders of magnitude larger. This does not falsify the method, but it means the abstract's strongest quantitative claim is not established, even in simulation. The reader's stated weakest assumption is simulator-to-experiment transfer, which is a separate and also valid concern. However, the reader's rationale also explicitly notes the lack of confidence intervals and test-set cutoff selection, so there is partial overlap with my concern. I keep the verdict CONDITIONAL rather than moving to ACCEPT or REJECT: the central simulation-based finding is plausible, but the confidence statement needs to be made statistically principled, and an out-of-distribution or experimental transfer test would be needed to support the broader applicability claim.","tokens_in":26627,"tokens_out":7805,"duration_ms":85150,"concrete_test":"Re-run the full-regime evaluation with a strict three-way split: 80% training, 10% validation, 10% untouched test. Select Ccutoff using only the validation set, freeze it, evaluate once on the test set, and report the number of passing test devices together with the exact Clopper-Pearson one-sided 95% upper bound for FP at each cutoff. If the number of passing devices at the cutoff used for the P>0.9998 claim is below about 15,000, or if the resulting upper bound exceeds 0.0002, the headline probability should be replaced by the computed bound (for example, FP < 0.005). This directly settles whether the observed zero false positives are consistent with the claimed FP < 0.0002.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim, stated in the abstract and Section IV.A, is that the classifier can push the false-positive probability below 0.0002. The evidence is a finite-sample frequency: entries of 1−FP = 1.0000 in Tables II–IV correspond to zero observed false positives in a single 10% test split, with no confidence intervals. Given roughly 20,000–30,000 total realizations (Section III.C), the test set is about 2,000–3,000 devices. At the full-regime cutoff Ccutoff=−0.8, the passing fraction is 0.3077, so only a few hundred test devices pass. Zero failures in a few hundred trials gives a one-sided 95% upper bound on the false-positive probability of roughly 0.3–1%, not 0.02%. To support FP < 0.0002 at 95% confidence with zero failures would require about 15,000 passing devices. Moreover, the cutoff is chosen after examining the same test data used for reporting performance, as stated in Section IV.A: 'using our test data, we do not find a single instance.' This makes the reported fidelity an optimistic selected value rather than a pre-registered threshold. The claim of 'arbitrary confidence' therefore outruns the statistical evidence, even within the simulated setting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a Vision Transformer on simulated four-terminal differential conductance traces (GLL, GRR, GRL, GLR) of disordered Majorana nanowires. The network is tasked with predicting the continuous scattering-matrix topological visibility TV, classifying devices as topologically non-trivial, reconstructing the TV phase diagram in (B, μ), and inferring two LDOS-based alternative indicators I1 and I2. The authors report RMSE values between 0.15 and 0.62 across three disorder regimes and claim that by tuning a continuous cutoff the false-positive probability of declaring a device topological can be pushed below 0.0002 (P>0.9998). They conclude that the conductance-only method should be usable for analyzing future experiments.","tokens_in":26932,"tokens_out":7213,"duration_ms":72411,"significance":"If the central claims hold, the method offers a practical way to extract topological information from routine four-terminal conductance measurements, including indicators that are not directly measurable experimentally. The use of a 3D-patched Vision Transformer with a multi-path tree decoder is a nontrivial architectural adaptation, and the simulation pipeline is based on standard KWANT transport calculations with publicly available data-generation code (Ref. [50]). The supervised-learning core is credible: held-out conductance traces are predicted with RMSE values of roughly 0.15-0.5, and the phase diagrams show good qualitative agreement. However, the headline statistical claim about arbitrary confidence is not supported by the finite test set, and the experimental-applicability conclusion goes beyond the evidence provided.","major_comments":[{"comment":"The claim that the false-positive probability can be pushed below 0.0002 ('P>0.9998') is not statistically supported. The reported entries of 1−FP = 1.0000 correspond to zero observed false positives in a single 10% test split with the cutoff selected on that same test data (Sec. IV.A: 'using our test data, we do not find a single instance'). With roughly 20,000–30,000 total realizations, the test set contains only a few thousand devices, and the number of passing devices at the relevant operating points is a few hundred to at most ~1,500. Zero failures in N trials gives a one-sided 95% upper bound of about 3/N on the failure probability, which is an order of magnitude larger than 0.0002. The abstract's 'up to arbitrary confidence (P>0.9998)' should be replaced by a statement with a confidence interval, or the test set must be enlarged so that zero observed failures actually supports the claimed bound.","section":"§IV.A, Tables II–IV"},{"comment":"The false-negative metric is misdefined. The table captions state FN = P(TV < Ccutoff | TPred > 0), which conditions on a predicted positive TV, not on the complement of passing (TPred ≥ Ccutoff). The standard false-negative rate for a pass/fail decision is P(TPred ≥ Ccutoff | TV < Ccutoff), i.e., the fraction of truly topological devices that are rejected. The reported '1−FN' values therefore do not quantify the risk that topological devices are discarded, which is exactly the question the text claims to address when discussing whether the method rejects too many devices. Please report a standard confusion matrix with sensitivity and specificity so that the trade-off between false positives and false negatives is unambiguous.","section":"§IV.A, definitions of FN in Tables II–IV"},{"comment":"The conclusion that the technique 'should be usable to analyze future experiments' is an extrapolation beyond the evidence. All tests are performed on simulated conductance traces generated from the same Hamiltonian family, with fixed barrier height (15 mV), fixed length (3 μm), fixed material parameters, and with disorder and spin-orbit coupling drawn from the training ranges. There is no out-of-distribution test, no variation of barrier height or device length, and no experimental data. The network may be learning the specific simulation manifold, and no evidence is given that real device traces lie on that manifold. The claims should be restricted to the simulated parameter ranges, and the manuscript should explicitly state that transfer to different geometries or barrier settings requires retraining or an explicit domain-adaptation test.","section":"§V and abstract/conclusion"}],"minor_comments":[{"comment":"In the full-regime table, the row for Ccutoff = −0.9 lists P(Passing) = 0.0010, which is inconsistent with the gradual progression from 0.3077 at Ccutoff = −0.8 and is never discussed in the text. Please check whether this is a typo and ensure the table matches the narrative.","section":"Table III"},{"comment":"The simulated conductance traces are repeatedly called 'measurements,' which can blur the distinction between simulated training data and experimental data. Consider reserving 'measurements' for actual experimental data and using 'simulated conductance traces' for the KWANT outputs.","section":"§II"},{"comment":"The rescaling definitions I2 = min(I2^(4) − 1, 1) and I1 = min(I1^(4)/0.1 − 1, 1) are not bounded below as stated in the text ('between 1 and −1'). Unless an additional lower truncation is applied, these quantities can be arbitrarily negative. Please clarify the exact clipping operation.","section":"§III.A"},{"comment":"The network was trained from random initialization, but no information is given about the variance of the reported metrics across random seeds or hyperparameter choices. Given the small test set, reporting seed-averaged results with standard deviations would substantially strengthen the quantitative claims.","section":"§III.C"},{"comment":"The trained model is not publicly released; Ref. [49] states that the authors 'are happy to help' rather than providing a repository. For reproducibility, consider releasing the model weights with a DOI-stable archive.","section":"Ref. [49]"},{"comment":"The phrase 'arbitrary confidence' is used repeatedly (abstract, Sec. IV.A, conclusion). Since the attainable confidence is bounded by finite test statistics and the selected cutoffs, 'tunable confidence' or 'high confidence' would be more precise.","section":"Throughout"},{"comment":"The edge length is denoted ϵ in the text but appears as ξ in the definition of I2; please make the notation consistent.","section":"§III.A"},{"comment":"Every figure caption begins with a stray '(a)', likely a LaTeX artifact. Please remove it unless subfigures are actually intended.","section":"Figure captions"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is from a group with extensive prior work on disorder and Majorana nanowires, and the simulation-based ML result is plausible. The main concern for the editor is the statistical overclaim: the 'P>0.9998' statement is an empirical frequency on a finite test set with cutoff selection on the same data, not a validated error bound. In addition, the false-negative metric is nonstandard and the experimental-applicability claim goes beyond the tested in-distribution simulation manifold. These issues are fixable with revised statistics and tempered claims, hence major revision rather than rejection. I would also suggest the editor ask for a validation set separate from the test set used for cutoff selection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: don't quote this paper's 'P>0.9998' headline, but do take the underlying method seriously. The core claim—that a vision transformer trained on simulated four-terminal differential conductance can predict the continuous scattering invariant TV, the full (B, mu) phase diagram, and the LDOS-based indicators I1 and I2 across disorder regimes—is credible and reproducible in principle. The held-out RMSE values (0.15–0.40 depending on regime) support that the network really learned the KWANT mapping.\n\nWhat's actually new: prior ML work on Majorana nanowires (Cheng et al.) used XGBoost to classify only high- and low-disorder extremes. This paper extends to the continuous invariant, the intermediate disorder regime, full phase diagrams, and the two LDOS indicators that have not been learned before. The 3D vision transformer adaptation is a technical extension, but a sensible one. The paper is also honest that validation is purely simulated and that topology is ambiguous in short disordered wires.\n\nThe soft spots are in the headline statistics. The claim of FP < 0.0002 is an empirical frequency: zero false positives in the test set at a cutoff chosen after looking at that same test set. With only a few hundred devices passing at the Ccutoff=-0.8 operating point, the one-sided 95% upper bound on FP is on the order of 0.3–1%, not 0.02%. The paper's 'arbitrary confidence' language is not supported. This is fixable: report a confidence interval, use a separate validation set for cutoff selection, or state the bound honestly.\n\nThe other soft spot is transfer to experiment. The method is trained on one Hamiltonian family with fixed material parameters, disorder model, barrier height, and length. Real devices will deviate from this distribution in ways the paper does not test. The authors acknowledge this implicitly, but the conclusion goes further than the evidence: 'should be usable to analyze future experiments' is a hope, not a demonstrated result.\n\nNone of this sinks the paper. Within the simulated setting, the supervised learning result is solid, and the paper is transparent about its limitations. It deserves a serious referee, and I'd accept it with revisions. The right referee should push on the statistical claim and ask for at least one out-of-distribution check before the experimental suggestion is taken literally.\n\nFor a reader: if you work on Majorana diagnostics or ML for quantum devices, this is worth reading and citing for the method, not for the 0.9998. I'd bring it to a reading group session on ML in condensed matter.\n\nBest,","headline":"Solid, reproducible ML study of Majorana nanowire diagnostics; the headline P>0.9998 claim outruns the statistics.","tokens_in":27480,"tokens_out":2599,"would_cite":true,"duration_ms":27736,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A vision transformer trained on simulated conductance traces can read out the topological state of a Majorana nanowire, including indicators that are not directly measurable.","keywords":["Majorana zero modes","disordered nanowires","vision transformer","topological invariant","conductance spectroscopy","machine learning","topological visibility","local density of states"],"falsifier":"Feed the trained network four-terminal conductance traces from a device whose true topology is known independently, for instance by a direct numerical solution of its full Hamiltonian or by braiding measurements, and compare the predicted $TV$ and phase diagram with the exact values; any systematic disagreement would falsify the claim that conductance alone determines topology. A simpler out-of-distribution check is to vary the barrier height or wire length outside the training ranges and see whether the prediction error collapses.","tokens_in":26367,"feed_emoji":"⚛️","tokens_out":6488,"duration_ms":64147,"temperature":0.7,"pith_summary":"The paper argues that a vision transformer trained only on simulated four-terminal conductance measurements can determine the topological state of a Majorana nanowire, including the continuous scattering invariant, the full phase diagram in magnetic field and chemical potential, and two local-density-of-states indicators that experiments cannot measure directly. The motivation is that disorder routinely produces zero-bias conductance peaks that mimic Majorana zero modes, so standard experimental signatures are unreliable. If correct, the method would let experimentalists declare a device topologically non-trivial with an arbitrarily low false-positive probability by tuning a cutoff, and locate the parameter region worth using for braiding. The whole procedure is validated in simulation across low, moderate, and high disorder, with the moderate regime matched to current experimental devices.","feed_headline":"Conductance traces alone reveal a nanowire's topological phase","feed_subtitle":"It maps the full (B, mu) phase diagram and hidden LDOS indicators, all from routine conductance data.","key_machinery":"The load-bearing object is a generalized Vision Transformer operating on a three-dimensional input image whose channels are the four conductance traces $G_{LL}$, $G_{RR}$, $G_{RL}$, and $G_{LR}$, and whose axes are bias voltage, magnetic field, and chemical potential. The network applies three-dimensional patching, additive positional encoding, four transformer blocks, and then splits into 100 small multilayer perceptron heads, one per $(B,\\mu)$ point, each producing a continuous indicator; a separate network is trained for each indicator. The continuous output lets the user impose a cutoff $C_{\\mathrm{cutoff}}$ such that only devices predicted below it are declared topological, converting raw accuracy into tunable confidence. Training data come from numerical solutions of the standard nanowire Bogoliubov–de Gennes Hamiltonian with Gaussian disorder, randomized spin-orbit coupling, and finite-temperature convolution.","core_discovery":"The central claim is that the four measured differential conductances of a nanowire, local at each end and nonlocal end-to-end, contain enough information to determine the scattering-matrix topological invariant $TV$, its sign, and its full $(B,\\mu)$ phase diagram, even when disorder amplitude, correlation length, and spin-orbit coupling are unknown. The paper further claims that the same conductance inputs determine the LDOS-based operational indicators $I_1$ and $I_2$, which assess whether end-localized Majoranas are usable for fusion and braiding, not merely present. Because $TV$ is predicted continuously rather than as a binary label, a user can set a passing cutoff and make the probability of falsely declaring a trivial device topological arbitrarily small, with reported fidelities above $0.9998$ at the most stringent cutoffs. The authors conclude that conductance data alone should suffice to analyze future experiments.","pith_inferences":["Editorial inference: the network's success implies the four conductances carry enough mutual information to fix the scattering invariant; if so, the same architecture should transfer to other topological platforms where transport spectra encode a topological index, such as Josephson junctions or higher-order topological insulators.","Editorial inference: because the paper finds prediction errors behave like a Gaussian filtering of the true phase diagram, a low-pass filtered version of the network output could serve as a conservative operating map, with the caveat that smoothing may erase narrow topological patches.","Editorial inference: the transfer claim would be directly testable by intentionally training on one material parameter set and testing on conductance traces generated with a moderately different barrier height or wire length; the paper reports no such out-of-distribution test.","Editorial inference: since the network never sees disorder parameters, its predictions are only as reliable as the training distribution's coverage of experimental reality; active learning on real devices would be the natural next step."],"forward_implications":["An experimentalist can use routine conductance measurements to declare a device topological with false-positive probability below $0.0002$, by choosing a sufficiently negative cutoff on the predicted $TV$.","The full topological phase diagram in $(B,\\mu)$ can be reconstructed from conductance alone, identifying the parameter window where a device should be operated for fusion and braiding experiments.","The LDOS-based indicators $I_1$ and $I_2$, which cannot be measured directly, can be inferred from the same conductance data, giving a check on whether end Majoranas are localized and usable.","The method replaces arbitrary visual thresholds in protocols like the topological gap protocol with a quantitative, tunable decision rule.","Retraining the same architecture with different material parameters would extend the technique to other nanowire platforms without changing the method."],"supporting_citations":[{"why":"supplies the vision transformer architecture that the paper generalizes to three-dimensional conductance inputs.","marker":"[39]"},{"why":"defines the scattering invariant whose sign indicates whether the wire is topologically non-trivial.","marker":"[27]"},{"why":"introduces topological visibility as the continuous scattering indicator that the network predicts.","marker":"[28]"},{"why":"shows why short disordered wires need more than the sign of the invariant, motivating continuous predictions.","marker":"[29]"},{"why":"defines the LDOS-based Majorana indicators $I_1$ and $I_2$ that the network learns from conductance data.","marker":"[31]"},{"why":"supplies the numerical quantum transport solver that generates all conductance training and test data.","marker":"[34]"},{"why":"provides the experimental InAs-Al device context and the topological gap protocol that set the moderate-disorder regime and parameter ranges.","marker":"[18]"},{"why":"is the prior machine-learning baseline on zero-bias peak measurements that the extremes-regime results are compared against.","marker":"[33]"},{"why":"is prior work on extracting disorder properties from Majorana nanowire conductance, which this paper extends by predicting topological indicators directly.","marker":"[25]"},{"why":"supplies the nonlocal conductance framework used to distinguish topological from trivial behavior in realistic disordered systems.","marker":"[37]"}],"fun_headline_variants":["AI reads Majorana topology from conductance traces","Transformer predicts nanowire topological phase from data","Conductance alone yields full Majorana phase diagram","Neural network maps disorder-robust topological phases","Deep learning distinguishes Majorana signals from noise"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole transfer to experiment rests on the assumption that real nanowire conductance traces are statistically close to the simulated training distribution, meaning the same Hamiltonian form and the same ranges of disorder, spin-orbit coupling, barrier strength, temperature, and wire length.","fun_headline_variants_meta":{"raw":{"variants":["AI reads Majorana topology from conductance traces","Transformer predicts nanowire topological phase from data","Conductance alone yields full Majorana phase diagram","Neural network maps disorder-robust topological phases","Deep learning distinguishes Majorana signals from noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000797,"raw_usage":{"total_tokens":3545,"prompt_tokens":1022,"completion_tokens":2523,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":2453}},"tokens_in":638,"tokens_out":2523,"duration_ms":19108,"temperature":1.0,"reasoning_tokens":2453,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:19:02.718575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed the trained network four-terminal conductance traces from a device whose true topology is known independently, for instance by a direct numerical solution of its full Hamiltonian or by braiding measurements, and compare the predicted $TV$ and phase diagram with the exact values; any systematic disagreement would falsify the claim that conductance alone determines topology. A simpler out-of-distribution check is to vary the barrier height or wire length outside the training ranges and see whether the prediction error collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the vision transformer architecture that the paper generalizes to three-dimensional conductance inputs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the scattering invariant whose sign indicates whether the wire is topologically non-trivial."},{"cited_title":"Thamm and B","cited_arxiv_id":null,"evidence_quote":"introduces topological visibility as the continuous scattering indicator that the network predicts."},{"cited_title":"Akhmerov, J","cited_arxiv_id":null,"evidence_quote":"shows why short disordered wires need more than the sign of the invariant, motivating continuous predictions."},{"cited_title":"Das Sarma, J","cited_arxiv_id":null,"evidence_quote":"defines the LDOS-based Majorana indicators $I_1$ and $I_2$ that the network learns from conductance data."},{"cited_title":"Peeters, T","cited_arxiv_id":null,"evidence_quote":"supplies the numerical quantum transport solver that generates all conductance training and test data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the experimental InAs-Al device context and the topological gap protocol that set the moderate-disorder regime and parameter ranges."},{"cited_title":"Pan and S","cited_arxiv_id":null,"evidence_quote":"is the prior machine-learning baseline on zero-bias peak measurements that the extremes-regime results are compared against."},{"cited_title":"Das Sarma and H","cited_arxiv_id":null,"evidence_quote":"is prior work on extracting disorder properties from Majorana nanowire conductance, which this paper extends by predicting topological indicators directly."},{"cited_title":"Das Sarma, J","cited_arxiv_id":null,"evidence_quote":"supplies the nonlocal conductance framework used to distinguish topological from trivial behavior in realistic disordered systems."}],"review_version":1}