REVIEW 3 major objections 5 minor 35 references
Response to recent comments on Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supp. Info. for Nature 638, 651-655 (2025)
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims that two external critiques fail to identify any flaw in the topological gap protocol's false discovery rate estimate.
desk verdict A competent rebuttal that adds useful robustness checks but overclaims 'no flaws' while leaving the simulation-to-device distributional match untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the false discovery rate (FDR): the probability that the TGP flags a trivial region as topological, estimated by running the protocol on simulated transport data calibrated to independently measured device disorder (localization length over 1 $\mu$m). The rebuttal's machinery is the transfer argument: because the simulations are drawn from the same distribution as real devices, a low simulated FDR bounds the experimental FDR. The two new tables are the operative mechanism: Table I shows that using the experimental setting `average_over_cutter=True` leaves the FDR statistically unchanged, and Table II shows that cropping $B_{\max}$ to 2.5 T keeps the FDR below 9% — together with a threshold-sensitivity analysis showing that changes in $G_{\mathrm{th}}$ of less than 20% move the FDR by only a few percent.
What would settle it
Re-run the released TGP code on the published simulated datasets with every implementation detail set to the experimental pipeline — cutter averaging on, corrected symmetric bias cropping, and magnetic-field ranges cropped to $B_{\max} \leq 2.5$ T as in Table II — and count false positives under the paper's own topological criterion; if the fraction of false positives among roughly 700 regions of interest exceeds 8%, the central claim fails. The code repository and data paths named in the paper make this check directly executable.
Extended reading notes
Core claim
On the paper's own terms, the claim to be defended is that no flaw has been shown in the TGP's false discovery rate: re-running the protocol with the experimental setting `average_over_cutter=True` yields at most one false positive in roughly 700 regions of interest (Table I), and cropping the magnetic-field range to $B_{\max} \leq 2.5$ T changes the FDR bound only slightly (Table II). The paper upholds the original claims that a TGP pass indicates a topological region with probability above $1-8\%=92\%$ for the SLG-$\beta$ stack and above $94\%$ for the DLG-$\epsilon$ stack, that measurement-range variations are explained by cooldown-to-cooldown disorder changes and stage-1 cluster sizes, and that the acknowledged bias-cropping bug in the Nature supplement alters extracted gap values by less than 5 $\mu$eV in more than 96% of pixels without changing the parity result. It concludes that the comments attack a 'smoking gun' methodology the papers never used.
Load-bearing premise
Everything rests on the assumption that the simulated datasets used to calibrate the protocol are drawn from the same probability distribution as the real device data; if the simulations overstate how representative the devices are, the claimed FDR bound could be too low.
Editorial extensions
If this is right
- If the rebuttal is correct, the <8% FDR bound (and <6% for the DLG-$\epsilon$ stack) survives, so regions that pass the TGP remain high-confidence operating points for topological qubit tune-up.
- The distinction between statistical and smoking-gun tests becomes the operative frame: hyperparameter sensitivity of the TGP is not evidence of bias, and critiques must engage with the FDR itself.
- The acknowledged bias-cropping bug in the Nature supplement is contained: gap values shift by under 5 $\mu$eV in more than 96% of pixels, and the flux-dependent bimodal random-telegraph-signal parity result is not affected.
- Measurement-range differences between devices and cooldowns are presented as natural consequences of stage-1 cluster sizes and device drift, with the simulation range distribution in Fig. 2 containing the experimental ranges.
Reading between the lines
- One could press further than the paper does: the robustness checks vary one parameter at a time, so a joint worst-case sweep over $G_{\mathrm{th}}$, cutter averaging, bias cropping, and $B_{\max}$ would directly test whether the FDR bound remains below 8% under combined stress.
- The same calibration logic could be exported to other material platforms: if simulations matched to measured localization length yield a similarly low FDR for other nanowire systems, the protocol would be a general tune-up tool rather than a device-specific one.
- The debate implicitly raises a question the paper leaves open: the true-positive criterion uses the scattering invariant together with stable zero-bias peaks and gap reopening, but the false-negative rate is explicitly unquantified; a reader wanting to use the TGP for discovery, not just tune-up, would want that number.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a point-by-point rebuttal by the Microsoft Quantum collaboration to two comments by H. F. Legg (arXiv:2502.19560 and arXiv:2503.08944) on the topological gap protocol (TGP) as used in Phys. Rev. B 107, 245423 (2023) and in the Supplementary Information of Nature 638, 651-655 (2025). The rebuttal addresses seven points from Ref. 3 and several from Ref. 4, covering the definition of the transport gap, the threshold parameter Gth, the difference between simulated and experimental analysis (average_over_cutter), measurement ranges and their effect on TGP outcomes, the definition of 'topological' via the scattering invariant, and the interpretation of conductance data in the parity-readout devices. The manuscript provides new simulation statistics in Tables I and II, a stability analysis of ROI2s under magnetic-field-range cropping in Fig. 3, and a reproducibility statement with code links. The central claim is that no flaws have been identified in the FDR estimate (below 8%) and that the objections raised in Refs. 3 and 4 are unfounded.
Significance. If the rebuttal is correct, it defends the reliability of the TGP as a statistical tuning tool, preserving the conclusions of two high-profile experimental papers. The manuscript's strengths include the provision of reproducible code and package environment files, new simulation tables showing FDR bounds under changed analysis settings, and a direct stability analysis of individual ROI2s under magnetic-field-range modifications. The point-by-point format is useful for the community. However, the significance is tempered by two load-bearing caveats identified in the manuscript itself: the admitted cropping bug in the published SI for the Nature paper, and the explicitly quoted but not quantitatively tested assumption that simulated data are drawn from the same probability distribution as experimental data. These caveats mean that the strong 'no flaws' claim in the abstract is not fully established as written.
major comments (3)
- [Technical Response to Ref. 4, paragraph beginning 'We note that the published version of the TGP data...'] The manuscript admits a data-processing bug in the published Supplementary Information for Ref. 2: the bias range was incorrectly cropped, and correcting it changes TGP outcomes for device B, including the appearance of a new SOI2 for one cutter and an increase in ROI2 size. This admission is difficult to reconcile with the abstract's claim that 'no flaws have been identified in our estimate of the FDR' and that the objections are 'unfounded.' At least one objection (Ref. 4) identified a genuine error in the published data analysis. The rebuttal should explicitly narrow the 'no flaws' claim to the FDR estimate itself and acknowledge that the published SI contained a flaw, or provide a detailed argument for why an error that changes TGP outcomes does not affect the FDR bound. As written, the strong claim overreaches the manuscript's own findings.
- [Quotation from Ref. 1 in 'Technical Response to Ref. 3', point 2] The FDR transfer from simulation to experiment relies on the assumption, quoted in the manuscript, that 'the simulated data is drawn from the same probability distribution as the data produced by real devices.' The new robustness analysis in Tables I and II and Fig. 3 varies parameters inside the simulation model (average_over_cutter, magnetic field range) but does not test the distributional match between simulated and experimental conductance features. Independent localization-length measurements (Fig. 5) constrain disorder but do not establish that TGP-relevant features--zero-bias peak statistics, nonlocal gap shapes, Andreev-enhanced subgap conductance, junction asymmetries--are drawn from the same distribution. The rebuttal should either provide a quantitative comparison between simulated and experimental conductance distributions or explicitly state that the FDR bound is conditional on this untested assumption. Without such a test, the claim that 'no flaws have been identified' is not fully supported.
- [Technical Response to Ref. 3, point 1 (threshold Gth)] The manuscript states that 'changes in Gth of less than 20% result in only minor variations in the FDR (i.e. a few percent).' This statement is load-bearing for the robustness of the FDR estimate, but no quantitative evidence is provided for this specific claim. Figure 1 shows the presence/absence of ROI2s for a single measurement as Gth varies from 0.04 to 0.06 Gmax, but it does not report FDR values or statistics over the simulation ensemble. If the FDR is to be claimed robust against Gth variations, the manuscript should include a simulation table or plot showing FDR as a function of Gth, or should rephrase the statement to describe the observed stability without assigning a precise few-percent bound.
minor comments (5)
- [Abstract] The phrase 'the objections in arXiv:2502.19560 and arXiv:2503.08944 are unfounded' is too broad given the admitted cropping bug in the published SI. Recommend softening to 'the objections do not affect the FDR estimate' or similar.
- [Section 'Technical Response to Ref. 4', paragraph on measurement ranges] The text says the experimental stage 2 ranges are 'well within' the simulated distribution shown in Fig. 2, but the figure does not overlay the experimental values. Adding such an overlay would make the claim directly verifiable.
- [Table II footnote [15]] The explanation that 'most of the ROI2s are at high fields for the DLG-ε stack' is helpful but could be expanded to state whether this is a property of the simulated model or a consequence of the chosen parameter ranges.
- [Section 'Technical Response to Ref. 3', point 4] The discussion of the scattering invariant as a finite-size criterion is clear, but the manuscript refers to 'Fig. 32 in our paper' without a self-contained reproduction; since this is a response, a brief description of the relevant panels would aid readers who do not have Ref. 1 open.
- [Throughout] The manuscript is written as a collective reply under 'Microsoft Quantum' with a long author list in a footnote. For reproducibility and transparency, it would help to indicate which authors produced the new simulations and tables.
Circularity Check
No significant circularity: the FDR estimate is calibrated against an external scattering-invariant label and the transfer-to-experiment assumption is explicit, not a hidden input.
full rationale
The paper is a rebuttal rather than a new derivation; its central assertion is that the FDR estimate from Ref. [1] survives the comments in arXiv:2502.19560 and arXiv:2503.08944. The FDR was computed by applying the TGP to simulated conductance data whose ground-truth labels come from the scattering invariant det(r) < 0, an external criterion drawn from the published literature (e.g., Ref. [16] of the paper), not from the TGP output itself. The transfer from simulations to experiment is explicitly conditional: the paper quotes Ref. [1] as stating that the <8% bound holds 'provided that the simulated data is drawn from the same probability distribution as the data produced by real devices.' That is a stated assumption rather than a circularly hidden input. The new robustness checks (Tables I and II, Fig. 3) vary hyperparameters and magnetic-field ranges inside the simulation ensemble, and they do not reuse the conclusion as a premise. Self-citations to Refs. [1,2] point to published papers with a public code repository [14], which the paper explicitly invites readers to use for validation; under the review rules, code-reproduced results count as independent support. The admitted bias-range bug in the Nature Supplementary Information is disclosed and its quantitative effect is assessed rather than used to redefine the outcome. No 'prediction' in the response reduces by construction to a fitted parameter, a renamed input, or a self-citation chain. The main weakness—that the simulation-to-experiment distributional match is not directly tested—is a validity risk, not circularity, and it does not make the derivation self-referential.
Assumptions & free parameters
free parameters (3)
- Gth (conductance threshold for gap extraction) =
0.05 * Gmax (exp(-3))
- average_over_cutter =
True for experimental, False for simulated data (original); response says correct value is True
- Stage 2 magnetic field range Bmax =
~3 T in simulations, 2.5 T in robustness test
assumptions (3)
- domain assumption Scattering invariant det(r) < 0 reliably identifies topological phase in finite-sized disordered nanowires
- domain assumption Simulated transport data is drawn from the same probability distribution as experimental data
- standard math Binomial confidence intervals for zero false positives give valid FDR upper bounds
Cite this review
Pith. "Pith review of Response to recent comments on Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supp. Info. for Nature 638, 651-655 (2025)." pith.science (2026). https://pith.science/paper/3QAL3FNP
@misc{pith2026250413240,
author = {Pith},
title = {Pith review of: Response to recent comments on Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supp. Info. for Nature 638, 651-655 (2025)},
year = {2026},
howpublished = {\url{https://pith.science/paper/3QAL3FNP}},
note = {Machine review of arXiv:2504.13240}
}
read the original abstract
The topological gap protocol (TGP) is a statistical test designed to identify a topological phase with high confidence and without human bias. It is used to determine a promising parameter regime for operating topological qubits. The protocol's key metric is the probability of incorrectly identifying a trivial region as topological, referred to as the false discovery rate (FDR). Two recent manuscripts [arXiv:2502.19560, arXiv:2503.08944] engage with the topological gap protocol and its use in Phys. Rev. B 107, 245423 (2023) and Subsection S4.3 of the Supplementary Information for Nature 638, 651-655 (2025), although they do not explicitly dispute the main results of either one. We demonstrate that the objections in arXiv:2502.19560 and arXiv:2503.08944 are unfounded, and we uphold the conclusions of Phys. Rev. B 107, 245423 (2023) and Nature 638, 651-655 (2025). Specifically, we show that no flaws have been identified in our estimate of the false discovery rate (FDR). We provide a point-by-point rebuttal of the comments in arXiv:2502.19560 and arXiv:2503.08944.
Figures
Reference graph
Works this paper leans on
-
[1]
Identification of the ‘gap’ differs between publication and released code [3]. This is incorrect. There is no difference [1]. In addition, it is worth emphasizing that the transport gap requires a careful definition in a finite-sized disordered system, as discussed at length in Ref. 1, but seemingly overlooked in Ref. 3, which uses quotation marks around ...
-
[2]
The TGP applied to experiments is not the same TGP applied to theoretical simulations [3]. The difference [ 1] is very minor and is easily changed by modifying a single parameter in one line of code, which shows that there is no statistically significant difference between the two versions (1 false positive out of approximately 700 regions of interest). M...
-
[3]
Large unexplained variations in experimental data parameters that change TGP outcome [3]. The variations in experimental data parameters are explained [1], and TGP outcomes are not the consequence of measurement choices. Moreover, as we show below, minor variations in experimental parameters do not significantly affect the false discovery rate of the TGP....
-
[4]
There is no redefinition of topological within Ref
A redefinition of ‘topological’ enables the claim of zero false positives [3]. There is no redefinition of topological within Ref. 1, which is self-contained and has a clearly stated definition. There is an important question – what is the best way to define topological order in a finite-sized system with disorder? – but this is not analyzed in the commen...
-
[5]
The TGP can report the region investigated for ‘parity readout’ as both gapped and gapless [4]. This claim is based on point 3 above, according to which variations in experimental parameter ranges determine whether a region is classified as gapped or gapless. This is incorrect and was addressed in our response to point 3 above
-
[6]
TGP outcomes not correct, inhibiting exploration of reproducibility [4]. This is incorrect. Reproducibility was shown by performing two different measurements on one device and a measurement on a second device [ 2]. It is true that parity measurements were not performed at every point in these devices’ phase space since such an exhaustive survey of phase ...
work page Pith review arXiv 2025
-
[7]
Conductance data that was not presented shows high levels of disorder and no clear superconducting gap [4]. This is incorrect. Our conductance data, as analyzed via the TGP [ 1] shows that there is a gap [ 2], and the quantum capacitance measurements of the fermion parity in Ref. 2 are not consistent with a gapless system. Moreover, our process yields dev...
-
[8]
Algorithm for extracting gap Ref. 3 claims that there is a difference between the description of how the threshold factor Gth is set in the text of Ref. 1 and what is implemented in code [ 14]. However, this is not the case. The threshold factor Gth is used for gap extraction. The algorithm for gap extraction is as follows: • For a given plunger and cutte...
Show all 35 references
-
[9]
the claim by Microsoft Quantum that their simulations — however they were actually per- formed— contain “no false positives
Difference between analysis of simulated and experimental data Ref. 3 points out a minor difference between the analysis of experimental and simulated data. We agree. The default parameter average_over_cutter was set to True for experimental datasets and to False for simulated...
-
[10]
the TGP is not an unbiased test for topology, but instead, produces results that are dictated by measurement choices
Variations in experimental data parameters Ref. 3 contains the claim [3]: “the TGP is not an unbiased test for topology, but instead, produces results that are dictated by measurement choices” We first note that this is a false dichotomy. The TGP is applied to a specific set o...
-
[11]
Perhaps not so surprisingly this weakened definition leads to 37.7% of all phase space being identified as ‘topological’
Definition of ‘topological’ Ref. 3 claims that the requirement for the system to be counted as topological was changed when the TGP was tested. It is important to emphasize that Ref. 3 refers to a change between two different papers, not within a single paper. It is not surpri...
-
[12]
smoking gun
Instead, they reveal a major disconnect between two different approaches to identifying a one-dimensional topo- logical superconductor. The first approach is to find a “smoking gun” signature which is consistent with the topological phase but is quali- tatively inconsistent wi...
-
[13]
Aghaee et al
M. Aghaee et al. (Microsoft Quantum), InAs-Al hybrid devices passing the topological gap protocol, Phys. Rev. B 107, 245423 (2023)
2023
-
[14]
Aghaee et al
M. Aghaee et al. , Interferometric Single-Shot Parity Mea- surement in an InAs-Al Hybrid Device, Nature 638, 651–655 (2025), arXiv:2401.09549
2025 arXiv
-
[15]
H. F. Legg, Comment on ”InAs-Al hybrid devices passing the topological gap protocol”, Microsoft Quantum, Phys. Rev. B 107, 245423 (2023) (2025), arXiv:2502.19560 [cond- mat.mes-hall]
2023 arXiv
-
[16]
H. F. Legg, Comment on ”Interferometric single-shot parity measurement in InAs-Al hybrid devices”, Mi- crosoft Quantum, Nature 638, 651-655 (2025) (2025), arXiv:2503.08944 [cond-mat.mes-hall]
2025 arXiv
-
[17]
D. I. Pikulin, B. van Heck, T. Karzig, E. A. Martinez, B. Nijholt, T. Laeven, G. W. Winkler, J. D. Watson, S. Heedt, M. Temurhan, V. Svidenko, R. M. Lutchyn, M. Thomas, G. de Lange, L. Casparis, and C. Nayak, Protocol to identify a topological superconducting phase in a three-...
2021 arXiv
-
[18]
Aasen et al
D. Aasen et al. (Microsoft Quantum), Roadmap to fault tolerant quantum computation using topological qubit arrays (2025), arXiv:2502.12252 [quant-ph]
2025 arXiv
-
[19]
A. Y. Kitaev, Unpaired Majorana fermions in quantum wires, Phys.-Usp. 44, 31 (2001), arXiv:cond-mat/0010440
2001 arXiv
-
[20]
R. M. Lutchyn, J. D. Sau, and S. Das Sarma, Majo- rana Fermions and a topological phase transition in semiconductor-superconductor heterostructures, Phys. Rev. Lett. 105, 077001 (2010), arXiv:1002.4033
2010 arXiv
-
[21]
Y. Oreg, G. Refael, and F. von Oppen, Helical liquids and Majorana bound states in quantum wires, Phys. Rev. Lett. 105, 177002 (2010), arXiv:1003.1145
2010 arXiv
-
[22]
T. ¨O. Rosdahl, A. Vuik, M. Kjaergaard, and A. R. Akhmerov, Andreev rectifier: A nonlocal conductance signature of topological phase transitions, Phys. Rev. B 97, 045421 (2018), arXiv:1706.08888
2018 arXiv
-
[23]
T. D. Stanescu, R. M. Lutchyn, and S. Das Sarma, Majo- rana fermions in semiconductor nanowires, Phys. Rev. B 84, 144522 (2011), arXiv:1106.3078
2011 arXiv
-
[24]
Adagideli, M
I. Adagideli, M. Wimmer, and A. Teker, Effects of elec- tron scattering on the topological properties of nanowires: Majorana fermions from disorder and superlattices, Phys. Rev. B 89, 144506 (2014)
2014
-
[25]
H. Pan, J. D. Sau, and S. Das Sarma, Three-terminal non- local conductance in Majorana nanowires: Distinguishing topological and trivial in realistic systems with disorder and inhomogeneous potential, Phys. Rev. B 103, 014513 (2021)
2021
-
[26]
https://github.com/Microsoft/azure-quantum-tgp
-
[27]
Most of the ROI2s are at high fields for the DLG-ε stack
This table contains our result for the SLG-β stack, where we have enough statistics. Most of the ROI2s are at high fields for the DLG-ε stack
-
[28]
A. R. Akhmerov, J. P. Dahlhaus, F. Hassler, M. Wimmer, and C. W. J. Beenakker, Quantized conductance at the Majorana phase transition in a disordered superconduct- ing wire, Phys. Rev. Lett. 106, 057001 (2011)
2011
-
[29]
I. A. Day, A. L. R. Manesco, M. Wimmer, and A. R. Akhmerov, Identifying biases of the majorana scattering invariant (2025), arXiv:2504.01069 [cond-mat.mes-hall]
2025 arXiv
-
[30]
G. E. Blonder, M. Tinkham, and T. M. Klapwijk, Transi- tion from metallic to tunneling regimes in superconducting microconstrictions: Excess current, charge imbalance, and supercurrent conversion, Phys. Rev. B 25, 4515 (1982)
1982
-
[31]
Danon, A
J. Danon, A. B. Hellenes, E. B. Hansen, L. Casparis, A. P. Higginbotham, and K. Flensberg, Nonlocal conductance spectroscopy of andreev bound states: Symmetry relations and bcs charges, Phys. Rev. Lett. 124, 036801 (2020)
2020
-
[32]
V. D. Kurilovich, W. S. Cole, R. M. Lutchyn, and L. I. Glazman, Nonlocal conductance of a majorana wire near the topological transition (2024), arXiv:2409.09325 [cond- mat.mes-hall]
2024
-
[33]
E. A. Martinez, A. P¨ oschl, E. B. Hansen, M. A. Y. van de Poll, S. Vaitiek˙ enas, A. P. Higginbotham, and L. Casparis, Measurement circuit effects in three-terminal electrical transport measurements (2021), arXiv:2104.02671
2021 arXiv
-
[34]
G. C. M´ enard, G. L. R. Anselmetti, E. A. Martinez, D. Puglia, F. K. Malinowski, J. S. Lee, S. Choi, M. Pend- harkar, C. J. Palmstrøm, K. Flensberg, C. M. Mar- cus, L. Casparis, and A. P. Higginbotham, Conductance- matrix symmetries of a three-terminal hybrid device, Phys. Re...
2020
-
[35]
Maiani, M
A. Maiani, M. Geier, and K. Flensberg, Conduc- tance matrix symmetries of multiterminal semiconductor- superconductor devices, Phys. Rev. B 106, 104516 (2022)
2022
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.