{"id":"24fc1d37-9a8c-4c7c-8ce0-b112cdd76fa3","arxiv_id":"2411.17947","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A simulated neural-network classifier is claimed to cut axion-haloscope integration time by 5,000x, but the figure follows from an invalid accuracy-to-SNR conversion.","lead":"The authors simulate a haloscope cavity and amplifier noise, then train a one-neuron neural network to spot mock axion signals in the noise. They claim this cuts the integration time needed for detection by about 5,000 times, but that gain is an artifact of how they translate classifier accuracy into signal-to-noise ratio.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. 7's accuracy-to-SNR conversion is invalid and internally inconsistent: Table I maps 47.7% accuracy (worse than chance) to SNR_equiv=0.639 and a 3.4e5-fold time improvement, so the headline 5e3 factor is an artifact of the conversion.","rationale":"The central claim is the 5e3 integration-time improvement, and the single equation connecting FNN accuracy to a time factor is Eq. 7. All other components (BI-RME3D simulation, cavity design, FNN) are inputs to that equation; even if they are perfectly correct, the claim fails if the mapping is wrong. The reader's weakest_assumption identifies exactly this. I agree and would add one concrete point: the formula is not merely uncalibrated, it is self-refuting. Table I converts a below-chance accuracy of 47.7% into a positive equivalent SNR and a 3.4e5-fold improvement, which is a reductio ad absurdum. Since T=(z/SNR0)^2 at fixed accuracy, the conversion would also imply unbounded time improvements as SNR0 tends to zero. A proper treatment would report ROC curves, fix a false-alarm probability, and compare with the matched filter, which is optimal for a known signal in Gaussian noise. Without that, the headline factor is an artifact, and the REJECT verdict stands.","tokens_in":8707,"tokens_out":5505,"duration_ms":51053,"concrete_test":"For the 1000-average, Tsys=1.2 K case (SNR0=0.036, acc=99.3%), generate 10^4 independent test traces from the same simulation and compute the FNN's full ROC curve by varying the output threshold. Then compute the Neyman-Pearson matched-filter (likelihood-ratio) detector on the same traces. If the FNN's ROC does not coincide with or beat the matched-filter ROC at false-alarm probability 10^-4, the quantity SNRequiv=2.697 in Table I is not a valid SNR and Eq. 7 must be replaced by a proper detection-theoretic SNR definition; either way, this test separates the neural-network gain from the classical optimal-detection gain that any analysis could exploit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III defines T=(SNRequiv/SNR0)^2 with SNRequiv obtained as z_{(1+acc)/2}, the Gaussian quantile of a central probability. This is not a valid detection-statistics link. Binary classification accuracy is a threshold-dependent summary of a ROC curve; it does not map to a shift parameter of a Gaussian detector unless the classifier statistic is Gaussian and the threshold and class priors are specified. The internal inconsistency is decisive: with acc=47.7% (below chance), the formula assigns SNRequiv=0.639 and T=3.4e5, i.e., a detector worse than random is claimed to improve integration time by five orders of magnitude. Moreover, for any fixed accuracy, T=(z/SNR0)^2 diverges as SNR0 approaches zero, so arbitrarily large improvements can be manufactured by lowering SNR0. The claimed 5e3 improvement (SNR0=0.036, acc=99.3%, SNRequiv=2.697) therefore rests entirely on an unjustified conversion. Even if the simulations, BI-RME3D model, and FNN accuracy are correct, the paper provides no false-alarm rate, no detection-efficiency curve, and no comparison with the optimal matched filter, so the time improvement is not physically meaningful.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes using a feedforward neural network to detect axion signals in simulated haloscope data, claiming a 5·10^3 reduction in the integration time needed to reach a given signal-to-noise ratio. The simulation chain combines the BI-RME3D modal method with generated resonant and amplifier noise, and the network performs a boolean classification ('axion present' or 'noise only'). The central quantitative claim is obtained from Eq. (7), which converts classifier accuracy into an equivalent Gaussian SNR and then into a time-improvement factor T. The paper reports accuracies for two system temperatures and several averaging levels, with the largest claimed improvement being a factor of a few thousand at SNR0 = 0.036.","tokens_in":8890,"tokens_out":3652,"duration_ms":35109,"significance":"If the claimed time improvement were physically valid, it would be a practically important result for axion haloscope searches, since it would dramatically shorten the exposure time needed to reach a given sensitivity. The direction of using simulation-trained classifiers is worth exploring, and the use of BI-RME3D to generate realistic signal-plus-noise samples is a reasonable starting point. However, the headline result depends entirely on Eq. (7), whose accuracy-to-SNR conversion is not a valid detection-statistics link. The paper provides no false-alarm rate, no detection-efficiency curve, and no comparison with a matched filter. The internal inconsistency in Table I—where an accuracy below chance still produces a large positive time improvement—shows that the claimed sensitivity gain is an artifact of the conversion rather than a physical effect. The manuscript therefore does not establish its main claim.","major_comments":[{"comment":"The conversion of classifier accuracy to equivalent SNR is unjustified. SNRequiv is defined as the Gaussian quantile z_{(1+acc)/2}, but binary classification accuracy is a threshold-dependent summary that depends on the prior probability of the axion class and on the distribution of the classifier score. It does not, by itself, determine a detection sensitivity or a false-alarm rate. Without a derivation or an explicit link to detection statistics, T is not a physically meaningful time improvement.","section":"Section III, Eq. (7) and Table I"},{"comment":"Table I contains an internal inconsistency: with SNR0 = 0.0011 and accuracy = 47.7% (worse than chance), Eq. (7) gives SNRequiv = 0.639 and T = 3.4e5. A classifier that performs below chance cannot improve sensitivity, so the accuracy-to-SNR mapping is not merely ad hoc but self-contradictory. This row shows that the time-improvement factor is an artifact of the conversion.","section":"Table I, row '1 average'"},{"comment":"Because T = (SNRequiv/SNR0)^2, for any fixed accuracy T diverges as SNR0 tends to zero. The claimed factor of 5e3 is therefore not robust: the same formula yields T = 3.4e5 at SNR0 = 0.0011 with nearly chance accuracy, and arbitrarily large improvements could be manufactured by lowering the input SNR. This scaling undermines the significance of the reported numerical values.","section":"Section III, Eq. (7)"},{"comment":"The paper does not compare the neural network's performance with the optimal matched filter or with the standard radiometer analysis on the same simulated data. It also reports no false-alarm probability, no detection-efficiency curve, and no threshold analysis for the boolean decision. Without these elements, the equivalent SNR claimed for the network cannot be validated as a detection statistic.","section":"Section III"}],"minor_comments":[{"comment":"The input representation given to the neural network is not described in enough detail: it is unclear whether the network receives time-domain samples, power spectra, or complex phasors, and at what resolution. Specifying this would improve reproducibility.","section":"Section II"},{"comment":"The accuracy values are reported as single numbers without error bars or multiple training seeds. Since the accuracy is based on 1000 tests, binomial uncertainties are non-negligible, and the propagated uncertainty on T would be helpful.","section":"Figure 5 and Table I"},{"comment":"The text says 'with a SNR = 0.036 the accuracy of the neural network is 99.3%, close to an equivalent SNR of 3', but the corresponding value is 2.697, and the paper should state whether this is considered acceptable and why.","section":"Section III"}],"recommendation":"reject","confidential_remarks":"The central claim of the manuscript rests entirely on Eq. (7), and the accuracy-to-SNR conversion is not a valid detection-statistics link. The internal inconsistency in Table I (an accuracy below chance yielding a huge time improvement) indicates that the reported factor of 5e3 is an artifact of the conversion. The underlying simulation approach may be useful, but the main result cannot be salvaged without a new detection-theoretic analysis, which is beyond the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the electromagnetic modeling is credible, but the headline 5e3 time improvement is an artifact of an invalid accuracy-to-SNR conversion. The paper does not demonstrate that a neural network shortens haloscope exposures.\n\nWhat is worth keeping: the simulation chain is not hand-wavy. They use BI-RME3D to model the cavity, add resonant noise and a LNA, and include the axion lineshape from a Maxwell-Boltzmann distribution. The single-neuron FNN is simple, but the accuracy curves are plausible, and the paper is honest that training on real experimental data would eat the time savings. As a first application of ML to axion sensitivity, it has a legitimate seed value.\n\nThe soft spot is load-bearing. Equation (7) defines T = (SNRequiv/SNR0)^2, with SNRequiv computed as the Gaussian quantile corresponding to classifier accuracy. That mapping is not a detection statistic. Binary classification accuracy is a threshold-dependent summary of a ROC curve; it does not translate into a SNR unless the underlying score distribution, threshold, and priors are specified. The internal inconsistency is decisive: Table I lists accuracy 47.7% at SNR0=0.0011 and quotes SNRequiv=0.639 and T=3.4e5. A classifier worse than random is claimed to improve integration time by five orders of magnitude. Also, for any fixed accuracy, T diverges as SNR0 goes to zero, so arbitrarily large improvements can be manufactured by lowering SNR0. The specific 5e3 factor (acc=99.3%, SNR0=0.036, SNRequiv=2.697) therefore rests entirely on the conversion.\n\nThere is also no matched-filter baseline, no false-alarm rate or detection-efficiency curve, no error bars on the 1000-test accuracies, and no released code or data. A comparison against the optimal matched filter is the obvious missing control; without it, we do not know if the network adds anything beyond what standard spectral analysis already does.\n\nBottom line: the paper is a reasonable simulation study with an unsupported headline. It deserves a serious referee, mainly so the authors are pushed to replace Eq. 7 with a proper detection-statistics link and add a matched-filter comparison. I would not cite it in its present form, but I might bring it to a reading group as a cautionary example of what happens when a classifier accuracy is converted to a physical sensitivity.","headline":"A plausible simulation study whose headline 5e3 time improvement is an artifact of an invalid accuracy-to-SNR conversion.","tokens_in":9560,"tokens_out":3079,"would_cite":false,"duration_ms":28360,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a single-neuron feedforward network, trained on simulated haloscope noise, reaches in about half an hour a detection sensitivity that would otherwise take 100 days of integration — a 5,000-fold reduction in exposure…","keywords":["axion dark matter","haloscope","neural network","signal-to-noise ratio","thermal noise","BI-RME3D","feedforward network","low-noise amplifier"],"falsifier":"The claim would be settled by measuring the network's false-positive rate on real thermal noise that certainly contains no axion: at a claimed ${\\rm SNR}_{\\rm equiv}\\approx 2.7$, a Gaussian reading predicts a two-sided false-alarm probability of roughly 0.7%, so if the network flags 'axion present' far more often on pure noise data, the accuracy-to-significance mapping is invalid and the time improvement does not exist. A complementary check is to re-run the same training on genuine off-resonance spectra from a working haloscope with a synthetic axion line injected at ${\\rm SNR}=0.036$; if accuracy does not reach the simulated 99.3%, the improvement is an artifact of the simulated noise statistics.","tokens_in":105,"feed_emoji":"🧠","tokens_out":10789,"duration_ms":209159,"temperature":0.7,"pith_summary":"This paper is the first attempt to apply a neural network directly to the axion dark-matter search, with the aim of improving sensitivity rather than just automating data analysis. The authors simulate the dominant noise sources of a haloscope read-out chain — cavity resonant noise and the first low-noise amplifier — using an exact modal method, and train a feedforward network with a single hidden neuron to decide, from each recorded radio-frequency trace, whether an axion signal is present or only noise. They report that at a per-trace signal-to-noise ratio of about 0.036 the network classifies correctly in 99.3% of test cases, which they equate to the significance a conventional analysis would reach only after roughly 5,000 times longer integration. If correct, this means a measurement point that takes 100 days to reach signal-to-noise ratio 3 could instead be completed in about half an hour, or the same exposure could probe smaller axion-photon couplings. The practical value of the claim is that it applies to any haloscope experiment and can run in parallel with standard analysis pipelines.","feed_headline":"Neural network cuts axion-detection time 5,000-fold","feed_subtitle":"A single-neuron classifier on simulated haloscope noise reaches in 30 minutes what takes 100 days of averaging.","key_machinery":"The argument is carried by three pieces. First, the BI-RME3D modal method, an exact full-wave technique that turns the axion-photon coupling into circuit quantities: the axion is a current source $I_a$ added to a random Gaussian noise current $I_n$, and the detected voltage is the phasor $V_c=(I_a+I_n)/(Y_w+Y_c)$, so the delivered power $P_w$ carries magnitude and phase information. Second, the Dicke radiometer equation ${\\rm SNR}=(P_w/k_BT_{\\rm sys})\\sqrt{t/\\Delta\\nu}$, which sets how integration time $t$ converts into signal-to-noise ratio. Third, the feedforward neural network — one hidden neuron, one boolean output — trained on 5000 simulated examples per SNR and tested on 1000, whose accuracy (the fraction of correct 'axion present or not' decisions) is the quantity the whole claim hangs on. The time-improvement factor is then defined by $T=({\\rm SNR}_{\\rm equiv}/{\\rm SNR}_0)^2$, where ${\\rm SNR}_{\\rm equiv}$ is read off as the number of standard deviations of a standard Gaussian distribution that corresponds to the network's accuracy.","core_discovery":"The central claim is that a very simple neural network can pull an axion signal out of thermal noise far below the signal-to-noise ratio that conventional power averaging requires. The authors model the axion as a current source in a rectangular cavity with the Boundary Integral-Resonant Mode Expansion (BI-RME) modal method, add Gaussian fluctuations normalized to the cavity noise power $k_B T_{\\rm cav}\\Delta\\nu$, include the amplifier in the system temperature $T_{\\rm sys}$, and present the network with the resulting complex voltage phasors. For $T_{\\rm sys}=1.2$ K and 1000 averages, an input SNR of 0.036 yields 99.3% classification accuracy; the paper converts this accuracy into an equivalent Gaussian significance of about 2.7 standard deviations and, through the relation $T=({\\rm SNR}_{\\rm equiv}/{\\rm SNR}_0)^2$ based on the Dicke radiometer scaling, obtains a time improvement factor of roughly $5\\times 10^3$. The concrete consequence stated in the paper is that an experiment needing 100 days of integration to reach ${\\rm SNR}=3$ could reach that same sensitivity in about half an hour with the neural network, and that slightly larger input SNRs push accuracy toward 100%.","pith_inferences":["My reading is that the quoted 5,000-fold saving is a single-operating-point number: it is computed at ${\\rm SNR}_0=0.036$, and it would shrink substantially if the accuracy-to-Gaussian mapping were replaced by directly measured false-alarm and false-dismissal rates, which the paper does not report.","The paper trains and tests on simulated noise only; a natural extension is to validate the single-neuron classifier on recorded off-resonance noise from a running haloscope, where drifts, standing waves, and non-stationary amplifier noise — effects the authors list as future work — would test whether the simulated accuracy survives contact with real data.","Because the network outputs a yes/no decision rather than an estimator, the equivalent-SNR construction silently assumes optimal decision thresholds; building the full receiver-operating-characteristic curve across SNR values would give a threshold-independent measure of how much integration time is actually saved.","If the method holds up on real noise, it effectively converts part of the exposure-time budget into offline training cost, so the fair comparison is not 100 days versus 30 minutes but 100 days versus 30 minutes plus the one-time cost of building a faithful noise simulator."],"forward_implications":["A haloscope measurement point that needs 100 days of integration to reach ${\\rm SNR}=3$ could instead reach that sensitivity in about half an hour with the trained network.","With the same exposure time as today, a scan would reach better sensitivity and probe lower values of the axion-photon coupling constant $g_{a\\gamma\\gamma}$.","The technique is simple enough to run in parallel with standard spectrum-processing pipelines in any current axion experiment, either as the primary readout or as a cross-check.","Accuracy approaches 100% for only slightly larger input SNR, so the benefit grows quickly as the input signal strengthens.","The same method transfers to other ultra-feeble-signal searches, notably high-frequency gravitational wave haloscopes, where exposure time is limited by the duration of the signal."],"supporting_citations":[{"why":"Supplies the BI-RME3D modal method used to model the cavity and to compute the axion and noise currents as phasors.","marker":"[12]"},{"why":"Provides the machine-learning precedent for low-SNR detection whose accuracy behavior the authors say their results match.","marker":"[13]"},{"why":"Establishes the wide-band full-wave modal analysis of axion-photon coupling in microwave resonators that this paper adapts to include noise currents.","marker":"[29]"},{"why":"Supplies the radiometer equation that defines the signal-to-noise ratio and sets the exposure-time scaling on which the improvement factor is built.","marker":"[30]"},{"why":"Defines the haloscope detection principle — a resonant cavity converting axions to photons in a magnetic field — that the simulated read-out chain implements.","marker":"[10]"},{"why":"Provides the standard axion-search analysis procedure and signal lineshape that the simulation's Maxwell-Boltzmann frequency distribution follows.","marker":"[27]"}],"fun_headline_variants":["Neural net spots axions 5,000x faster","AI finds dark matter axion 5,000x quicker","Deep learning cuts axion detection time 5,000-fold","Network improves axion sensitivity 5,000x","Neural network detects axions at ultra-low SNR, 5,000x faster"],"cache_read_input_tokens":11520,"weakest_assumption_plain":"The entire improvement rests on treating the network's classification accuracy as a Gaussian detection significance — a 99.3% accuracy is taken to mean an equivalent signal-to-noise ratio of about 3 — because the claimed 5,000-fold time saving is just the square of that ratio divided by the input signal-to-noise ratio.","fun_headline_variants_meta":{"raw":{"variants":["Neural net spots axions 5,000x faster","AI finds dark matter axion 5,000x quicker","Deep learning cuts axion detection time 5,000-fold","Network improves axion sensitivity 5,000x","Neural network detects axions at ultra-low SNR, 5,000x faster"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000228,"raw_usage":{"total_tokens":1464,"prompt_tokens":921,"completion_tokens":543,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":454}},"tokens_in":537,"tokens_out":543,"duration_ms":5163,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:40:57.416799+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The claim would be settled by measuring the network's false-positive rate on real thermal noise that certainly contains no axion: at a claimed ${\\rm SNR}_{\\rm equiv}\\approx 2.7$, a Gaussian reading predicts a two-sided false-alarm probability of roughly 0.7%, so if the network flags 'axion present' far more often on pure noise data, the accuracy-to-significance mapping is invalid and the time improvement does not exist. A complementary check is to re-run the same training on genuine off-resonance spectra from a working haloscope with a synthetic axion line injected at ${\\rm SNR}=0.036$; if accuracy does not reach the simulated 99.3%, the improvement is an artifact of the simulated noise statistics.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the machine-learning precedent for low-SNR detection whose accuracy behavior the authors say their results match."}],"review_version":1}