{"id":"2cba1cce-9bb9-4f85-b904-7b465157401d","arxiv_id":"2509.00976","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Two neural-network emulators, 21cmLSTM and 21cmKAN, are reported to accurately mimic simulated global 21 cm signals and enable fast Bayesian parameter constraints for lunar far-side cosmology experiments.","lead":"This proceedings paper describes two neural-network emulators, 21cmLSTM and 21cmKAN, that rapidly reproduce simulated global 21 cm signals from the early universe. It argues these tools can speed up Bayesian analyses of future lunar far-side radio observations, with the KAN version training in about 10 minutes on a typical GPU.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The leap from same-simulation mock recovery to real lunar far-side observations is the weakest link: validation on 21CMGEM test signals with added Gaussian noise cannot certify unbiased physical constraints once foregrounds and instrument systematics enter.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: validation on mock signals drawn from the same simulation suite used for training is not sufficient evidence that real lunar far-side observations will yield unbiased constraints. This is a scope limitation, not an internal contradiction. The paper is a short IAU proceedings contribution that summarizes results from two companion papers; the quantitative evidence for the accuracy and speed claims lives in those peer-reviewed papers, so the central emulator-performance claim is plausible and likely correct. However, the 'reliable physical constraints' claim depends on the broader data-analysis pipeline, and the paper explicitly conditions on systematics being properly addressed. My proposed end-to-end test would directly probe whether the emulator plus the assumed noise model transfers to realistic observational conditions. Since this concern does not invalidate the emulator's accuracy or speed claims but does limit the strength of the 'unbiased constraints' statement as presented, the reader's CONDITIONAL verdict remains appropriate. No change to the verdict is needed.","tokens_in":5592,"tokens_out":4165,"duration_ms":56631,"concrete_test":"Build an end-to-end LuSEE-Night mock: inject a known global 21 cm signal from an independent simulator (e.g., ARES or 21cmFAST with parameters outside the 21CMGEM training grid) into a physically motivated foreground model (e.g., Galactic synchrotron with spatially varying spectral index) and apply a realistic instrument response (antenna beam, bandpass, calibration gain errors) plus noise at the expected lunar far-side level. Run the full signal-extraction and 21CMKAN nested-sampling pipeline on 50-100 noise realizations and test whether the 68% credible intervals contain the true parameters in approximately 68% of runs and whether the posterior mean is biased by more than the statistical uncertainty. If coverage fails, the emulator is not the bottleneck; the validation protocol is.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 21CMLSTM/21CMKAN enable 'quick and reliable physical constraints' for lunar far-side observations rests on mock-signal validation (Figure 3), in which 21CMKAN fits three 21CMGEM test signals with added Gaussian noise of 5, 10, and 25 mK. Because 21CMGEM is the same simulation suite used to generate training data, this establishes interpolation accuracy within one simulation model, not transfer to the real sky. The paper itself flags this in Section 1: 'assuming systematic effects are properly addressed.' Real lunar far-side observations will include foreground residuals, antenna beam chromaticity, calibration errors, and potentially non-white noise; none of these enter the mock-signal tests. If those systematics are imperfectly removed, the emulator's posterior will be biased regardless of emulation accuracy. Additionally, the posterior-recovery demonstration uses only three noise realizations, which is insufficient to establish frequentist unbiasedness or coverage over the prior volume. The load-bearing claim that needs support is not simply 'the emulator is fast and accurate on 21CMGEM' but 'the full analysis pipeline with this emulator produces unbiased physical constraints on data that resemble LuSEE-Night observations.' The paper's own caveat marks this as an acknowledged limitation rather than an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This proceedings paper describes two neural-network emulators for the global 21 cm signal: 21CMLSTM, a recurrent LSTM-based emulator, and 21CMKAN, a Kolmogorov-Arnold-network-based emulator. The authors claim that 21CMLSTM is the most accurate global 21 cm emulator to date, leveraging the temporal correlation of the signal, and that 21CMKAN achieves comparable accuracy while training about 75 times faster and evaluating about 5 times faster. They present an error histogram comparing these emulators with 21CMVAE, and a posterior-recovery demonstration in which 21CMKAN fits three 21CMGEM mock signals with added Gaussian noise of 5, 10, and 25 mK. The paper concludes that the emulators enable quick and reliable physical parameter constraints for upcoming lunar far-side experiments, with 21CMKAN able to complete end-to-end training and inference in under 30 minutes.","tokens_in":5843,"tokens_out":3745,"duration_ms":48287,"significance":"If the stated performance is correct, these emulators would be practically valuable for Bayesian inference pipelines with expensive physical simulations, particularly for LuSEE-Night and FarView. The manuscript's strengths are its clear motivation, the public availability of both emulator codes, and the direct comparison to an existing community emulator (21CMVAE). The qualitative architectural argument for using LSTM and KAN for the global 21 cm signal is reasonable. However, the evidence presented in this manuscript alone is thin: the accuracy comparison is qualitative, the unbiased-constraints claim rests on only three noise realizations, and all validation is internal to the 21CMGEM simulation suite. The paper is best read as a summary of the companion papers (Dorigo Jones et al. 2024, 2025), and the reader must consult those for the details needed to verify the central claims.","major_comments":[{"comment":"The claim of 'unbiased physical parameter posterior distributions' is supported by only three mock signals with added Gaussian noise. Three noise realizations are insufficient to establish unbiasedness or correct coverage over the prior volume. Please report quantitative bias metrics (e.g., mean/median difference between true and posterior values), credible-interval coverage on a larger held-out test set, and statistics over many noise realizations. Alternatively, rephrase the claim as 'recovered input parameters for three example signals.'","section":"Section 2 / Figure 3"},{"comment":"The validation is entirely internal to the 21CMGEM simulation suite: training and test signals come from the same simulator. Real lunar far-side observations will include foreground residuals, beam chromaticity, calibration errors, and possibly non-white noise. The manuscript's own caveat 'assuming systematic effects are properly addressed' acknowledges this limitation. The title's 'Reliable Physical Constraints' and the abstract's 'unbiased physical parameter constraints' overreach. Either add a systematics-inclusive validation (e.g., foreground-plus-instrument simulations) or explicitly scope the claims to noise-only mock observations.","section":"Section 1 / Figure 3"},{"comment":"The accuracy comparison is qualitative. The text states that 21CMLSTM is 'the most accurate emulator' and that 21CMKAN has 'similar accuracy,' but no numerical error statistics are given in this manuscript. The speed figures (3.7 ms, 75x, 5x) are point estimates with no uncertainty. Please report the mean/median and scatter of the relative RMS errors, training times, and evaluation times, or cite specific tables in the companion papers, so the headline claims can be checked without rereading the full papers.","section":"Section 2 / Figure 2"}],"minor_comments":[{"comment":"The code names are typeset inconsistently: '21CMLSTM' vs '21 CMLSTM' and '21CMKAN' vs '21 CMKAN'. Please use \texttt consistently throughout.","section":"Abstract and Section 2"},{"comment":"The red histogram is described as 'approximate error' for 21CMVAE. Please specify the source of the approximation (e.g., digitized values from Bye et al. 2022) and how much uncertainty that introduces.","section":"Figure 2"},{"comment":"The statement that 21CMKAN can complete 'end-to-end training and inference in under 30 minutes' should specify what is included (training, nested sampling, number of live points, hardware) and ideally report a range over runs.","section":"Section 2"},{"comment":"The factors '75 times faster' and '5 times faster' depend on hardware and stopping criteria. Please state the GPU model and whether the comparison uses matched training epochs, early stopping, or target accuracy.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings contribution that largely summarizes the companion papers. The central claims are likely correct given the detailed studies, but this manuscript, as a standalone, does not provide enough evidence for the strength of its abstract claims. The authors are honest about the systematics caveat, but the abstract and title should be aligned with the actual scope of the validation. If the venue expects concise summaries, adding a few numbers and a coverage statement would make the paper self-contained enough."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on this one. It's a short IAU proceedings piece that summarizes the authors' own 21cmLSTM and 21cmKAN emulators. If you're expecting new results, there aren't any--the performance numbers and figures come from Dorigo Jones et al. 2024 and 2025. What it does well is explain, in plain terms, why an LSTM and a KAN are natural fits for the global 21 cm signal, and it points to public code. That's genuinely useful for someone entering the field or scoping an analysis pipeline.\n\nThe central claim--that these are the most accurate/fastest emulators--is plausible, and the companion papers are peer-reviewed. But this manuscript itself gives you no numbers to check: no error statistics, no architecture details, no training curves. It's a pointer, not evidence. For a proceedings paper that's acceptable, as long as the referee verifies the claims actually appear in the cited papers.\n\nThe bigger soft spot is the validation story. The mock-signal recovery in Figure 3 uses test signals drawn from the same 21CMGEM simulation suite that made the training data, with three noise realizations. That establishes interpolation accuracy within one simulation model, not transfer to a real lunar far-side observation. The authors do flag this in Section 1 ('assuming systematic effects are properly addressed'), so it's an acknowledged limit rather than an overreach. Still, a reader should not come away thinking the emulator alone guarantees reliable constraints on LuSEE-Night data. The foreground/beam/calibration pipeline is separate, and the paper doesn't test it.\n\nAll that said, the work is honest and the code is public. The citations are appropriate, and the self-citation isn't a red flag here because the claims are about the authors' own released tools. I'd send this to a referee for a proceedings volume: the referee's job is light but real--check that the companion papers support the performance claims and that the caveat about systematics stays visible. For a regular research journal I'd expect more self-contained evidence. For the community, this is a useful overview, not a landmark.","headline":"A useful proceedings summary of the authors' own 21 cm emulators, not new science; the validation is honest but internal to 21CMGEM, so the real payoff depends on a careful foreground pipeline.","tokens_in":6424,"tokens_out":2845,"would_cite":false,"duration_ms":34721,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper demonstrates that two neural-network emulators—one LSTM-based, one KAN-based—can replace slow physics simulations of the global 21 cm signal and recover unbiased cosmological parameters from mock lunar far-side observations, with","keywords":["global 21 cm signal","neural network emulator","LSTM","Kolmogorov-Arnold Network","Bayesian inference","cosmic dawn","epoch of reionization","lunar far-side cosmology"],"falsifier":"Take a set of global 21 cm spectra generated by a different simulation code or by an independent instrument simulator that includes calibration errors and residual foregrounds, then run the published 21CMKAN training and nested-sampling pipeline on them; if the recovered posterior distributions exclude the true parameters by more than the expected statistical scatter, the unbiased-constraint claim is falsified.","tokens_in":5442,"feed_emoji":"🌙","tokens_out":8688,"duration_ms":98904,"temperature":0.7,"pith_summary":"The paper makes the case that emulation, not simulation, should carry the computational load of global 21 cm cosmology inference. It introduces two neural-network emulators that map physical model parameters to the sky-averaged 21 cm brightness-temperature curve: an LSTM network that exploits the signal's smooth evolution across frequencies, and a KAN network that learns the curve's functional form directly. The KAN version matches the LSTM's emulation accuracy while training 75 times faster and evaluating 5 times faster, so a complete training-plus-Bayesian-inference run takes under 30 minutes. Fitted to mock signals representing upcoming lunar far-side observations, both emulators recover unbiased physical parameters. The speed matters because every competing theoretical model of the early universe needs its own emulator trained before its parameters can be constrained.","feed_headline":"AI emulator makes 21 cm cosmology fits 75 times faster","feed_subtitle":"A KAN-based model matches the top LSTM emulator's accuracy and finishes training and inference in under 30 minutes.","key_machinery":"The global 21 cm signal is a 1D function of frequency, so the key machinery is choosing network architectures that respect that structure. 21CMLSTM is a recurrent Long Short-Term Memory network whose gated memory cells propagate information across adjacent frequency bins, capturing temporal correlation; its accuracy comes at the cost of sequential training. 21CMKAN is a Kolmogorov-Arnold Network, which replaces fixed node activations with learned activation functions on the edges and sums them at nodes; it explicitly learns the smooth functional transformations of the 21 cm curve with far fewer parameters. The emulators act as surrogate models in nested-sampling likelihood calls, converting","core_discovery":"The central claim is that 21CMLSTM and 21CMKAN provide state-of-the-art speed and accuracy for emulating the global 21 cm signal, with 21CMKAN reducing the training bottleneck to minutes. The paper demonstrates this by training on a suite of global 21 cm spectra generated by a semi-numerical simulation, comparing emulation error across trials against an existing variational-autoencoder emulator, and performing nested-sampling fits of mock signals with added 5, 10, and 25 mK noise. In those fits the posteriors are unbiased and the emulator's signal realizations bracket the true signal within 1σ. Because the emulators are publicly available and architecture-agnostic to the underlying signal mo","pith_inferences":["A stress test not reported in the paper: train on one simulation suite and attempt to recover parameters from signals generated by an independent 21 cm code or with a different foreground model; if posteriors shift, the unbiased claim is specific to the training distribution.","The KAN architecture's speed and functional-learning design likely transfer to other smooth 1D summary statistics in cosmology and astrophysics, such as galaxy luminosity functions or matter power spectra, where expensive simulations currently bottleneck inference.","If the mock-noise validation holds with real instrument calibration, real-time Bayesian fits during lunar observations become feasible, allowing observing strategy to adapt to the data.","A direct comparison on public data against the previous variational-autoencoder emulator is already suggestive; an end-to-end reanalysis of existing ground-based 21 cm measurements would test whether the unbiased-recovery performance persists under unmodeled foregrounds."],"forward_implications":["Upcoming lunar far-side radio experiments can run end-to-end Bayesian parameter constraints in under 30 minutes per model, making emulator training a non-bottleneck step in the analysis pipeline.","The unbiased-recovery result at noise levels from 5 to 25 mK means the emulators are positioned for the signal-to-noise conditions expected in global 21 cm measurements.","Because each theoretical model of cosmic dawn and reionization requires its own emulator, the 10-minute training time of 21CMKAN makes it practical to compare many models and parameterizations.","The transparent, interpretable KAN architecture allows users to inspect which physical parameters most strongly shape the signal, aiding model validation beyond point predictions."],"supporting_citations":[{"why":"Introduces 21CMLSTM, the LSTM-based emulator whose accuracy the paper adopts as the state-of-the-art baseline.","marker":"J. Dorigo Jones et al. 2024"},{"why":"Introduces 21CMKAN, the KAN-based emulator whose 75-times-faster training and similar accuracy are the paper's central result.","marker":"J. Dorigo Jones et al. 2025"},{"why":"Supplies the suite of global 21 cm mock signals used to train and test both emulators and the variational-autoencoder baseline against which accuracy is compared.","marker":"C. H. Bye et al. 2022"},{"why":"Defines the Kolmogorov-Arnold Network architecture that 21CMKAN is built on.","marker":"Z. Liu et al. 2025"},{"why":"Defines the LSTM recurrent architecture that 21CMLSTM uses to capture temporal correlations.","marker":"S. Hochreiter & J. Schmidhuber 1997"},{"why":"Extends LSTM with the forget gate mechanism the emulator's memory cells rely on.","marker":"F. A. Gers et al. 2000"},{"why":"Earlier global 21 cm emulator that the paper positions as a baseline for accuracy and speed comparisons.","marker":"A. Cohen et al. 2020"},{"why":"Another prior emulator that the paper cites as an existing surrogate model, defining the accuracy level to surpass.","marker":"H. T. J. Bevins et al. 2021"}],"fun_headline_variants":["AI emulator speeds lunar 21 cm cosmology fits 75x","Quick 21 cm constraints from AI emulator for Moon missions","KAN-based emulator nails 21 cm signal in under 30 min","Lunar far side 21 cm models get 75x faster with AI","AI makes global 21 cm analysis quick and reliable"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"That a mock signal drawn from the same simulation suite the emulator was trained on, with simple additive noise, is representative of real lunar far-side observations once foregrounds and instrument systematics are removed.","fun_headline_variants_meta":{"raw":{"variants":["AI emulator speeds lunar 21 cm cosmology fits 75x","Quick 21 cm constraints from AI emulator for Moon missions","KAN-based emulator nails 21 cm signal in under 30 min","Lunar far side 21 cm models get 75x faster with AI","AI makes global 21 cm analysis quick and reliable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000713,"raw_usage":{"total_tokens":3040,"prompt_tokens":737,"completion_tokens":2303,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":2212}},"tokens_in":481,"tokens_out":2303,"duration_ms":18470,"temperature":1.0,"reasoning_tokens":2212,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:59:37.761903+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of global 21 cm spectra generated by a different simulation code or by an independent instrument simulator that includes calibration errors and residual foregrounds, then run the published 21CMKAN training and nested-sampling pipeline on them; if the recovered posterior distributions exclude the true parameters by more than the expected statistical scatter, the unbiased-constraint claim is falsified.","supporting_citations":[{"cited_title":"2025, in The Thirteenth International Conference on Learning Representations","cited_arxiv_id":null,"evidence_quote":"Defines the Kolmogorov-Arnold Network architecture that 21CMKAN is built on."},{"cited_title":"1997, Long Short-Term Memory, Neural Computation, 9, 1735, doi:10","cited_arxiv_id":null,"evidence_quote":"Defines the LSTM recurrent architecture that 21CMLSTM uses to capture temporal correlations."},{"cited_title":"The Wrath of KAN: Enabling Fast, Accurate, and Transparent Emulation of the Global 21 cm Cosmology Signal","cited_arxiv_id":"2508.11752","evidence_quote":"Earlier global 21 cm emulator that the paper positions as a baseline for accuracy and speed comparisons."}],"review_version":1}