{"id":"3a8931f8-c38b-4945-95b6-07da418911b5","arxiv_id":"2505.22574","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A transformer-based emulator reproduces CAMB CMB TT, TE, and EE power spectra within cosmic variance errors across a wide Lambda-CDM parameter space, with outlier fractions below 10% for future survey configurations.","lead":"This paper trains transformer-based neural networks to emulate CMB power spectra computed by the CAMB code, achieving accuracy below cosmic variance limits across a broad Lambda-CDM parameter range. The work shows that attention mechanisms reduce outlier predictions compared to standard MLP emulators, making them viable for next-generation CMB surveys like CMB-S4 and CMB-HD.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation against CAMB AccuracyBoost=1.5 is not sufficient: Appendix A shows this reference is numerically unstable in the high-A_s region, and the testing cut log(10^10 A_s)<3.5 excludes exactly that region, so the claimed large-volume precision is not demonstrated against a more accurate CAMB…","rationale":"The reader's weakest assumption identifies the most load-bearing concern: the accuracy reference itself is not stable. Every headline number in the abstract and conclusion is a comparison to CAMB with AccuracyBoost=1.5, and Appendix A demonstrates that this setting deviates from a more accurate setting under the cosmic-variance metric, especially at high A_s. The post-hoc testing cut on log(10^10 A_s) removes the problematic region, so the claimed outlier fractions do not cover the full prior volume the paper says it emulates. If the AB=1.5 artifacts are systematic, the transformer could memorize them, making the validation circular with respect to the chosen CAMB settings. This concern is more fundamental than the CMB-HD multipole-range limitation, because it affects the interpretation of the accuracy metric for all experiments, not just the coverage of one forecast. It is also not addressed by the ACT DR6 test, which uses the same reference. A direct test against AccuracyBoost=2.5 over the full A_s prior would settle the issue. The rest of the paper is internally consistent: the loss-function and architecture comparisons are careful, the interpolation test to low temperature is reassuring, and the runtime claims are plausible. The method is promising, but the headline precision claim needs validation against a more accurate CAMB setting over the full claimed prior. I therefore agree with the reader's conditional verdict and see no reason to move it.","tokens_in":34656,"tokens_out":5262,"duration_ms":64117,"concrete_test":"Generate a test set from the same Ttest=128 Gaussian prior (and the full uniform prior without the A_s cut) using CAMB with AccuracyBoost=2.5 (or 3.0) and all other settings fixed; evaluate the already-trained baseline transformer and recompute <Delta chi^2>_med and the outlier fraction for log(10^10 A_s)>3.5 and for the full sample. If the outlier fraction exceeds 10% or the median Delta chi^2 rises by more than a factor of two relative to the AB=1.5 validation, the claimed precision is tied to the AB=1.5 reference rather than to the underlying physics. As a complementary check, compute the correlation between the emulator's error (vs AB=1.5) and the AB=1.5-vs-AB=1.8 difference vector; a strong correlation indicates the model has absorbed the numerical artifacts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the attention-based emulator matches CAMB to within cosmic variance over a large LCDM volume, with outlier fractions below 10% for Planck, SO, S4, and CMB-HD. Every quantitative validation is referenced to CAMB outputs generated with AccuracyBoost=1.5. Appendix A and Figure 14 show that this setting deviates from AccuracyBoost=1.8 by amounts that are visible under the cosmic-variance covariance, with deviations growing as log(10^10 A_s) increases. To avoid the most unstable region, Section III A imposes a post-hoc testing cut of log(10^10 A_s)<3.5, although the training prior extends to 3.6 (uniform) or 4.0 (Gaussian). This is load-bearing for two reasons. First, the headline outlier fractions and the 'maximized volume of applicability' are demonstrated only on a restricted testing prior, not on the full volume claimed in the abstract. Second, if the AB=1.5 artifacts are systematic rather than random, the transformer can learn them; the paper itself warns in Appendix A that the emulator 'will attempt to learn these numerical instabilities.' In that case the emulator would look accurate against AB=1.5 but biased relative to a more accurate Boltzmann code or the true sky, undermining the claim that it meets the precision criteria of future CMB experiments. The ACT DR6 test in Appendix B does not resolve this, because it uses the same CAMB reference. Thus the reported <10% outlier rates are conditional on the accuracy of the AB=1.5 reference, and that reference is shown to be unreliable in part of the claimed parameter space.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents an extensive comparison of neural-network emulators for CAMB TT, TE, and EE power spectra over a broad flat-ΛCDM prior, focusing on a dot-product-attention transformer baseline with ResMLP blocks, the H(x) activation, the L4 loss, and an A_s e^{-2τ} rescaling pre-processing step. Accuracy is measured by the median Δχ² and by the outlier fraction with Δχ²>0.2 relative to CAMB, using a cosmic-variance covariance and forecast/real covariances for Planck, Simons Observatory, CMB-S4, and CMB-HD, as well as ACT DR6 chains. The central claim is that with roughly 2×10^5–4×10^5 training points the attention-based emulators achieve outlier fractions below 10% and can therefore replace CAMB in real MCMC analyses without importance-sampling corrections.","tokens_in":35013,"tokens_out":6735,"duration_ms":80827,"significance":"If the claims hold, this is a useful and timely contribution: it provides a systematic architecture comparison, demonstrates that dot-product attention reduces outlier fractions relative to ResMLP-only models at fixed training-set size, evaluates several loss functions and pre-processing choices, and includes validation on Planck and ACT DR6 chains as well as on forecast covariances for next-generation experiments. The interpolation test down to T=1 (Section IV G, Figure 13) and the small outlier fractions on real chains are genuine strengths. The main caveat is that all quantitative validation is referenced to CAMB with AccuracyBoost=1.5, which the paper itself shows has numerical artifacts above cosmic variance in the high-A_s region, and the testing sets exclude part of that region; this tempers the strong claim of a 'maximized volume of applicability' but does not eliminate the practical value of the emulator for posterior volumes near the current fiducial cosmology.","major_comments":[{"comment":"The reference CAMB setting (AccuracyBoost=1.5) is shown in Figure 14 to deviate from AccuracyBoost=1.8 by amounts visible under the cosmic-variance covariance, with the deviations growing as log(10^10 A_s) increases, and Section III A excludes log(10^10 A_s)>3.5 from the testing set for exactly this reason. Since the emulator is trained on AB=1.5 outputs and Appendix A warns that the emulator will attempt to learn these numerical instabilities, the reported outlier fractions are conditional on a reference that is itself not accurate to cosmic variance over the full claimed training volume. The volume-of-applicability claim in the abstract and conclusion should either be re-scoped to the tested region or supported by re-validation with a higher AccuracyBoost setting over the full prior, including the high-A_s region.","section":"Appendix A, Section III A, Table I, Figures 5 and 15"},{"comment":"The Planck and ACT DR6 chain validations compare emulator outputs to CAMB outputs generated with the same AB=1.5 setting used for training, so they establish interpolation accuracy against that code version rather than against a more accurate Boltzmann calculation or the true sky. Given Appendix A, the paper should include a direct comparison of the AB=1.5 reference itself against a higher-accuracy CAMB or CLASS setting over the posterior volume of those real-data chains, so that reference error is separated from emulator error and cannot be absorbed into the emulator in a way that biases real-data analyses.","section":"Appendix B, Figure 16"},{"comment":"The headline result that a transformer reaches ê(Δχ²>0.2)<10% with approximately 4×10^5 training points is reported for testing sets drawn at Ttest=128 with the A_s cut discussed above, except that the caption of Figure 5 states Ttest=256. Because the abstract and conclusion quote the 4×10^5 point as the threshold, the testing conditions used for that number should be stated consistently and exactly in the main text, and the corresponding outlier fraction should be quoted under the same conditions.","section":"Section IV B and Section IV G"}],"minor_comments":[{"comment":"The Figure 5 caption says the model was tested with Ttest=256, while the text in Section IV and Figure 6 use Ttest=128; please reconcile this inconsistency.","section":"Figure 5 and Section IV B"},{"comment":"The text says a stringent cut of log(10^10 A_s)<3.5 is applied to the testing set, but Table I shows the Gaussian test set extends to log(10^10 A_s)=4.0 and only the uniform test set is cut at 3.5; please clarify which cut applies to each sampling scheme.","section":"Section III A and Table I"},{"comment":"For the CMB-HD-like forecast, the validation is restricted to ℓ≤5000 because the emulator is trained only to ℓ=5000, whereas the experiment is planned to reach ℓ~20000; the abstract and main text should state that the CMB-HD claim applies only to the emulated multipole range.","section":"Appendix B, Figure 15"},{"comment":"The damping-tail pre-processing test is presented as preliminary, using only three free parameters, a ResMLP architecture, and 1.8×10^4 training points; this should be flagged clearly wherever it is cited in the main text so it is not read as part of the baseline transformer validation.","section":"Appendix F and Section IV A"},{"comment":"Please add a data and code availability statement for the trained models, training sets, and analysis scripts, as the reproducibility of the numerical comparisons would be greatly enhanced by releasing them.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent about its own limitations, and the main technical risk is that the emulator accuracy is measured against a CAMB setting that the authors themselves show to be numerically unstable in part of the claimed training volume. This is fixable by re-validating or re-scoping the volume claim, so I recommend major revision rather than rejection. The practical value of the emulator for analyses near the Planck/ACT posterior is likely to survive such a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a serious methodological paper and I think it deserves a real referee. The new piece is genuine: taking the self-attention transformer architecture from the weak-lensing and clustering emulators in Parts I and II and applying it to CMB TT, TE, EE spectra, with careful work on pre-processing, loss functions, activation functions, and sampling strategies. The outlier-focused validation, including the comparison against ResMLP at several training-set sizes and the tests on Planck and ACT DR6 chains, is more thorough than what most emulator papers do. The authors are also transparent about limitations, which I credit.\n\nThe soft spot is exactly the one the stress test flags, and it is real but not fatal. All validation is referenced to CAMB with AccuracyBoost=1.5. Appendix A shows that this setting differs from AccuracyBoost=1.8 by amounts visible under cosmic variance in the high-As region, and the testing set applies a cut at log(10^10 As)<3.5 to avoid that region. So the headline 'less than 10% outliers over a large volume' is demonstrated only on a restricted volume, not the full training prior. The ACT DR6 test does not fix this because it uses the same CAMB reference. This is not a hidden flaw; the authors state it plainly. But the abstract and conclusion state the claim more broadly than the evidence supports. A referee should ask for the emulator to be retrained or at least re-validated against AB=2.5 or another reference in the high-As region, or for the claim to be reworded to match the tested volume.\n\nTwo smaller soft spots. First, no code or trained emulator is released, which limits reproducibility and reuse; for an emulator paper that is a real cost. Second, the CMB-HD forecast stops at ell=5000 because that is the training range; the paper says so, but the abstract listing CMB-HD should carry that caveat more explicitly.\n\nWho is this for? Anyone building or using CMB emulators for MCMC, and people thinking about how to validate ML surrogates against reference code numerical accuracy. I would not be surprised if the main architecture choice becomes a standard benchmark. The citation pattern looks fair; prior CosmoPower and the earlier parts of this series are properly acknowledged.\n\nRecommendation: send it to peer review. It is a useful, honest contribution, and the CAMB-reference issue is addressable with targeted extra validation rather than a fundamental flaw.","headline":"A solid, honest extension of transformer emulators to CMB power spectra; the headline accuracy claim is slightly overbroad because the hardest part of the parameter volume is cut from the testing set, but the central methodological result holds.","tokens_in":35584,"tokens_out":1788,"would_cite":true,"duration_ms":26259,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["85A40","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Attention-based emulators reproduce CAMB's CMB power spectra within cosmic-variance errors out to multipole 5000 with a few hundred thousand training spectra.","keywords":["CMB power spectra","cosmic microwave background","neural network emulator","transformer architecture","attention mechanism","cosmic variance","CAMB","Lambda-CDM cosmology"],"falsifier":"Regenerate the validation power spectra with a higher CAMB accuracy setting (AccuracyBoost 2.5 or above) and recompute the outlier fraction for the same trained emulators, restricting to $\\log(10^{10}A_s)\\leq 3.5$ and $\\ell \\leq 5000$; if the fraction with $\\Delta\\chi^2 > 0.2$ rises above 10 percent, the claim as stated is falsified for that configuration.","tokens_in":34447,"feed_emoji":"🔭","tokens_out":14389,"duration_ms":142766,"temperature":0.7,"pith_summary":"The paper sets out to show that an emulator built on self-attention (transformer) layers can reproduce the CMB temperature and polarization power spectra that the Boltzmann code CAMB computes, with errors small enough that no importance-sampling correction is needed when the emulator replaces CAMB in a likelihood analysis. The target is demanding: for a cosmic-variance-limited measurement, the fraction of test cosmologies whose emulated spectra differ from CAMB by more than $\\Delta\\chi^2 = 0.2$ should stay below 10 percent, out to multipole $\\ell = 5000$ and across a wide $\\Lambda$CDM parameter volume. The paper finds that dot-product attention does this with roughly $2\\times10^5$ to $4\\times10^5$ training spectra for each of the four next-generation experimental configurations considered, and that the attention mechanism is the ingredient that suppresses the outlier tail relative to plain multilayer-perceptron architectures. If the claim holds, a costly bottleneck of Markov-chain cosmological inference, namely repeated Boltzmann-code evaluations taking minutes per spectrum, is replaced by a sub-second neural evaluation.","feed_headline":"Attention transformers emulate CMB spectra within cosmic variance","feed_subtitle":"A few hundred thousand training spectra keep outlier cosmologies below 10 percent from Planck to CMB-HD","key_machinery":"The central object is the scaled dot-product self-attention block: the multipole vector is split into $N$ channels of length $d$, linearly mapped to query, key, and value matrices $Q = W_Q X$, $K = W_K X$, $V = W_V X$; the block computes $\\mathrm{Softmax}(QK^T/\\sqrt{d})V$, so each output channel is a weighted mixture of all channels, with weights drawn from the similarity of the query and key vectors. This operation lets the network exploit correlations among different parts of the CMB spectrum, and the paper's comparison to ResMLP-only models attributes the suppression of the outlier tail to this mechanism. Around it, three supporting pieces carry much of the accuracy: dividing the spectra by $A_s e^{-2\\tau}$ to remove the dominant amplitude degeneracy; the learned activation $h(x)$ of Eq. (6), which interpolates between linear and sigmoid-gated behavior with trainable parameters; and the $L_4$ loss, the square root of the cosmic-variance $\\Delta\\chi^2$ computed on the rescaled spectra, which weights outliers linearly instead of quadratically. Tempered Gaussian sampling of the training cosmologies, with the training distribution at temperature $T_{\\rm train}=256$ and testing at $T_{\\rm test}=128$, provides the wide-but-physical parameter coverage.","core_discovery":"On the paper's own terms, the discovery is that the scaled dot-product attention mechanism -- the same operation used in language-model transformers, applied here by splitting the power-spectrum vector into 16 channels -- changes the scaling of emulator accuracy with training-set size. For every training-set size tested, the attention-based model yields a lower median $\\Delta\\chi^2$ and a lower fraction of outlier cosmologies than a ResMLP without attention; the ResMLP would need hundreds of thousands of additional CAMB spectra to reach the same outlier fraction. Combined with three supporting choices -- rescaling the spectra by $A_s e^{-2\\tau}$ before training, using the learned activation $h(x)$, and training on the loss $L_4 = \\langle\\sqrt{\\Delta\\tilde\\chi^2_{XY}}\\rangle$ evaluated on the rescaled spectra -- the transformer reaches the target of fewer than 10 percent outliers with $\\Delta\\chi^2 > 0.2$ for the Planck, Simons Observatory, CMB-S4, and CMB-HD configurations examined. The same conclusion is reached by a 1D convolutional architecture at the largest training sets, and the attention-free transformer performs nearly as well as dot-product attention, while three cheaper approximations to the attention matrix (linear, latent, and locality-sensitive hashing) perform worse.","pith_inferences":["Beyond the paper: because the reference spectra are CAMB outputs at AccuracyBoost 1.5, the claimed outlier fractions should be read as relative to that specific numerical target; retraining or retesting against a higher-accuracy setting could shift the required training-set size, especially for $\\log(10^{10}A_s)>3.5$ or high multipoles.","Beyond the paper: the $\\Delta\\chi^2=0.2$ threshold is a fixed number taken from prior survey practice; an experiment with much smaller error bars, such as a CMB-HD-class survey, may need a stricter threshold, which would likely push the required training-set size above the quoted $4\\times10^5$.","Beyond the paper: the demonstrated transfer to weak lensing and clustering suggests that a single attention-based emulator architecture, trained with a rescaled cosmic-variance loss, could serve as a common backend for multi-probe analyses, although the paper does not itself train a combined multi-probe data vector.","Beyond the paper: the symbolic-regression damping-tail rescaling in Appendix F gives about a factor of two improvement with only $1.8\\times10^4$ training points; adopting such analytic pre-processing as standard could reduce training-data costs further than the paper's headline numbers."],"forward_implications":["A cosmology pipeline can replace a CAMB evaluation costing about 100 seconds on nine CPU cores with a transformer emulator costing $0.01$--$0.1$ seconds on one core, without an importance-sampling correction step.","For Planck-like and Simons Observatory-like analyses, around $2\\times10^5$ training spectra suffice; for CMB-S4-like and CMB-HD-like precision, around $4\\times10^5$ are needed to hold the outlier fraction below 10 percent.","Tempered Gaussian sampling reaches the target with roughly three to five times fewer training spectra than uniform sampling, so the training-data bottleneck can be reduced by using a Fisher-informed, correlated sampling distribution.","The outlier suppression is architectural, not just data-driven: dot-product attention and its attention-free variant outperform linear, latent, and locality-sensitive-hashing attention, and attention outperforms a ResMLP of comparable size.","The same attention/loss recipe transfers to other probes: a $3\\times2$pt weak-lensing and galaxy-clustering emulator reaches the $\\Delta\\chi^2$-based threshold with $10^5$ training points, and the appendix extends the approach to supernova distances."],"supporting_citations":[{"why":"Computes the TT, TE, and EE power spectra that are the emulator's training and testing target.","marker":"[5]"},{"why":"Defines the scaled dot-product attention operation that is the central mechanism of the transformer blocks.","marker":"[26]"},{"why":"Introduces the self-attention-based emulator for optical shear that this series extends to CMB power spectra.","marker":"[20]"},{"why":"Supplies the ResMLP/transformer comparison, the tempered Gaussian sampling, and the loss-function framework used here.","marker":"[21]"},{"why":"Provides the earlier CMB emulator design, including the MSE loss and the H(x) activation this work adapts.","marker":"[16]"},{"why":"Introduces the learnable h(x) activation function that the paper finds outperforms Tanh.","marker":"[34]"},{"why":"Supplies the default CAMB accuracy settings, raised to Accuracy-Boost 1.5 for the reference spectra.","marker":"[45]"},{"why":"Provides the ACT DR6 likelihood and covariance used to validate the emulator against current data.","marker":"[44]"},{"why":"Supplies the fiducial Planck cosmology and the covariance/binning framework for the Planck-like forecast.","marker":"[30]"}],"fun_headline_variants":["Attention emulators hit CMB precision for future surveys","Attention-based emulators match cosmic variance for CMB","Attention nets cut CMB emulator outliers below 10%","Transformer attention tames CMB spectra emulation","Attention emulators meet CMB survey precision goals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy claim rests on treating the reference computations produced by the Boltzmann code CAMB with its numerical-accuracy knob set to 1.5 as the truth; if the emulator learns numerical artifacts that appear at that setting, especially at high $A_s$, its true error against a more accurate computation would be larger than the reported cosmic-variance-level error.","fun_headline_variants_meta":{"raw":{"variants":["Attention emulators hit CMB precision for future surveys","Attention-based emulators match cosmic variance for CMB","Attention nets cut CMB emulator outliers below 10%","Transformer attention tames CMB spectra emulation","Attention emulators meet CMB survey precision goals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1381,"prompt_tokens":1099,"completion_tokens":282,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":206}},"tokens_in":715,"tokens_out":282,"duration_ms":3573,"temperature":1.0,"reasoning_tokens":206,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:03:14.135154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate the validation power spectra with a higher CAMB accuracy setting (AccuracyBoost 2.5 or above) and recompute the outlier fraction for the same trained emulators, restricting to $\\log(10^{10}A_s)\\leq 3.5$ and $\\ell \\leq 5000$; if the fraction with $\\Delta\\chi^2 > 0.2$ rises above 10 percent, the claim as stated is falsified for that configuration.","supporting_citations":[{"cited_title":"Attention-based Neural Network Emulators for Multi-Probe Data Vectors Part I: Forecasting the Growth-Geometry split","cited_arxiv_id":"2402.17716","evidence_quote":"Introduces the self-attention-based emulator for optical shear that this series extends to CMB power spectra."},{"cited_title":"Attention-Based Neural Network Emulators for Multi-Probe Data Vectors Part II: Assessing Tension Metrics","cited_arxiv_id":"2403.12337","evidence_quote":"Supplies the ResMLP/transformer comparison, the tempered Gaussian sampling, and the loss-function framework used here."}],"review_version":1}