{"id":"53e9e2e8-58e1-42f7-be11-d8914efbb172","arxiv_id":"2608.06343","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A foreground-robust CMB lensing B-mode template combining SPT-3G lensing reconstructions with a Planck CIB tracer achieves residual lensing amplitude 0.48 over ell=20-200, the best delensing efficiency reported to date.","lead":"This paper builds maps of gravitational lensing distortion of the cosmic microwave background using SPT-3G and Planck data, and uses them to predict and subtract lensing-induced B-mode polarization, a major contaminant in searches for primordial gravitational waves. The authors report the most effective such delensing template so far, reducing residual lensing power to about 48% of its original level over degree scales.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Foreground-robustness claim rests on in-sample Agora validation: the GMVph hardening profile and the foreground-bias test share the same Agora tSZ model, so a mismatch between Agora and the real foreground sky is not excluded by the simulation tests.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the foreground-bias validation is in-sample with respect to the Agora foreground model used for the GMVph hardening profile, and the data difference tests have limited power. My review agrees with this assessment. The central claim of foreground robustness is essential to the paper's stated purpose of providing a template for upcoming BICEP delensing analyses. If the real foreground sky differs from Agora in profile shape or in its correlations with the lensing signal, the residual bias could exceed the claimed 0.1 sigma level, and the template would no longer be 'foreground-robust' in the strong sense asserted. The proposed concrete test, using an independent foreground simulation with a different tSZ profile, would directly test the generality of the hardening and bias validation. Since this concern is the same one identified by the reader, and the reader's conditional verdict already reflects the need for additional out-of-sample validation, no change to the verdict is required. I therefore recommend keeping the verdict as CONDITIONAL (equivalently, UNCHANGED relative to the reader's recommendation).","tokens_in":35408,"tokens_out":4062,"duration_ms":47948,"concrete_test":"Rerun the foreground-bias measurement of Fig. 9 for the GMVph + CIB template using an independent set of non-Gaussian foreground simulations with a different tSZ profile than Agora, for example WebSky or the Sehgal et al. simulations, while keeping the GMVph hardening kernel fixed to the Agora-based profile. If the normalized bias over 20 <= ell <= 200 remains below 0.1 sigma across these out-of-sample foreground models, the in-sample concern is resolved. If the bias varies substantially across models or exceeds 0.1 sigma for any model, the claim of foreground robustness must be qualified as model-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the GMVph + CIB template is foreground-robust, with residual bias below 10% of the statistical uncertainty on the template power spectrum (Abstract, Sec. IV E, Fig. 9), depends on validation against Agora simulations whose non-Gaussian foreground model is the same one used to define the profile-hardening kernel. Specifically, Sec. III A 1 b states that the GMVph estimator is hardened using a profile matching a modified tSZ power spectrum profile from Agora [19], and Sec. IV E measures the foreground-induced bias using Agora simulations that contain that same tSZ model. The validation is therefore in-sample: the estimator is tuned to the exact foreground statistics used to test it. If the real tSZ profile, or the correlation between foregrounds and the lensing signal, differs from Agora, the residual bias after hardening could be larger than the reported ~0.05 sigma average. The data-based difference tests (Fig. 8, bottom panel) do provide an out-of-sample check, but their error bars are large; the paper itself notes that the GMVph vs. PP consistency is 'given the size of the difference error bars.' A bias at the level of several tenths of sigma in the template auto-spectrum would not be excluded by these tests, and would matter for the upcoming BICEP delensing application. Thus the load-bearing assumption is that Agora's foreground model adequately represents the real sky for the purposes of profile hardening and bias estimation, and this assumption is not independently tested in the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper constructs CMB lensing B-mode templates for delensing primordial B-mode searches, using SPT-3G D1 E-mode maps and lensing-potential reconstructions (standard GMV, profile-hardened GMV, and polarization-only PP) combined with the Planck 545 GHz CIB map. The central result is A_lens^res = 0.48 over 20 <= ell <= 200 for the GMVph + CIB template, claimed to be the highest delensing efficiency to date, with foreground-induced bias below 10% of the template power-spectrum statistical uncertainty. The validation uses 500 Gaussian simulations and 10 Agora patches with realistic non-Gaussian foregrounds; data auto-spectra are consistent with Gaussian simulations with PTEs of 0.47-0.97.","tokens_in":35734,"tokens_out":10864,"duration_ms":115537,"significance":"If robust, this is a timely and important result for the BICEP/SPT delensing program. The paper's strengths are the realistic end-to-end pipeline, the explicit comparison of three estimators with different foreground immunities, the use of matched Agora realizations to isolate non-Gaussian foreground effects, and the detailed multipole-cut robustness tests. The PTEs from data-versus-simulation comparisons are healthy. The main weakness is that the headline foreground-robustness claim rests partly on an in-sample Agora validation and on a cross-spectrum quantity that is presented as a template-power-spectrum bias.","major_comments":[{"comment":"The claim that the GMVph + CIB template has residual foreground bias below 10% of the statistical uncertainty on the template power spectrum is not directly supported by the plotted quantity. Fig. 9 shows the Agora-versus-Gaussian difference in the template cross-spectrum with the input B field, C^{BLT B_in}, normalized by the standard deviation of the template auto-spectrum C^{BLT BLT}; the foreground bias in the template auto-spectrum itself could differ. Since the abstract and Section IV E quote the power-spectrum bias, please either report the Agora-based bias in C^{BLT BLT} for GMVph + CIB or restate the claim as applying to the cross-spectrum bias. The data PTE for the auto-spectrum in Fig. 5 is a different, though supportive, test.","section":"Abstract and Sec. IV E (Fig. 9)"},{"comment":"The foreground-bias validation is in-sample with respect to the foreground model. The GMVph hardening profile is defined using Agora's tSZ model, and the bias test uses 10 patches cut from the same single full-sky Agora realization. If the real tSZ profile, or the correlation between foregrounds and the lensing field, differs from Agora, the residual bias after hardening could exceed the quoted ~0.05 sigma average. The bottom panel of Fig. 8 provides an out-of-sample data check, but as the authors note, the difference error bars are large, so a bias of several tenths of a sigma in the template spectrum would not be excluded. I recommend adding sensitivity tests that vary the hardening profile or use an independent non-Gaussian foreground simulation, and in any case qualifying the Abstract and Conclusion to state that the sub-0.1 sigma bias is demonstrated under the Agora foreground model.","section":"Sec. III A 1 b and Sec. IV E"}],"minor_comments":[{"comment":"The sigma(r) forecast for BK18 is internally inconsistent. If the total sigma(r)=0.009 and the no-lensing sigma(r)=0.004, the lensing contribution in quadrature is sqrt(0.009^2 - 0.004^2) = 0.008, not 0.005; moreover, since A_lens^res is defined as a power ratio in Eq. (8), the delensed lensing sigma(r) should scale as sqrt(A_lens^res), not A_lens^res. Please redo this toy calculation and update the quoted 29% improvement.","section":"Sec. IV F"},{"comment":"The uncertainties on A_lens^res quoted in the text (0.021-0.025) are not shown in the table; including them would allow a proper comparison of the tracer variants.","section":"Table I"},{"comment":"The phrase 'the highest delensing efficiency lensing template to date' should be supported by explicit A_lens values for the previous BICEP/SPTpol [8] and ACT DR6 [14] templates, rather than left as an unquantified claim.","section":"Abstract and Sec. IV C"},{"comment":"The PTE labels such as 'PTE (No CIB): 0.68 (+ CIB): 0.96' are ambiguous; please clearly separate the no-CIB and with-CIB cases.","section":"Fig. 5 caption/labels"}],"recommendation":"major_revision","confidential_remarks":"The paper is publishable after the validation caveats are addressed. I am not recommending rejection because the central template construction is well done and the data-simulation agreement is strong. The in-sample nature of the Agora test is the main risk; if the authors cannot provide an independent foreground test, the paper should be reframed as demonstrating robustness under the Agora model."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you need to know: this is a well-executed delensing paper. The genuinely new piece is applying the profile-hardened GMV lensing reconstruction from SPT-3G D1 combined with a Planck CIB tracer to build a lensing B-mode template, reaching A_lens^res ~0.48 averaged over ell=20-200. That's a real step over previous templates, and the paper says it will feed into the upcoming BICEP delensing analysis, which is plausible.\n\nWhat it does well: the formalism is standard but clearly presented, and the Wiener filtering/tracer combination follows prior work (Smith et al., Yu et al., BICEP/SPTpol, ACT DR6) without overclaiming. The validation is genuinely thorough: data template auto-spectra match Gaussian-simulation means with PTEs from 0.47 to 0.97, and the robustness tests to multipole cuts are systematic. The PP estimator is a good foreground-immune cross-check. Appendix A shows auto- and cross-spectra agree within 3%, supporting near-optimal filtering. The sigma(r) impact estimate is carefully labeled as a toy scenario.\n\nNow the soft spots, in proportion. The main one is the in-sample nature of the foreground-robustness claim. The GMVph estimator is hardened using a profile from Agora's tSZ model, and the foreground-bias validation in Fig. 9 uses Agora simulations containing that same model. So the test is not independent. The paper does include data difference tests (Fig. 8), but those have limited power; a residual bias at the few tenths of sigma level in the template auto-spectrum is not excluded. I don't think this sinks the paper, but it needs to be acknowledged, and ideally the authors should test with a different foreground model or vary the hardening profile shape. Second, the 'highest delensing efficiency to date' claim is stated without a direct comparison table to prior delensing results; given different sky coverage and metrics, a table would let the reader verify it. Third, the paper relies on the companion O26 lensing paper that is still in prep, which makes reproduction harder, though that's normal for collaboration work. Also, the Agora validation uses only 10 patches from one full-sky realization, limiting the generality of the foreground statistics.\n\nOverall, the central A_lens^res result is credible: it's simulation-derived, but the data auto-spectra agree with simulations, so the metric is well founded. The citation pattern looks fine. This paper will be useful to people in B-mode delensing and the BICEP/SPT joint analysis. It deserves a serious referee. I'd send it out, but with a request that the authors address the in-sample validation and provide a proper comparison table.","headline":"Solid delensing paper with a real improvement in template efficiency; the foreground-robustness claim is undercut by in-sample Agora validation, but it deserves serious review.","tokens_in":36889,"tokens_out":2658,"would_cite":true,"duration_ms":27032,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A lensing template made from a profile-hardened CMB lensing map plus a cosmic infrared background tracer removes about half the lensing B-mode power, with foreground bias under a tenth of the statistical uncertainty.","keywords":["CMB lensing","B-mode polarization","delensing","primordial gravitational waves","foreground hardening","cosmic infrared background","tensor-to-scalar ratio","SPT-3G lensing"],"falsifier":"Recompute the GMVph-minus-polarization-only template difference bandpowers on several independent, deeper CMB maps; if the coherent offset exceeds 0.1$\\sigma$ of the template's statistical uncertainty, the foreground-bias claim is falsified, while a null result within the quoted error keeps it.","tokens_in":35214,"feed_emoji":"🔭","tokens_out":11676,"duration_ms":108400,"temperature":0.7,"pith_summary":"This paper constructs a template for the lensing-induced B-mode polarization of the cosmic microwave background (CMB) and shows that it can be made both highly efficient and safe against extragalactic foreground contamination. The template combines a lensing-potential map reconstructed from South Pole Telescope polarization data with a cosmic infrared background (CIB) map from the Planck satellite, using a profile-hardened quadratic estimator that suppresses foreground-induced biases. The authors find that this combined template leaves about 48 percent of the lensing B-mode power on degree scales, the lowest residual yet achieved, while keeping foreground-induced bias below 10 percent of the statistical uncertainty. This matters because lensing B-modes are currently the dominant noise floor for searches for primordial gravitational waves: removing half the lensing power directly sharpens the tensor-to-scalar ratio constraint.","feed_headline":"Lensing template cuts leftover B-mode power to 48 percent","feed_subtitle":"Foreground bias stays under 10 percent of the statistical error, sharpening the search for primordial gravitational waves.","key_machinery":"The load-bearing object is the gradient-order lensing template $B^{\\rm LT}_{\\ell m} = \\sum g^{\\rm EB}_{\\ell \\ell' L} E^{\\rm WF}_{\\ell' m'} \\phi^{\\rm WF}_{L M}$, a harmonic-space convolution of Wiener-filtered CMB E-modes with a Wiener-filtered lensing-potential tracer. The tracer is a combined map whose scale-dependent weights are chosen to maximize its correlation with the true lensing potential, and whose CMB part is a profile-hardened GMV quadratic estimator that subtracts a nuisance 'source' field built from an assumed contaminant profile. The CIB map supplies an external tracer that stays well correlated with lensing at high multipoles where the CMB reconstruction becomes noise dominated. The efficiency metric is the residual lensing amplitude $A_{\\rm lens}^{\\rm res}(\\ell) = 1-(\\rho^B_\\ell)^2$ when the Wiener filters are optimal, where $\\rho^B_\\ell$ is the correlation between the template and the true lensing B-mode field.","core_discovery":"On the paper's own terms, the central result is that a lensing B-mode template built from a profile-hardened global-minimum-variance (GMVph) CMB lensing reconstruction combined with a CIB tracer is both the most efficient delensing template to date and effectively free of foreground bias. In the fiducial configuration the template achieves a residual lensing amplitude $A_{\\rm lens}^{\\rm res} \\simeq 0.48$ averaged over $20 \\leq \\ell \\leq 200$, meaning roughly half of the lensing B-mode power is removed; adding the CIB tracer improves the residual by about 20–25 percent relative to the CMB-reconstruction-only template. The paper further argues, from simulations with realistic non-Gaussian foregrounds and from data difference tests, that foreground-induced bias in the GMVph + CIB template is below 10 percent of the statistical uncertainty, with an average level near 0.05$\\sigma$. Because the template spectra measured in data agree with the Gaussian simulation predictions, the authors take the simulation-derived delensing efficiency as a reliable forecast for real data.","pith_inferences":["If the real extragalactic foreground sky differs from the simulation model used to calibrate the hardening and the bias test—for instance in the tSZ profile shape or in correlations with the lensing signal—the sub-0.1$\\sigma$ bias bound could be optimistic; this is the main residual risk and will be tested as deeper data accumulate.","The same hardened-plus-external-tracer combination should carry over to other CMB surveys with deeper polarization maps, where the polarization-only cross-check becomes a genuinely powerful null test.","Because the CIB tracer contributes high-multipole lensing modes that CMB reconstruction misses, adding further large-scale-structure tracers (deeper CIB maps, galaxy lensing) could push $A_{\\rm lens}^{\\rm res}$ below 0.4.","A decisive future check is a GMVph-minus-polarization-only difference measured on several independent patches with enough integration to reach sub-0.1$\\sigma$ precision; a coherent non-zero difference would falsify the claimed foreground robustness."],"forward_implications":["The GMVph + CIB template removes about 52 percent of the degree-scale lensing B-mode power, leaving residual lensing at $A_{\\rm lens}^{\\rm res}\\simeq 0.48$, so a directly delensed map retains barely half the lensing contamination.","Adding the CIB tracer to the CMB-only reconstruction improves delensing efficiency by about 20–25 percent across all estimator choices considered.","Foreground-induced bias in the baseline template sits below a tenth of the statistical uncertainty, so the template can be used without paying a foreground-bias penalty in near-term delensing analyses.","Delensing with this template on the current leading B-mode dataset would cut the lensing contribution to $\\sigma(r)$ from about 0.005 to 0.0024, a roughly 29 percent reduction in the total uncertainty.","The validated pipeline—hardened CMB lensing reconstruction plus an external tracer—is the recipe for foreground-safe delensing in upcoming CMB experiments."],"supporting_citations":[{"why":"The latest tensor-to-scalar-ratio constraint motivates delensing and provides the $\\sigma(r)$ baseline used to project the improvement.","marker":"[1]"},{"why":"This earlier delensing demonstration on real data sets the previous template baseline that this work improves upon.","marker":"[8]"},{"why":"This work establishes the gradient-order harmonic-space template formalism, giving the coupling kernel used to build the lensing template.","marker":"[11]"},{"why":"This analysis shows that using lensed rather than unlensed E-modes in the template is preferable, justifying the E-mode input choice.","marker":"[13]"},{"why":"The SPT-3G D1 lensing reconstruction supplies the GMV, GMVph, and polarization-only lensing-potential maps used as internal tracers.","marker":"[15]"},{"why":"This work provides the multitracer Wiener-filter combination formalism used to merge the CMB reconstruction with the CIB tracer.","marker":"[18]"},{"why":"These multicomponent simulations with realistic non-Gaussian foregrounds are the validation set behind the foreground-bias measurements.","marker":"[19]"},{"why":"This is the 545 GHz cosmic infrared background map used as the external lensing tracer in the combination.","marker":"[21]"},{"why":"This study identifies extragalactic foreground contamination in temperature-based lens reconstruction, motivating the profile-hardened approach.","marker":"[28]"},{"why":"This paper develops the bias-hardened lensing formalism that the profile-hardened GMV procedure implements.","marker":"[32]"}],"fun_headline_variants":["Foreground-proof lensing template halves B-mode contamination","New lensing template removes half of lensing B-mode power","Delensing hits 52% with foreground bias under 10% error","Lensing template slashes B-mode leftovers to 48%","Robust lensing template cuts B-mode power by over half"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the single full-sky realization of non-Gaussian extragalactic foregrounds used in the paper's simulations—including the tSZ profile that defines the profile-hardened estimator—faithfully represents the real foreground sky's shape and its correlations with the lensing signal.","fun_headline_variants_meta":{"raw":{"variants":["Foreground-proof lensing template halves B-mode contamination","New lensing template removes half of lensing B-mode power","Delensing hits 52% with foreground bias under 10% error","Lensing template slashes B-mode leftovers to 48%","Robust lensing template cuts B-mode power by over half"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2808,"prompt_tokens":1136,"completion_tokens":1672,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":752,"completion_tokens_details":{"reasoning_tokens":1586}},"tokens_in":752,"tokens_out":1672,"duration_ms":11090,"temperature":1.0,"reasoning_tokens":1586,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:48:46.315768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the GMVph-minus-polarization-only template difference bandpowers on several independent, deeper CMB maps; if the coherent offset exceeds 0.1$\\sigma$ of the template's statistical uncertainty, the foreground-bias claim is falsified, while a null result within the quoted error keeps it.","supporting_citations":[{"cited_title":"5 GNILC is designed to separate Galactic thermal dust emission from CIB anisotropies by exploiting both their spectral and spatial differences","cited_arxiv_id":null,"evidence_quote":"The latest tensor-to-scalar-ratio constraint motivates delensing and provides the $\\sigma(r)$ baseline used to project the improvement."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This earlier delensing demonstration on real data sets the previous template baseline that this work improves upon."},{"cited_title":"Tristram, A","cited_arxiv_id":null,"evidence_quote":"This work establishes the gradient-order harmonic-space template formalism, giving the coupling kernel used to build the lensing template."},{"cited_title":"Kusaka, J","cited_arxiv_id":null,"evidence_quote":"This analysis shows that using lensed rather than unlensed E-modes in the template is preferable, justifying the E-mode input choice."},{"cited_title":"Grav- itational Lensing, Astronomy and Astrophysics641, A8 (2020)","cited_arxiv_id":null,"evidence_quote":"The SPT-3G D1 lensing reconstruction supplies the GMV, GMVph, and polarization-only lensing-potential maps used as internal tracers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This work provides the multitracer Wiener-filter combination formalism used to merge the CMB reconstruction with the CIB tracer."},{"cited_title":"Baleato Lizancos, A","cited_arxiv_id":null,"evidence_quote":"These multicomponent simulations with realistic non-Gaussian foregrounds are the validation set behind the foreground-bias measurements."},{"cited_title":"Omori and SPT-3G Collaboration, SPT-3G D1: Quadratic-Estimator Lensing Reconstruction and Cos- mology,in prep.(2026)","cited_arxiv_id":null,"evidence_quote":"This is the 545 GHz cosmic infrared background map used as the external lensing tracer in the combination."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This study identifies extragalactic foreground contamination in temperature-based lens reconstruction, motivating the profile-hardened approach."}],"review_version":1}