{"id":"60c5f37f-0a92-4985-8441-cb03d023dc43","arxiv_id":"2509.03720","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Improved foreground cleaning and a 1% sky mask reduce the significance of two CMB anomalies (low correlation, local-variance asymmetry) from about 3 sigma to about 2 sigma.","lead":"The authors re-examine five well-known large-scale CMB temperature anomalies using new foreground-cleaned maps that allow a 1% sky mask, and find that two anomalies, the low angular correlation and the local-variance asymmetry, drop from about 3 sigma to about 2 sigma significance. The result matters because it weakens the case that these anomalies point to new physics beyond the standard cosmological model.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ALV significance reduction hinges on cleaning-validation shifts of up to 2.5 percentage points that are not directionally reported; the claim is not yet robust.","rationale":"The reader's weakest assumption is that the foreground-cleaned maps faithfully recover the CMB at 1° resolution. My concern is a specific, internally documented facet of that assumption: the cleaning validation shows shifts in ALV p-values that are large enough to potentially reverse the headline reduction. This is more concrete than the general map-fidelity worry because it is stated explicitly in Sec. 3.5 and is directly testable with the authors' existing simulations. The S1/2 claim is more robust (shifts ≤0.9% against p-values of ~6%), so I focus on ALV as the load-bearing element. The appropriate verdict remains CONDITIONAL, as the reader concluded: the analysis is careful and the internal results are plausible, but the ALV conclusion should be resolved by reporting cleaned-CMB p-values before accepting the abstract's claim that both anomalies reduce to ~2σ. No verdict adjustment is needed; my test would settle the condition.","tokens_in":22377,"tokens_out":6284,"duration_ms":66460,"concrete_test":"Recompute ALV p-values using the 10^4 cleaned-CMB simulations for the four foreground-cleaned maps with the 1% mask and full sky, and report these p-values alongside the CMB-only p-values, including the sign of the shift. If any cleaned-CMB 1%-mask p-value falls below 0.5%, the claimed reduction to ~2σ fails for that map; if all remain above 1%, the claim is robust. Additionally, increase the cleaned simulation set from 10^4 to 10^5 to reduce sampling noise, and verify whether the p-value ordering across maps is preserved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the local-variance asymmetry (ALV) drops from ~3σ to ~2σ with the 1% mask is not fully supported by the paper's own cleaning validation. In Sec. 3.5, the authors report shifts in PTE between CMB-only and cleaned-CMB simulations of up to Δp ≤ 2.5% for the 1% mask and Δp ≤ 4.3% for the full sky, yet they do not report the direction of these shifts or the cleaned-CMB p-values themselves. The quoted ALV p-values for the 1% mask are only 2.2–2.8%. A shift of 2.5 percentage points is therefore comparable to the p-value itself: if the cleaning-induced shift makes the anomaly more significant, the 1%-mask ALV could remain at ~3σ, contradicting the abstract's claim of reduced significance. The paper explicitly acknowledges the shift is 'not small' but proceeds to quote p-values from CMB-only simulations. Because ALV is one of the two anomalies highlighted in the abstract, the headline conclusion relies on an unresolved systematic uncertainty that is of the same order as the claimed effect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper re-evaluates five well-known CMB large-angle anomalies using four new foreground-cleaned, low-resolution WMAP+Planck maps from a companion paper that require only a 1% sky mask. For each anomaly statistic, p-values are computed from 10^5 Gaussian statistically-isotropic ΛCDM simulations and compared across the 1% mask, the Planck common mask (26% masked), and the full sky, as well as against the Planck 2018 component-separated maps. The main finding is that the significance of the low real-space correlation (S1/2) and of the local-variance asymmetry (ALV) drops from about 3σ with the common mask to about 2σ with the 1% mask; the northern variance, parity asymmetry, and quadrupole-octopole alignment remain at roughly 3σ, 2σ, and 3σ respectively. The paper concludes that individual anomalies do not by themselves strongly favor new physics over ΛCDM.","tokens_in":22655,"tokens_out":6189,"duration_ms":68823,"significance":"If correct, the paper is a valuable reassessment of the CMB anomaly landscape: it shows that a substantially larger usable sky fraction, enabled by improved foreground cleaning, reduces the significance of two frequently cited anomalies. The analysis has notable strengths: 10^5 Gaussian statistically-isotropic simulations, consistent mask handling across estimators, an explicit validation of the approximate S1/2 expression (Eq. 6), and dedicated cleaned-simulation tests for CMB-foreground chance correlation. The paper also correctly stresses the a-posteriori nature of the statistics and the need for look-elsewhere corrections. The central limitation is that for ALV the cleaning-induced PTE shift is comparable to the quoted p-values and is not directionally reported, so the headline reduction from ~3σ to ~2σ is not yet robust. I do not see a circularity problem: the anomaly statistics are computed directly from maps and compared with external Monte Carlo simulations.","major_comments":[{"comment":"The ALV cleaning-validation result is load-bearing for the abstract's claim that ALV drops from ~3σ to ~2σ. The paper reports Δp ≤ 4.3% (full sky) and Δp ≤ 2.5% (1% mask) when PTE is computed from cleaned-CMB rather than CMB-only simulations, and states that the shift is 'not small'. However, the quoted 1%-mask ALV p-values are only p=2.2–2.8% (Table 1). A 2.5 percentage-point shift is the same order as the p-value itself, and without reporting the direction it is possible that the cleaned-simulation p-value is ≤0.3%, i.e. still ~3σ. This would directly contradict the abstract. The manuscript must report the cleaned-CMB p-values and the direction of the shifts, and explain why the conclusion is unchanged. Note also that Sec. 2.2's blanket statement that the shifts are 'small or do not impact the conclusions' is inconsistent with Sec. 3.5's characterization of the ALV shift as 'not small'","section":"Sec. 3.5"},{"comment":"The central results rest on the four foreground-cleaned maps and the 1% mask from Nofi et al. (2025a), cited as 'in prep'. The manuscript states that the maps and masks are 'provided', but they are not public at the time of writing, and the cleaning validation in Sec. 2.2 applies the same cleaning procedure to simulated maps, so it cannot independently establish the residual-foreground level. If residual or over-subtracted foreground power survives at low multipoles, the reduced S1/2 and ALV significances could be an artifact of added power masking the anomaly. Please clarify the availability/release plan for the maps and masks and, if possible, provide an independent residual-foreground test beyond self-cleaned simulations.","section":"Sec. 2.1 / 2.2"}],"minor_comments":[{"comment":"Minor typo: 'Commander, NILC, SEVEM, und SMICA' should be 'Commander, NILC, SEVEM, and SMICA'.","section":"Sec. 2.1"},{"comment":"The companion-paper reference lists 'Bennet, C., L.'; the author name in the manuscript is Bennett. Please correct the spelling in the reference list.","section":"References"},{"comment":"The exclusion of outlier maps for ALV (Commander, SEVEM, and for the full sky also 94 GHz) is motivated by dipole directions pointing toward the Galactic plane, but the selection criteria are qualitative. Table 1 shows a wide spread among component-separated maps (e.g. SEVEM p=0.00% at 1% mask versus p=2.77% for the 143 GHz cleaned map). A formal, pre-defined outlier criterion would strengthen the claim of 'good agreement' among foreground-cleaned maps.","section":"Sec. 3.5"},{"comment":"Equation (9) is typeset in an unusual way (the r=1 evaluation and the gradient notation are cramped). The definitions of the multipole vectors and OAVs are otherwise clear, but a cleaner equation would help.","section":"Sec. 3.3"},{"comment":"The bottom panel of Fig. 2 and the discussion of S_μ are informative, but it would be useful to state explicitly that the p-values in the μ-scan are not corrected for the look-elsewhere effect; the paper later makes this point in Sec. 4, but a pointer there would prevent over-interpretation.","section":"Sec. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The ALV direction issue is the key blocking point: it is fixable within the manuscript's scope by reporting cleaned-CMB p-values and shift directions, but until then the abstract's headline claim about reduced ALV significance is not supported. The editor may also want to ensure that the companion papers (Nofi et al. 2025a,b) are available before acceptance, since the current manuscript relies on them for the maps, masks, and cleaning validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a careful, useful reassessment of five standard CMB anomalies on new foreground-cleaned maps, and the S1/2 result—significance drops from ~3σ with the 26% common mask to ~2σ with the 1% mask—looks solid. The ALV result is more fragile, and the paper's own cleaning test leaves a real open question about whether it actually drops to ~2σ.\n\nWhat's new is the quantitative evaluation of five anomalies on four foreground-cleaned maps that retain 99% of the sky, plus the same statistics on Planck component-separated maps and three sky cuts. That is a genuinely useful comparison. The methodology is careful: 10^5 Gaussian ΛCDM simulations, a validated approximation for S1/2, and a cleaned-simulation test for CMB-foreground chance correlation. The S1/2 conclusion holds across cleaning procedures and is consistent with the QML Cℓ. That part deserves credit.\n\nThe soft spot is ALV. In Sec. 3.5 the authors report that the PTE shift between CMB-only and cleaned-CMB simulations is up to 2.5 percentage points for the 1% mask—comparable to the quoted p-values of 2.2–2.8%—but they do not report the direction of the shift. If the cleaning makes the anomaly more significant, the 1%-mask ALV could still be ~3σ, contradicting the abstract. The paper acknowledges the shift is \"not small\" but quotes the CMB-only p-values anyway. That is a real unresolved systematic for one of the two headline claims.\n\nTwo smaller issues. First, the maps and masks are not public; they are in companion papers \"in prep.\" That limits independent verification. Second, the p-value ranges are quoted after excluding outlier maps (Commander, SEVEM, and the 94 GHz map for some statistics), and those exclusions are post hoc. Not fatal, but it makes \"good agreement across cleaning procedures\" overstate what the table actually shows.\n\nFor S1/2 and the other three statistics, I think the conclusions hold up. For ALV, the paper needs to report the direction of the cleaning-induced shifts and show the cleaned-simulation p-values explicitly; then the claim would be as solid as the S1/2 one.\n\nWho should read this: anyone working on CMB anomalies or large-scale isotropy tests, and model-builders who use anomaly significances as motivation. It deserves a serious referee. I would send it to review with a request for the missing ALV information and public maps.","headline":"Careful, useful reassessment: the S1/2 significance drop to ~2σ with a 1% mask holds up, but the ALV drop is not yet robust because the cleaning-validation shift is comparable to the p-value and its direction is unreported.","tokens_in":23178,"tokens_out":3468,"would_cite":true,"duration_ms":33505,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["98.70.Vc","98.80.Es"],"model":"deepseek-v4-flash","headline":"Using nearly full-sky maps built with improved foreground cleaning, this paper shows that the low large-angle correlation and the local-variance asymmetry drop from about 3σ to about 2σ — too weak, individually, to challenge the standard co","keywords":["cosmic microwave background","CMB anomalies","foreground cleaning","large-angle correlations","hemispherical asymmetry","statistical isotropy","low multipoles","parity asymmetry"],"falsifier":"Take the best-fit foreground templates from the cleaning, add them back to the cleaned maps, and re-measure S1/2 and ALV: if the ~3σ significance under the 1% mask returns, the reduced significance is a cleaning artifact; if it stays near ~2σ, the mask drove the result. Alternatively, an independent low-resolution cleaning built from different template sets (e.g., pure high-frequency dust and low-frequency synchrotron channels, without morphology fitting) applied with the same 1% mask would settle whether the result reproduces. The companion paper's 10^4 cleaned simulations bound the chance-co","tokens_in":22268,"feed_emoji":"🌌","tokens_out":8538,"duration_ms":75347,"temperature":0.7,"pith_summary":"This paper re-tests five well-known large-scale CMB temperature anomalies — low angular correlation, parity asymmetry, quadrupole–octopole alignment, low northern variance, and local-variance asymmetry — using new foreground-cleaned maps that mask only 1% of the sky instead of the usual 26%. The authors show that the choice of sky cut changes the verdict for two of the five: the low-correlation statistic S1/2 falls from p ≈ 0.2% (3.0–3.1σ) with the large Planck common mask to p ≈ 6–7% (1.8–1.9σ) with the 1% mask, and the local-variance asymmetry ALV falls from p ≈ 0.13–0.16% (3.2σ) to p ≈ 2.2–2.8% (2.2–2.3σ). The other three are stable across masks: northern variance near 3σ, parity asymmetry near 2σ, quadrupole–octopole alignment near 3σ. If the cleaned maps are trustworthy, two of the most cited anomalies are weaker evidence against statistically isotropic Gaussian ΛCDM than previously reported, and any alternative model would need to explain several anomalies at once or fit other measurements besides CMB temperature.","feed_headline":"Fuller CMB sky shrinks two anomalies to ~2 sigma","feed_subtitle":"Foreground cleaning that keeps 99% of the sky weakens the low-correlation and asymmetry signals.","key_machinery":"The carrier of the argument is a set of four foreground-cleaned maps at 1° resolution (70, 94, 100, 143 GHz) built by fitting six archival foreground templates with free normalization and masking only 1% of the sky, which keeps the low-ℓ spherical-harmonic coefficients nearly orthogonal. Five a-posteriori statistics are then evaluated on three sky cuts — full sky, 1% mask, and the 26% Planck common mask: S1/2, the integral of the squared angular correlation function above 60°; R27, the ratio of even to odd multipole power; SQO, the alignment of quadrupole and octopole planes via Maxwell multipole vectors and oriented area vectors; σ²16, the northern-ecliptic-hemisphere pixel variance at Nsid","core_discovery":"The central claim is that the significance of two commonly studied CMB anomalies depends strongly on how much of the sky is masked, and that with nearly full-sky maps the anomalies weaken. Using four foreground-cleaned WMAP/Planck maps at 1° resolution that require only a 1% galactic mask, the paper finds that the low real-space correlation statistic S1/2 goes from a 3.0–3.1σ deviation (p = 0.19–0.24%) under the 26%-masked Planck common mask to 1.8–1.9σ (p = 5.8–7.3%) under the 1% mask, and that the local-variance asymmetry ALV goes from 3.2σ (p = 0.13–0.16%) to 2.2–2.3σ (p = 2.2–2.8%). The other anomalies are mask-stable: low northern variance stays near 3σ, parity asymmetry near 2σ, and th","pith_inferences":["If the ~2σ result holds under independent cleaning, it suggests that the historical 'missing angular correlation' was partly an artifact of the standard analysis convention: a 26% mask reshapes the low-ℓ mode structure and suppresses C(θ) more than the underlying sky does.","A sharp testable consequence: applying the same nearly full-sky cleaning to low-ℓ CMB polarization should preserve the quadrupole–octopole alignment if it is cosmological, while the correlation and variance-asymmetry signals should stay near ~2σ if they are mask artifacts.","The five anomalies may not share a single origin: the results split them into a mask-sensitive pair (S1/2, ALV) and a mask-stable trio (SQO, σ²16, R27), which argues against any one new-physics mechanism.","Once the cleaned maps and 1% mask are public, S1/2 can be recomputed with alternative estimators (e.g., QML or Bayesian Cℓ on the same sky cut); the paper's own comparison with the Planck QML Cℓ already points the same way."],"forward_implications":["The low large-angle correlation and the local-variance asymmetry are only ~2σ phenomena once 99% of the sky is usable, so by themselves they are weak evidence against statistically isotropic Gaussian ΛCDM.","The quadrupole–octopole alignment remains the strongest and most stable feature, at 3.2–3.5σ, essentially independent of cleaning method and sky cut.","Any alternative model must explain several anomalies at once, or describe a non-CMB measurement better, to be preferred over ΛCDM.","The explicit dependence of S1/2 on the integration limit µ and of R27 on ℓmax motivates look-elsewhere corrections, which would further reduce the significance of these statistics.","The paper's analysis implies that the standard 26%-mask convention itself contributed to the reported S1/2 and ALV signals: the 1%-mask results are consistently closer to the ΛCDM expectation."],"supporting_citations":[{"why":"Supplies the four foreground-cleaned maps at 1° resolution and the 1% mask that make the nearly full-sky analysis possible.","marker":"Nofi et al. (2025a)"},{"why":"Companion analysis that establishes the low quadrupole at 2.2σ and contributes the foreground-cleaned simulation test used to bound CMB–foreground chance correlation.","marker":"Nofi et al. (2025b)"},{"why":"Provides the component-separated maps (Commander, NILC, SEVEM, SMICA) used as the cleaning-method comparison baseline.","marker":"Planck Collaboration IV (2020)"},{"why":"Supplies the downgrading procedure, the low-northern-variance convention, and the local-variance methodology.","marker":"Planck Collaboration XVI (2016)"},{"why":"Defines the S1/2 statistic that measures the lack of large-angle correlation.","marker":"Spergel et al. (2003)"},{"why":"Introduces multipole vectors and oriented area vectors and the SQO alignment statistic.","marker":"Copi et al. (2004)"},{"why":"Defines the local-variance asymmetry ALV statistic in disks.","marker":"Akrami et al. (2014)"},{"why":"Provides prior anomaly significance baselines and the normalization scheme for the local-variance map.","marker":"Muir et al. (2018)"},{"why":"Prior Planck results on S1/2 and the anomaly significances this paper compares against.","marker":"Planck Collaboration VII (2020)"},{"why":"Provides the σ²16 estimator and recent anomaly significances used for comparison.","marker":"Jones et al. (2023)"}],"fun_headline_variants":["99% sky view drops two CMB anomalies to ~2σ","Full-sky CMB maps soften two anomalies to ~2σ","Nearly full CMB sky shrinks two anomalies to ~2σ","Reduced sky masking lowers two CMB anomalies to ~2σ","Using 99% of CMB sky cuts two anomalies to ~2σ"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the foreground-cleaned maps contain only true CMB temperature at degree scales outside the 1% mask — no residual or over-subtracted Milky Way foreground that could shift the large-scale statistics — since the paper's validation of the cleaning uses simulations cleaned by the same method it is checking.","fun_headline_variants_meta":{"raw":{"variants":["99% sky view drops two CMB anomalies to ~2σ","Full-sky CMB maps soften two anomalies to ~2σ","Nearly full CMB sky shrinks two anomalies to ~2σ","Reduced sky masking lowers two CMB anomalies to ~2σ","Using 99% of CMB sky cuts two anomalies to ~2σ"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2853,"prompt_tokens":899,"completion_tokens":1954,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":1861}},"tokens_in":643,"tokens_out":1954,"duration_ms":14855,"temperature":1.0,"reasoning_tokens":1861,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T10:43:01.417747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the best-fit foreground templates from the cleaning, add them back to the cleaned maps, and re-measure S1/2 and ALV: if the ~3σ significance under the 1% mask returns, the reduced significance is a cleaning artifact; if it stays near ~2σ, the mask drove the result. Alternatively, an independent low-resolution cleaning built from different template sets (e.g., pure high-frequency dust and low-frequency synchrotron channels, without morphology fitting) applied with the same 1% mask would settle whether the result reproduces. The companion paper's 10^4 cleaned simulations bound the chance-co","supporting_citations":[{"cited_title":"J., Huterer, D., & Starkman, G","cited_arxiv_id":null,"evidence_quote":"Introduces multipole vectors and oriented area vectors and the SQO alignment statistic."},{"cited_title":"2014, Astrophys","cited_arxiv_id":null,"evidence_quote":"Defines the local-variance asymmetry ALV statistic in disks."}],"review_version":1}