{"id":"4c67c843-4306-4317-8915-304c3fb4ef02","arxiv_id":"2507.12315","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Marked angular power spectra applied to HSC-Y1 weak lensing data yield S8 = 0.807 ± 0.024, about 43 percent tighter than standard power spectra.","lead":"This paper applies marked angular power spectra to galaxy lensing maps from the Hyper Suprime-Cam survey, reporting a 43 percent tighter constraint on the cosmic clustering parameter S8 than standard power spectra. It is the first such application to weak lensing data, offering a practical way to extract higher-order information from future surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Emulator validation is internal to the cosmo-varied suite; Appendix B shows the same pipeline overestimates S8 on the independent covariance suite, so the reported S8=0.807 and 1.43x gain may be dominated by a simulation-suite mismatch.","rationale":"The reader's weakest assumption points at emulator fidelity; I agree partially but sharpen it. The load-bearing issue is not generic small-scale emulator error; it is the demonstrable, unexplained mismatch between the two simulation suites that both enter the analysis—one for covariance, one for emulation. Appendix B is the paper's own external validation and it fails: S8 is systematically overestimated on the covariance suite for every tested scale cut. The authors acknowledge the offset and defer investigation, but still quote the result as fiducial. This makes the central numerical claim conditional at best. I would not reject: the method, the systematics tests, and the paper's transparency are valuable, and a calibration correction might bring S8 into line with other HSC analyses. The scale-mixing discussion in Sec 4.3 is honest and should be credited, but it undercuts the 'non-Gaussian information' framing of the improvement. Thus no verdict change from the reader's CONDITIONAL is needed, but the revision conditions should include an external emulator calibration test, not just internal leave-one-out validation.","tokens_in":23432,"tokens_out":5005,"duration_ms":62185,"concrete_test":"Take the mean covariance-suite data vector used in Appendix B, add the per-bandpower calibration residual from Fig B2 (true/emulated) to the HSC-Y1 data vector, and re-run the moped-compressed MCMC (Sec 3.6). If the inferred S8 shifts by more than 0.02 (the quoted 1σ error is 0.024) relative to 0.807, or if the improvement over C^κκ drops below ~1.2x, then the claimed constraint and 43% gain are dominated by the simulation-suite mismatch. A complementary check: repeat the same calibration test with the mark-B cross-spectrum removed, since Fig B2 shows the largest residuals there.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (S8=0.807±0.024, ~43% improvement) rests on comparing HSC-Y1 data to a GP emulator trained on the 100-cosmology cosmo-varied suite (Sec 3.5). The only validation of that emulator is leave-one-out within the same suite. Appendix B performs the crucial external check: it feeds the mean data vector of the independent covariance suite (Takahashi et al. 2017), the same suite used for the covariance, through the pipeline. The result is a systematic overestimate of S8 for every ℓmax tested, with the largest spectral residuals in the mark-B cross-spectrum (Fig B2), attributed to a known large-scale power deficit in the cosmo-varied suite. Because the real HSC data are processed with this same emulator, the reported S8 inherits this bias. Appendix D then shows real data exhibit a rising S8 with ℓmax that simulations do not reproduce, so the scale cut ℓmax=1500 is chosen in a regime where the data/simulation mismatch is already visible. These are not external model disagreements; they are internal inconsistencies in the calibration path of the central number. The 43% improvement is also partly an artifact of scale mixing: Sec 4.3 shows the Gaussian part of C^ΔΔ_ℓ is sensitive to P(k) up to ℓ≈3200, beyond the ℓmax=1500 applied to C^κκ_ℓ, so the comparison is not between statistics at identical information content. Unless the emulator is calibrated against an independent suite and the scale-mixing contribution is subtracted, neither the S8 value nor the improvement factor is established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the first application of marked angular power spectra to weak lensing data, using HSC-Y1 convergence maps. Marked fields are constructed by weighting the convergence field with three nonlinear mark functions (a Gaussian-process-derived mark, a smoothed-field mark, and a modified White 2016 mark), and their auto- and cross-spectra are combined with the standard convergence power spectrum. A Gaussian-process emulator trained on 100 cosmology-varied N-body simulations is used to model the summary statistics, and a covariance from 2268 pseudo-independent realizations is used for likelihood inference. The baseline analysis yields S8 = 0.807 ± 0.024 and claims a ~43% improvement in the S8 error bar over the standard power spectrum. The paper also tests sensitivity to baryonic effects, intrinsic alignment, photo-z errors, and multiplicative shear bias, and shows that the marked spectra contain bispectrum- and trispectrum-like contributions.","tokens_in":23796,"tokens_out":5978,"duration_ms":65393,"significance":"The paper is a proof-of-concept that marked angular power spectra can be applied to real weak-lensing data and combined with standard spectra to tighten cosmological constraints. If the result holds, it would be a useful addition to the toolkit of higher-order statistics for Stage-IV surveys. The simulation pipeline is detailed, the systematic tests are extensive, and the paper is unusually transparent about its own internal inconsistencies: the simulation-suite offset in Appendix B, the scale-mixing interpretation in Section 4.3, and the unexplained scale-dependent S8 trend in Appendix D. These are genuine strengths. However, the same passages show that the headline improvement factor and the quoted S8 value rest on an emulator that fails an external validation test and on a comparison that is not matched in information content. The central claim is therefore not yet established, although the underlying idea is promising.","major_comments":[{"comment":"The emulator is validated only by leave-one-out tests inside the cosmo-varied suite. Appendix B provides the crucial external check: feeding the covariance-suite mean data vector through the pipeline systematically overestimates S8 for every ℓmax tested, with the largest spectral residuals in the mark-B cross-spectrum (Fig. B2). Because the same emulator is used for the HSC-Y1 analysis, the reported S8 = 0.807 ± 0.024 inherits this bias. The paper acknowledges the offset but does not correct for it or marginalize over it. The central claim requires either recalibrating the emulator against an independent suite or quantifying the bias in units of the final uncertainty and adding it to the systematic budget.","section":"Section 3.5 and Appendix B"},{"comment":"The decomposition in Eq. (14) shows that the Gaussian (disconnected) part of C^ΔΔ_ℓ is sensitive to the power spectrum at multipoles ℓ ≲ 3200 for θ = 2', far beyond the ℓmax = 1500 applied to C^κκ_ℓ in the baseline comparison. The paper itself concludes that the apparent additional constraining power likely comes from Gaussian fluctuations on smaller scales, not from intrinsic non-Gaussianity. A direct comparison at matched effective information content is therefore not made, and the abstract's statement that marked spectra 'improve constraints on S8 by ≈43% compared to standard two-point power spectra' overstates the improvement as a purely non-Gaussian gain. The claim should be reworded to separate scale mixing from genuine higher-order information, or the power-spectrum baseline should be extended to the same effective scale.","section":"Section 4.3, Eq. (14), and the abstract"},{"comment":"The HSC-Y1 data show a consistent rise in S8 as ℓmax increases, for both marked and standard spectra, while neither simulation suite reproduces this trend. The baseline scale cut ℓmax = 1500 is chosen in a regime where this trend is already visible. Given the emulator bias demonstrated in Appendix B, the data-versus-simulation mismatch weakens the interpretation of the reported S8 as a cosmology measurement. The paper asserts the bias is within statistical uncertainties, but no quantitative comparison between the observed trend and the simulation-derived uncertainty is provided. The authors should demonstrate explicitly that the scale-dependence seen in Figure D1 is consistent with the covariance-suite validation once the Appendix B offset is accounted for.","section":"Appendix D and Figure D1"},{"comment":"The fiducial S8 value is quoted inconsistently: the abstract and Section 4.1 give S8 = 0.807 ± 0.024, while Section 5 gives S8 = 0.804 ± 0.023. Similarly, the abstract claims an improvement of ≈43%, Section 4.1 states '1.4× smaller error', and Section 5 quotes an improvement factor of ~1.43. These numbers refer to the same baseline analysis but are not mutually consistent in their presentation. The headline number and the claimed improvement factor should be stated once, consistently, with the associated error bars, and the abstract must match the body.","section":"Section 4.1, Section 5, and the abstract"}],"minor_comments":[{"comment":"The stated priors are 'uniform priors on Ωm of [0.1, 0.4] and S8 of [0.5, 0.1]'; the S8 interval is clearly a typo (it should presumably be [0.5, 1.0] or similar, based on Fig. 3), and should be corrected.","section":"Section 3.6"},{"comment":"The caption says 'The coloured regions show 1/3σ confidence interval of the underlying cosmology'; this is ambiguous and likely should read '1σ and 3σ confidence intervals'.","section":"Figure 4 caption"},{"comment":"The text above each bar is described as the 'improvement in the value of σ(S8) found from C^κκ_ℓ', but the y-axis is labelled σ(S8) and the bars appear to show absolute errors; the relationship between the labels and the bars should be clarified.","section":"Figure 8"},{"comment":"In the sentence 'they are almost affected by finite thickness effects', 'almost' appears to be a typo, likely 'also'; this should be corrected for clarity.","section":"Appendix B (last paragraph)"},{"comment":"In Eq. (14) and the surrounding text, the normalization of the marked field Δ(x) is dropped for simplicity, as noted in the footnote; it would be helpful to state explicitly in the main text that the 1/σ(κ_θ) factor is omitted throughout the n-point function expansion.","section":"Section 4.3"},{"comment":"The heatmaps in Figure C1 are informative, but the text states that 'white text flags deviations greater than 0.3σ' without explaining the colour scale for values below that threshold; adding a colour bar with the deviation units would improve readability.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"This paper is part of a coordinated series of non-Gaussian statistics applied to HSC-Y1, and it shares a known simulation-suite discrepancy with the companion papers. The methodological novelty is real, but the headline claims are not yet supported: the emulator's external validation fails (Appendix B), the improvement factor is inflated by scale mixing (Section 4.3), and the real-data S8 trend with scale (Appendix D) remains unexplained. I would require, as a condition for acceptance, that the authors (i) reprocess the covariance-suite validation through the full pipeline and either correct or marginalize over the resulting S8 bias, (ii) report the improvement factor after removing the Gaussian scale-mixing contribution or state it as an effective gain, and (iii) reconcile the inconsistent fiducial S8 values. The paper's transparency about these issues is commendable and the underlying approach is worth publishing after these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know first: this is the first time marked angular power spectra have been run on real weak lensing data, and the paper is worth reading for the method regardless of the headline number. The three mark choices are sensible, the GP emulator is a reasonable way to model a statistic with no fast analytic prediction, and the systematic tests (baryons, IA, photo-z, m-bias) are unusually extensive for a first application. The honest appendix text, which connects mark B to bispectrum and trispectrum pieces and flags the scale-mixing explicitly, is good scientific hygiene.\n\nThe soft spot is the central constraint. The emulator is validated only by leave-one-out on the same 100-cosmology suite. Appendix B is the independent check: throwing the covariance-suite mean through the same pipeline overestimates S8 at every lmax, with residuals concentrated in the mark-B cross spectrum. That is exactly the calibration path used for the real HSC data, so the quoted S8 = 0.807 ± 0.024 inherits an unquantified systematic shift. Appendix D shows real data have S8 rising with lmax while simulations do not; choosing lmax = 1500 sits in a regime where that mismatch is already visible. And Section 4.3 concedes that part of the gain is Gaussian small-scale leakage rather than intrinsic non-Gaussian information. The 43% improvement and the absolute S8 value are therefore both softer than the abstract suggests.\n\nThe authors know all of this and say it in the appendices, so I don't think there's any bad faith. The abstract overclaims, but the body is honest. The fix is straightforward: calibrate the emulator on the covariance suite, or fit a nuisance offset, and quote a systematic error that includes the simulation-suite difference.\n\nWho gets value: anyone working on higher-order lensing statistics, marked statistics, or simulation-based inference. It deserves a serious referee; the method is novel and the failure modes are instructional. I'd want the numerical claims reined in before publication, and ideally a reanalysis with the emulator trained or validated on an independent suite. This is the kind of paper I'd bring to reading group and cite for the method.","headline":"First marked-spectra weak lensing analysis: method is real, but the headline S8 is not yet calibrated.","tokens_in":24366,"tokens_out":2422,"would_cite":true,"duration_ms":27436,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Marked angular power spectra, applied here to real lensing data for the first time, tighten S8 by about 43 percent over the standard spectrum, giving S8 = 0.807 ± 0.024.","keywords":["marked angular power spectra","weak lensing convergence","S8 clustering amplitude","non-Gaussian information","Hyper Suprime-Cam Year 1","Gaussian process emulator","higher-order statistics","cosmological parameter constraints"],"falsifier":"Build the emulator from a single, higher-resolution simulation suite (or one including baryonic feedback), re-infer $S_8$ from the same HSC-Y1 data vector, and watch the $\\ell_{\\rm max}$ dependence: the Appendix D rise of $S_8$ with scale predicts that training on simulations with more small-scale power should flatten that trend and pull the fiducial value below $0.807$, while removing mark B, the largest source of simulation-suite residuals in Appendix B, should barely move the constraint if the multi-mark result holds.","tokens_in":23251,"feed_emoji":"🔭","tokens_out":16873,"duration_ms":153224,"temperature":0.7,"pith_summary":"This paper is the first to apply marked angular power spectra to real weak-lensing data. The idea is simple: before measuring angular power spectra of the convergence field $\\kappa$, weight the field by a non-linear mark function of its smoothed version, $\\Delta(\\kappa)=m(\\kappa_\\theta)\\,\\kappa$, so that different density environments contribute with different weights and higher-order information enters an otherwise standard two-point statistic. Using HSC-Y1 convergence maps, the authors combine three complementary mark functions and report that the joint marked auto- and cross-spectra tighten the clustering parameter $S_8\\equiv\\sigma_8\\sqrt{\\Omega_{\\rm m}/0.3}$ by about 43 percent compared with the standard power spectrum alone, giving $S_8=0.807\\pm0.024$. If correct, this makes marked spectra a computationally cheap way for current and future surveys to recover non-Gaussian information that two-point analyses miss.","feed_headline":"43 percent tighter S8 constraints from marked lensing spectra","feed_subtitle":"First use of weighted convergence spectra on HSC data gives S8 = 0.807 ± 0.024, central to the S8-tension debate.","key_machinery":"The central object is the marked convergence field $\\Delta(\\kappa)=m(\\kappa_\\theta)\\,\\kappa$, where $m$ is a mark function and $\\kappa_\\theta$ is the convergence field smoothed by a Gaussian of width $\\theta\\in\\{2',4',10'\\}$; its angular power spectrum is computed with the same mask-corrected pseudo-$C_\\ell$ machinery as the ordinary spectrum. Three mark functions probe different density environments: a Gaussian-process-shaped weight (A), the rescaled smoothed field (B), and a modified power law that up-weights underdensities (C). The cosmological prediction is carried by a Gaussian-process emulator trained on 100 cosmology-varied N-body simulations spanning $\\Omega_{\\rm m}$ and $\\sigma_8$, with each bandpower emulated separately, while the covariance is estimated from 2268 realisations of a fiducial-cosmology simulation suite and inverted with the Anderson--Hartlap correction. The emulator is what turns the measured spectra into a posterior on $S_8$ and $\\Omega_{\\rm m}$, so the paper's quoted improvements stand or fall on its fidelity.","core_discovery":"The paper's central claim is that marked angular power spectra are a practical higher-order statistic for weak lensing, not just a theoretical construct. Starting from HSC-Y1 convergence maps, the authors build nine marked fields from three mark functions (a Gaussian-process-shaped weight, the rescaled smoothed field $\\kappa_\\theta/\\sigma(\\kappa_\\theta)$, and a modified power law that up-weights underdense regions) at smoothing scales of 2, 4, and 10 arcminutes, and measure their auto-spectra and cross-spectra with the original $\\kappa$ field. Combining all three marks with the standard spectrum tightens $S_8$ to $0.807\\pm0.024$, a factor of 1.43 (32 percent smaller error bars) over the power spectrum alone under identical scale cuts, with the gain driven by the marks having different degeneracy directions in the $S_8$--$\\Omega_{\\rm m}$ plane. The paper further argues that part of the improvement is scale mixing: the non-linear mark lets small-scale Gaussian power leak into large-scale multipoles, so the marked spectra are a practical proxy for bispectrum- and trispectrum-like information rather than a pure measurement of non-Gaussianity, while the $m_B$ cross-spectrum is a projected bispectrum integral and the auto-spectrum splits into Gaussian and connected-trispectrum parts. Systematic tests on baryonic feedback, intrinsic alignments, photometric-redshift choice, and multiplicative shear bias keep $S_8$ shifts within about $0.4\\sigma$ at the adopted scale cuts.","pith_inferences":["If the 43 percent gain survives in Stage-IV surveys with much lower shape noise, marked spectra become one of the cheapest non-Gaussian additions to a lensing pipeline, since they reuse the standard power-spectrum code path and simulation-based emulators.","The scale-mixing decomposition suggests a design principle for future work: comparing marked-spectrum gains across smoothing scales separates Gaussian leakage from genuine non-Gaussian information, and could be used to engineer marks that maximize the truly non-Gaussian component.","The rising $S_8$ with $\\ell_{\\rm max}$ in the real data, absent in simulations, reads as a small-scale modelling deficit; if higher-resolution training shifts the fiducial value, the reported $0.807$ should be revised downward."],"forward_implications":["No single mark matches the trio: marks B and C have nearly orthogonal degeneracy directions in the $S_8$--$\\Omega_{\\rm m}$ plane, which the paper identifies as the source of the 1.43$\\times$ gain from combining them.","Part of the constraining power of marked spectra comes from Gaussian small-scale information leaking into large-scale multipoles through the non-linear mark, so the method is a practical proxy for higher-order correlations rather than a pure non-Gaussian statistic.","With smoothing scales and scale cuts chosen so that each tested systematic shifts $S_8$ by less than roughly $0.4\\sigma$, the baseline analysis keeps $\\ell_{\\rm max}=1500$ and the inferred $S_8$ stays within that tolerance for baryons, intrinsic alignments, photo-$z$ choice, and multiplicative shear bias.","Mark B dominates the $\\Omega_{\\rm m}$ constraining power, with an error about a third of the prior width, although the paper cautions that its $\\Omega_{\\rm m}$ value is affected by emulator accuracy.","Leaving out cross-correlations between tomographic redshift bins leaves information unused; the paper notes these could tighten constraints further."],"supporting_citations":[{"why":"Supplies the Gaussian-process mark function $m_A$ and the optimisation logic behind its shape.","marker":"Cowell et al. (2024)"},{"why":"Origin of the modified power-law mark $m_C$ that up-weights underdense regions; the paper adapts its safety function for projected fields.","marker":"White (2016)"},{"why":"Provides the 100 cosmology-varied N-body simulations (50 realisations each) on which the emulator is trained.","marker":"Shirasaki et al. (2021)"},{"why":"Companion HSC-Y1 analysis whose tailored mocks and systematic contamination procedures are reused to build and test the data vector.","marker":"Marques et al. (2024a)"},{"why":"Produces the 108 full-sky N-body simulations whose 2268 pseudo-independent realisations form the fiducial-cosmology covariance matrix.","marker":"Takahashi et al. (2017)"},{"why":"Supplies the Anderson--Hartlap factor that corrects the inverse covariance for the finite number of simulation realisations.","marker":"Hartlap et al. (2007)"},{"why":"Sets the HSC-Y1 multipole binning and the intrinsic-alignment amplitude range adopted for the systematic tests.","marker":"Hikage et al. (2019)"}],"fun_headline_variants":["Marked lensing spectra sharpen S8 by 43% in HSC Year 1","First marked-spectra lensing analysis tightens S8 to 0.807","HSC marked spectra yield 43% better S8 constraints","New lensing statistic cuts S8 error bars 43% on HSC data","Marked angular spectra boost S8 precision 43% in HSC-Y1"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands or falls on whether a Gaussian-process emulator trained on 100 gravity-only N-body simulations predicts the marked angular power spectra of the real universe on the scales used in the analysis, and the paper itself documents the risk: a systematic offset between its two simulation suites and a rise in the inferred $S_8$ with smaller scales in real data that the simulations do not reproduce.","fun_headline_variants_meta":{"raw":{"variants":["Marked lensing spectra sharpen S8 by 43% in HSC Year 1","First marked-spectra lensing analysis tightens S8 to 0.807","HSC marked spectra yield 43% better S8 constraints","New lensing statistic cuts S8 error bars 43% on HSC data","Marked angular spectra boost S8 precision 43% in HSC-Y1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000219,"raw_usage":{"total_tokens":1520,"prompt_tokens":1098,"completion_tokens":422,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":714,"tokens_out":422,"duration_ms":4921,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:48:48.908584+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build the emulator from a single, higher-resolution simulation suite (or one including baryonic feedback), re-infer $S_8$ from the same HSC-Y1 data vector, and watch the $\\ell_{\\rm max}$ dependence: the Appendix D rise of $S_8$ with scale predicts that training on simulations with more small-scale power should flatten that trend and pull the fiducial value below $0.807$, while removing mark B, the largest source of simulation-suite residuals in Appendix B, should barely move the constraint if the multi-mark result holds.","supporting_citations":[],"review_version":1}