{"id":"5b1f8a04-0593-4c37-ab0f-ebdc742ebadd","arxiv_id":"2509.05152","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Deep-field AnaCal computes shear responses from deep-field images rather than adding noise to wide-field images, meeting LSST bias requirements while raising effective galaxy density from 17 to 30 arcmin^-2.","lead":"This paper introduces a way to calibrate weak lensing shear measurements using deep-field images, so wide-field survey images do not need extra noise added to them. In simulations it meets the LSST bias requirements and raises usable galaxy density from 17 to 30 per square arcminute, reducing shear uncertainty by about 25 percent.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Deep-field response may not transfer to wide-field population: simulations draw both samples from the same galaxy catalog, leaving systematic depth/redshift differences untested.","rationale":"The reader's weakest_assumption identifies exactly the same gap: the deep-field response is assumed to apply to the wide-field population, and the simulations do not test systematic galaxy-property differences. This is the most load-bearing concern because it directly targets the core architecture of the method—separating where shapes are measured from where responses are measured. If the response is population-dependent, the method's unbiasedness on real data fails even if the simulation-level demonstration is correct. The paper's own future-work statement about reweighting deep-field galaxies to match wide-field properties reinforces that this is a known, unresolved issue. The concern does not invalidate the simulation results, which are internally consistent, but it does condition any claim of applicability to LSST-like surveys on additional tests that have not been performed. Therefore the reader's CONDITIONAL verdict remains appropriate; no verdict change is needed.","tokens_in":14419,"tokens_out":8383,"duration_ms":91638,"concrete_test":"Run the §3.2 blended-galaxy pipeline with a controlled population split: simulate wide-field images from a bright-limited catalog (e.g., F814W < 24) and deep-field images from a deeper catalog (e.g., F814W < 25.2), both drawn from the same COSMOS/descwl input, with the paper's PSF/noise-matching procedure unchanged. Measure the multiplicative bias via Eq. (20). If |m| > 3×10^-3 at 99.7% confidence, the transfer assumption fails. A cheaper check: on a fixed set of galaxies, compute ⟨R_dacal⟩ in Eq. (16) separately for the deep-field galaxy subsample and the wide-field galaxy subsample after applying identical renoising; if the fractional difference exceeds 3×10^-3, the bias is already present before shape noise averaging.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The estimator in Eq. (17) combines shape measurements from wide-field images with a shear response measured from deep-field images. The simulations in §4.1 and §4.2 construct deep-field images by 'randomly selected a new galaxy for every corresponding wide-field image' (§4.2), meaning the wide and deep samples have identical intrinsic galaxy-property distributions. Thus the tests verify only that the response transfers between independent realizations of the same underlying population, not between populations that differ systematically. In a real survey, deep fields are deeper (10× exposure) and detect fainter, higher-redshift galaxies than the wide-field shape sample. The response R in Eq. (16) depends on galaxy properties (S/N, size, morphology) through the nonlinear weights in Eq. (6). If the mean response of the deep-field sample differs from that of the wide-field sample, the ratio in Eq. (17) inherits a multiplicative bias from that mismatch. Section 4.3 quantifies only the statistical scatter in R due to finite area (sample variance), not a systematic offset induced by different selection functions. The paper's own concluding remark that future work will 'reweigh deep-field galaxies to match their properties' (§5) implicitly acknowledges this gap. No simulation in the paper varies the galaxy property distribution between wide and deep inputs while keeping PSF/noise matching fixed, so the unbiasedness claim is untested against the one systematic perturbation that the method's design specifically exposes it to.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Deep-Field Analytical Calibration (deep-field AnaCal), an extension of the AnaCal/FPFS shear estimation framework. The key idea is to measure galaxy shapes from wide-field images (avoiding the noise penalty of standard AnaCal) while computing the ensemble shear response from deeper-field images, with a noise/PSF matching procedure defined in Eqs. (14)-(16). The estimator is given in Eq. (17). The method is tested on isolated and blended galaxy simulations with LSST-like conditions. Reported results include multiplicative bias consistent with |m| < 3e-3 at 3σ, an improvement in effective number density from 17 to 30 arcmin^-2 relative to wide-field-only AnaCal, and a sample-variance calibration uncertainty of ≲0.3% for LSST Deep Drilling Fields. The paper argues that deep-field AnaCal preserves the statistical power of wide-field images while avoiding the need to inject an extra noise layer into them.","tokens_in":14736,"tokens_out":13743,"duration_ms":159110,"significance":"If the method works as claimed, it offers a practical and computationally efficient way to recover a substantial fraction of the statistical power lost in standard AnaCal's renoising procedure. The paper is careful in several respects: the simulations are controlled (fixed/variable noise, fixed/variable PSF, single- and two-component galaxy models, isolated and blended cases), no parameters are fitted to make the measured biases pass the requirement, and the code is publicly available. The O(γ^3) accuracy and noise-bias proofs are imported from prior work (Li et al. 2024a), which is appropriate. The main unresolved issue is whether the deep-field response can be transferred to the wide-field shape population in a real survey, where the two samples are not drawn from the same galaxy-property distribution. This is the load-bearing assumption of the method and is not tested in the present simulations.","major_comments":[{"comment":"The simulations do not test the transfer of the deep-field response to a wide-field shape sample with different galaxy properties. In §4.2 the deep-field images are made by \"randomly selecting a new galaxy for every corresponding wide-field image,\" so the wide and deep samples have the same intrinsic galaxy-property distribution. The same is true for the isolated-galaxy tests in §3.1, where the same galaxy catalog is used with lower noise. Thus the reported unbiasedness (|m|<3e-3) is established only for the case <R_wide> = <R_deep> by construction. In a real survey, deep fields are deeper and select fainter, higher-redshift, and possibly morphologically different galaxies than the wide-field shape sample. Since the response in Eq. (16) depends on galaxy properties through the nonlinear weights in Eq. (6), a systematic difference between the wide and deep populations would induce a multi","section":"§4.2 and §5, Eq. (17)"},{"comment":"All simulations provide the true PSF model and the true noise variance to the estimator. This is stated clearly, but it means the validation does not include PSF misestimation, which is a leading shear systematic. The abstract's claim that the method mitigates biases \"arising from ... the PSF\" is broader than what is tested. Please either add tests with an imperfect PSF model (e.g., wrong size or ellipticity) or temper the abstract/conclusions to make clear that the results assume a perfect PSF model and known noise variance.","section":"§3, first paragraph"}],"minor_comments":[{"comment":"The claimed statistical gain factor sqrt(2/(1+2/10)) ≈ 1.30 is described as reducing pixel noise variance by 30%. From the stated noise contributions, the ratio of variances is (σ_w²+2σ_d²)/(2σ_w²) = 0.6, i.e., a 40% variance reduction, and the corresponding uncertainty reduction is about 23%, not 30%. Also, the measured m uncertainty in Table 2 improves by a factor ≈2.1/1.3 ≈ 1.62, larger than the predicted factor. Please clarify which quantity the factor applies to (σ_γ vs. m uncertainty) and reconcile the numbers.","section":"§4.2 and Abstract"},{"comment":"The notation for the added noise layers is slightly inconsistent: Eq. (14) uses n'_deep and n''_deep, while Eq. (15) uses n'_wide and n'_deep, and the response Eq. (16) subtracts 2δν'_deep and δν'_wide. It would help readers to define explicitly which noise realizations are rotated by 90 degrees and which enter the shape vs. response images, to make the noise-matching construction easier to verify.","section":"Eq. (14)-(16)"},{"comment":"The n_eff comparison assumes equal areas for the wide and deep fields, as noted in the table caption. This is a useful idealized comparison, but the text should state more prominently that in a real survey the deep-field area is much smaller, so the n_eff gain applies only to a subset of the survey unless the response is transferred to the rest of the sky; the sample-variance discussion in §4.3 partially addresses this but not the population-mismatch point.","section":"Table 3 and §4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable incremental extension of Zhang et al. (2023) and the AnaCal framework, and the technical implementation appears sound. My main concern is that the central unbiasedness claim is validated only for the case where the wide and deep galaxy populations are statistically identical. Given that the authors themselves identify property matching as future work, I recommend a major revision requiring either an explicit population-mismatch simulation or a substantial qualification of the claims. The paper is otherwise well-suited to an astro-ph methods journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a well-executed extension of AnaCal that does what it claims in the simulations it runs, and the simulation campaign is more thorough than most shear-calibration papers. The headline gain—n_eff from 17 to 30 arcmin^-2, noise variance down 30%—is real under their assumptions. But the load-bearing assumption that deep-field responses apply to the wide-field shape sample is never tested directly, because both samples are drawn from the same galaxy catalog. That's the soft spot, and it's exactly what a referee should push on.\n\nWhat's new: the application of deep-field response calibration to the analytical FPFS/AnaCal framework (prior deep-field work was metacalibration), the explicit formalism for matching PSFs and noise across two fields, and the blended-galaxy validation with shear-dependent detection. They also quantify sample variance for the LSST DDF strategy using OpenUniverse2024, which is a nice touch. The code is public; the paper is honest about what it does and doesn't test.\n\nThe simulations themselves look clean. The isolated cases 1–4 give m consistent with zero at 3σ, and the blended test gives m=0.9±1.3e-3, within the |m|<3e-3 requirement. The claimed gain factor sqrt(2/(1+2/10))≈1.30 matches the reported uncertainty reduction. Good use of 90-degree rotation and paired shears to suppress shape noise.\n\nNow the soft spots. The big one: in §4.2, each deep-field image is built by randomly selecting a new galaxy for every wide-field image—same underlying population, just independent realizations. So the tests verify the response transfers between noise and PSF realizations, not between populations with different redshift, size, or S/N distributions. In a real survey, deep fields go fainter and detect different galaxies. If their mean response R differs from the wide-field sample's, Eq. (17) picks up a multiplicative bias. Section 4.3 only quantifies statistical scatter from finite area, not a systematic offset from selection. The paper's own last paragraph says future work will reweigh deep-field galaxies to match properties—that's an implicit admission that the mismatch is unresolved. This is a genuine gap, not a nitpick.\n\nMinor but worth saying: all tests use the true PSF model and true noise variance, only γ1 is validated (they argue γ2 should be comparable, but it's not shown), and the two-field noise-bias cancellation is imported from the authors' prior proof rather than re-derived here. Those are addressable.\n\nBottom line: for someone working on shear calibration for LSST, this is worth reading and citing as a promising method, and it deserves a serious referee. The referee should ask for simulations that vary galaxy properties between deep and wide inputs while keeping PSF/noise matched, plus tests with PSF misestimation and both shear components. I'd accept it with major revision, or maybe conditional acceptance, but not desk-reject.","headline":"Solid simulation-level demonstration of deep-field AnaCal, but the response-transfer assumption between different galaxy populations is untested.","tokens_in":15258,"tokens_out":2234,"would_cite":true,"duration_ms":22663,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A deep-field calibration strategy keeps 76 percent more galaxies for weak lensing.","keywords":["weak gravitational lensing","cosmic shear","shear calibration","multiplicative bias","deep-field imaging","analytical calibration","FPFS","LSST"],"falsifier":"A direct falsifier is a simulation where deep- and wide-field galaxy populations are deliberately mismatched in redshift or morphology while PSF and noise are matched; if the multiplicative bias then exceeds |m| > 3×10^-3, the response-transfer assumption fails. An observational check would overlay wide and deep images of the same patch and compare the deep-field AnaCal response against one measured directly from the wide-field data with matched noise.","tokens_in":14299,"feed_emoji":"🔭","tokens_out":6588,"duration_ms":64295,"temperature":0.7,"pith_summary":"This paper introduces a way to calibrate weak gravitational lensing shear measurements that keeps nearly all the statistical power of wide survey images instead of sacrificing a third of it. Standard analytical calibration requires adding an extra layer of noise to every image to cancel a noise bias, which doubles the pixel noise and lowers the effective depth. The new method measures galaxy shapes on wide-field images but computes the needed shear response from separate deep-field images, whose noise is doubled instead. In LSST-like simulations the multiplicative bias stays below 3×10^-3 at 99.7 percent confidence, and the effective galaxy number density rises from 17 to 30 per square arcmin. If real surveys behave like these simulations, this is a practical route to meeting the ten-year LSST calibration requirement without giving up depth.","feed_headline":"Deep-field calibration keeps 76% more galaxies for weak lensing","feed_subtitle":"Meets LSST's sub-percent shear bias target while lifting effective galaxy counts from 17 to 30 per arcmin^2.","key_machinery":"The central object is the deep-field AnaCal shear estimator, γ̂_α = ⟨ẽ_dacal, α⟩ / ⟨R̃_dacal, α⟩ + O(γ³), built from two matched noise-added observables: the wide-field ellipticity ẽ_dacal = e(ν̃_wide) and the deep-field response R̃_dacal. The identity that carries the argument is the noise-matching construction of Eqs. (14) and (15): deconvolving each field by its own PSF and reconvolving by a common larger PSF, then adding rotated noise realizations from the other field, gives both fields the same effective PSF and the same noise variance σ²_wide + 2σ²_deep. That lets the linear shear response of the deep field be used to calibrate the wide-field shape, with bias corrections derived analyt","core_discovery":"The paper's central claim is that shear can be estimated by the ratio of a wide-field measured ellipticity to a deep-field measured response, γ̂_α = ⟨ẽ_dacal, α⟩ / ⟨R̃_dacal, α⟩ + O(γ³), with both sides renoised so that the PSF and noise statistics of the two fields match. When the wide-field shape is measured on a noisy image and the response is measured on a deeper image, the noise-bias cancellation that standard AnaCal achieves by doubling wide-field noise is preserved, but the doubled noise is moved to the deep field, where it costs less. The claim is that this estimator is unbiased at the level required by ten-year LSST surveys — |m| < 3×10^-3 at 99.7 percent confidence — in both isolat","pith_inferences":["A natural next step the paper leaves implicit is to compute redshift-dependent responses from deep-field galaxies in redshift bins, which would extend deep-field AnaCal to tomographic shear without re-introducing wide-field noise.","The same noise-matching construction could be applied to multi-band or multi-epoch combinations, where the deep field is built from coadded visits; the benefit would grow when the wide-field noise is high relative to the deep field.","One testable consequence is that the statistical gain should track the ratio sqrt(2/(1+2r)) with r the deep-to-wide noise variance ratio; simulating other ratios would confirm the scaling and help observers choose deep-field exposure time.","If deep-field galaxy properties differ from wide-field ones in ways beyond sample variance — for example in redshift or morphology — the response transfer could become biased; reweighting deep galaxies to match wide-field distributions is a plausible fix that the same simulations could test."],"forward_implications":["Surveys using deep-field AnaCal can recover about 76 percent more effective galaxies (n_eff from 17 to 30 arcmin^-2 in ten-year LSST r-band simulations) than standard AnaCal on wide-field images, directly improving cosmic shear statistical precision.","With a ten-times deeper field, pixel noise variance in the shear estimate drops by about 30 percent and overall shear uncertainty by about 25 percent relative to standard wide-field AnaCal.","The multiplicative bias remains within |m| < 3×10^-3 at 99.7 percent confidence across isolated, blended, variable-noise, variable-PSF, and bulge+disk galaxy simulations, meeting the nominal ten-year LSST calibration requirement.","Sample variance from realistic LSST Deep Drilling Field coverage contributes an equivalent calibration uncertainty of ≲0.3 percent, small enough to stay within the LSST error budget.","The method generalizes the deep-field response idea to blended galaxies and shear-dependent detection bias, which earlier deep-field metacalibration did not cover."],"supporting_citations":[{"why":"Supplies the AnaCal noise-bias cancellation proof and the renoised response estimator that deep-field AnaCal modifies.","marker":"Li et al. 2024a"},{"why":"Defines the FPFS linear observables, detection and selection weight responses, and the shear-response chain the deep-field estimator inherits.","marker":"Li & Mandelbaum 2023"},{"why":"Introduces deep-field metacalibration, the direct precursor concept of computing the response from deep data; this paper extends it to AnaCal and blending.","marker":"Zhang et al. 2023"},{"why":"Provides the response-based shear estimation formalism and the ratio estimator that deep-field AnaCal adapts.","marker":"Sheldon & Huff 2017"},{"why":"Supplies the descwl-shear-sims blended-image simulation pipeline used for the blended validation.","marker":"Sheldon et al. 2023"},{"why":"Supplies the large-scale-structure-correlated galaxy catalog used to estimate the sample-variance scatter in the deep-field response.","marker":"OpenUniverse et al. 2025"},{"why":"Supplies the COSMOS HST galaxy catalog used to render isolated-galaxy simulations.","marker":"Mandelbaum et al. 2019"},{"why":"Sets the |m|<3×10^-3 multiplicative-bias requirement that defines the validation target.","marker":"The LSST Dark Energy Science Collaboration et al. 2018"},{"why":"Provides the LSST Deep Drilling Fields area and observing strategy used for the sample-variance assessment.","marker":"Bianco et al. 2025"}],"fun_headline_variants":["Deep-field calibration lifts weak-lensing galaxy yield by 76%","Deep-field method cuts shear uncertainty 25% for LSST","Deep-field AnaCal meets LSST bias target, adds 76% more galaxies","Deep-field calibration cuts shear noise 30% and boosts galaxy counts 76%"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The deep-field galaxies used to measure the shear response are assumed to represent the wide-field galaxies whose shapes are being calibrated, with any difference appearing only as random sample variance; the simulations always supply the true PSF and noise, so real-world PSF or population mismatches between the two fields are not tested.","fun_headline_variants_meta":{"raw":{"variants":["Deep-field calibration lifts weak-lensing galaxy yield by 76%","Deep-field method cuts shear uncertainty 25% for LSST","Deep-field AnaCal meets LSST bias target, adds 76% more galaxies","Deep-field calibration cuts shear noise 30% and boosts galaxy counts 76%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000959,"raw_usage":{"total_tokens":3994,"prompt_tokens":884,"completion_tokens":3110,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":3030}},"tokens_in":628,"tokens_out":3110,"duration_ms":24847,"temperature":1.0,"reasoning_tokens":3030,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:32:44.667099+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct falsifier is a simulation where deep- and wide-field galaxy populations are deliberately mismatched in redshift or morphology while PSF and noise are matched; if the multiplicative bias then exceeds |m| > 3×10^-3, the response-transfer assumption fails. An observational check would overlay wide and deep images of the same patch and compare the deep-field AnaCal response against one measured directly from the wide-field data with matched noise.","supporting_citations":[{"cited_title":"Deep-field Metacalibration","cited_arxiv_id":"2206.07683","evidence_quote":"Introduces deep-field metacalibration, the direct precursor concept of computing the response from deep data; this paper extends it to AnaCal and blending."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the COSMOS HST galaxy catalog used to render isolated-galaxy simulations."},{"cited_title":"The Rubin Observatory Survey Cadence Optimization Committee","cited_arxiv_id":null,"evidence_quote":"Provides the LSST Deep Drilling Fields area and observing strategy used for the sample-variance assessment."}],"review_version":1}