{"id":"44d5a4ad-0082-42b3-a2ab-eb8c9c9434ae","arxiv_id":"2501.17932","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A corrected pseudo-power spectrum formalism and a random forest masking method let CIBER recover unbiased sky fluctuation power with residual flat field errors below 20% of the uncertainty.","lead":"The CIBER collaboration presents a new way to correct flat field errors in rocket-based near-infrared background measurements, plus a machine-learning method to mask stars and galaxies more deeply. Both are tested on thousands of simulated observations and shown to recover the input sky fluctuations without bias.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Masking errors are excluded from the power-spectrum validation, and Table 2 shows deep-mask shot-noise biases exceeding the abstract's <10% claim, so the unbiased-recovery claim is not yet demonstrated end-to-end.","rationale":"The paper is methodologically careful: it derives an explicit FF-bias formalism, builds realistic mocks with read/photon noise, laboratory FF structure, ZL, ISL, DGL and EBL components, and validates the multiplicative bias correction on 1000 mock sets. I read the central claim as the combination of (i) unbiased sky-power recovery with the pseudo-C_ell pipeline, and (ii) <10% shot-noise errors from the new masking method. The second claim is contradicted by the paper's own Table 2 at the deepest cuts, and the first claim is not tested with any masking errors at all. Section 7's explicit assumption of perfect masking knowledge means the pipeline validation and the masking validation are never joined end-to-end. The reader's weakest_assumption correctly identified that masking errors are untested; the additional concrete point here is that the abstract's <10% statement is already in tension with Table 2, which strengthens the concern from 'untested' to 'internally inconsistent for the deepest masks.' This is addressable by adding an end-to-end mock test, and does not invalidate the formalism or the usefulness of the method; it does mean the headline claims are stronger than the presented evidence. A CONDITIONAL verdict remains appropriate, so I recommend no change to the reader's verdict.","tokens_in":31413,"tokens_out":3675,"duration_ms":40677,"concrete_test":"Rerun the Section 7 mock power-spectrum validation with masking errors emulated: for each mock, generate the mask from the random-forest predicted catalog applied to the simulated TRILEGAL+Helgason sources, including the incompleteness and magnitude scatter quantified in Table 2 and Figure 5 (e.g., randomly drop 2-8% of sources near threshold and perturb predicted magnitudes by the validation RMS), rather than from perfect knowledge of the injected sources. Then compare the recovered field-averaged C_ell against the input sky C_ell over 500 < ell < 2000. If the mean recovery remains within the mock dispersion and the FF-error uncertainty ratio stays below 20%, the concern is resolved; if residual shot-noise biases appear or the FF-error budget is exceeded, the Section 8 and abstract claims must be qualified to the tested masking depths.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the pipeline \"can recover unbiased power spectra\" is validated in Section 7 only under the stated assumption \"we do not directly emulate masking errors\" (Section 7). The source-masking method is validated separately, at catalog level, in Section 6.3, where Table 2 reports fractional shot-noise biases delta C_SN/C_SN of -5.2% to -21.7% for J-band and -2.2% to -16.4% for H-band on the COSMOS test set. The abstract says shot-noise errors remain below <10% \"at all masking depths considered,\" but J<19.0 gives -21.7% and H<18.5 gives -16.4%. These are not small compared with the 10% claim, and they are never propagated through the FF-corrected power-spectrum pipeline. Because the FF stacking estimator uses source masks, and Section 7.5 states that the matrix formalism breaks down in the presence of bright unmasked point sources, realistic mask incompleteness could couple into FF errors and bias recovered C_ell on precisely the scales where the paper claims <20% FF-induced uncertainty growth (500 < ell < 2000). The mock recovery tests therefore establish unbiasedness conditional on perfect masks, not for the actual masking implementation that will be used on CIBER data. This is an internal tension between the abstract and Table 2, not a disagreement with external consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This methodology paper presents the analysis framework for the fourth CIBER flight, with two main innovations: a pseudo-power-spectrum formalism that corrects for additive and multiplicative biases from an in-flight flat-field (FF) stacking estimator, and a random-forest-based source masking method that uses PanSTARRS and unWISE photometry to predict J- and H-band magnitudes, allowing deeper masking than 2MASS alone. The authors validate the pipeline on 1000 synthetic CIBER observations that include realistic sky signals, noise, masks, filtering, and injected lab-derived FFs. They report unbiased recovery of sky fluctuations except on the smallest angular scales, with residual FF errors increasing uncertainties by less than 20% on scales 500 < ell < 2000, and shot-noise errors below 10% at all masking depths considered. The paper is written as a methods paper preceding a companion analysis of CIBER data.","tokens_in":31755,"tokens_out":6980,"duration_ms":69702,"significance":"If the claims hold, this is a valuable methodological contribution for CIBER and for future NIR intensity-mapping experiments such as CIBER-2 and SPHEREx. The derivation in Appendix A is careful and the use of a large mock ensemble to quantify biases, covariances, and field weights is a clear strength. The random-forest masking approach is well motivated and includes an out-of-sample test on COSMOS, which is a useful check of distribution shift. The paper also correctly identifies and quantifies several non-trivial effects, such as the coupling of FF errors with masks and the need to include filtering in the mode-mixing matrix. However, the central 'unbiased recovery' claim is currently stated more strongly than the evidence supports, because the mock tests assume perfect masking and because several known residual biases are acknowledged in the text. The tension between the abstract's <10% shot-noise claim and the deeper-mask entries in Table 2 also needs to be resolved before the paper is ready for publication.","major_comments":[{"comment":"The abstract states that shot-noise errors remain below <10% 'at all masking depths considered,' but Table 2 reports fractional shot-noise biases of -21.7% for J<19.0 and -16.4% for H<18.5, and the text in Section 6.3 acknowledges 22% and 16% departures at the deepest cuts. These deepest cuts are precisely where the claimed two-magnitude improvement over 2MASS is demonstrated, so the <10% claim is not supported as written. Please either correct the abstract and Section 6.3 to state the depth-dependent range and explicitly report the deepest-cut values, or revise the masking method so that the shot-noise errors are below 10% at all depths claimed.","section":"Abstract and Section 6.3, Table 2"},{"comment":"The mock recovery tests explicitly assume perfect knowledge of source masking: 'we assume perfect knowledge for source masking, i.e., we do not directly emulate masking errors.' The central conclusion that the pipeline 'can recover unbiased power spectra' is therefore conditional on perfect masks. The source-masking method is validated separately at catalog level in Section 6.3, but the fractional shot-noise biases measured there are never propagated through the FF-corrected pseudo-C_ell pipeline. Since the FF stacking estimator depends on the masks (Section 5.2.2) and Section 7.5 states that bright unmasked point sources break the matrix formalism, mask incompleteness or impurity could couple into FF errors and bias C_ell on exactly the scales where the paper claims <20% FF-induced uncertainty growth (500 < ell < 2000). The validation would be complete if masking errors were injected into the mock pipeline using the measured completeness and purity of the predicted catalogs, or if a quantitative propagation of the Section 6.3 shot-noise errors to recovered C_ell were provided.","section":"Section 7"},{"comment":"The paper claims unbiased recovery 'for all but the smallest angular scales,' yet Section 7.2 reports a negative bias at the fifth bandpower at the 1-2 sigma level in both the delta[FF]=0 and delta[FF]!=0 cases, and a positive bias at ell>50000 in the delta[FF]!=0 case. The fifth bandpower is not one of the smallest angular scales, so the Section 8 claim is not supported as stated. Please quantify these biases (amplitude relative to statistical error, field dependence, and whether they persist with more realizations) and either adjust the conclusions to list these exceptions or reduce the biases with additional corrections.","section":"Section 7.2 and Section 8"},{"comment":"The varying-masking-depth analysis relies on an empirical switch between M^{mask+filter} and M^{mask+filter+FF} at (Jlim,Hlim)=15 because the matrix formalism breaks down with bright unmasked point sources. This is a reasonable pragmatic choice, but it means the FF bias correction is not exact for shallow cuts, and the text notes a slight underestimation at (Jlim,Hlim)=16. The robustness of the large-angle science results to this approximation should be stated explicitly in the conclusions, since the companion paper will use these masks and readers may otherwise infer that the full pipeline is uniformly validated across all masking depths.","section":"Section 7.5"}],"minor_comments":[{"comment":"There is a typo in the sentence describing the third-bandpower bias: 'however this is not seen in the The bias is not delta[FF] != 0 case' should be 'however this is not seen in the delta[FF] != 0 case.'","section":"Section 7.2"},{"comment":"The sentence 'In practice we use the J < 17.5 and H < 17.0 masks to calculate FFhat for all shallower masking cuts)' contains an unmatched parenthesis; please correct it.","section":"Section 7.5"},{"comment":"The column headers for completeness and purity are difficult to parse (the repeated 'C, P' groups). Please define each subcategory (e.g., 'predicted', 'PS+unWISE', 'PS only', 'unWISE only', '2MASS only') with a clear row/column structure, and state the units of the delta C_SN/C_SN column explicitly.","section":"Table 2"},{"comment":"Equation (30) is presented in the main text as the ratio hat C_{ell,j} / C^{true}_{ell,j}, but the derivation in Appendix A.2.2 defines this ratio after noise-bias subtraction. Please clarify in the text that Eq. (30) is the multiplicative factor that applies to the noise-debiased power spectrum, not the full observed-to-true ratio.","section":"Section 5.2.2, Eq. (30)"}],"recommendation":"major_revision","confidential_remarks":"This is a solid methods paper with a careful derivation and an extensive mock campaign. The main issues are internal: the abstract overclaims the shot-noise performance relative to Table 2, and the 'unbiased recovery' claim is not yet demonstrated end-to-end because masking errors are excluded from the mock pipeline. Both are fixable: the first by rewording or by improving the deepest mask cuts, and the second by adding masking-error injection or by explicitly framing the unbiased-recovery claim as conditional on perfect masks. The paper is within the scope of the journal and, after these revisions, would be a useful contribution for the CIBER and broader intensity-mapping community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague —\n\nWorth a look if you work on NIR extragalactic background light fluctuations or pseudo-C_l methods, but the abstract overstates one claim and the end-to-end validation has a gap. The new material is real: in-flight flat-field stacking estimator, its additive and multiplicative bias corrections, and a single mode-coupling matrix that folds in FF errors, filtering, and masks. The derivation in App. A is careful, and the Monte Carlo recipe is concrete enough to reproduce. The random forest masking technique, trained on UKIDSS UDS and tested on COSMOS, is a clean new application and reaches roughly two magnitudes deeper than 2MASS at catalog level. Mock validation is extensive (1000 realizations, multiple masking depths, per-field and averaged recovery) and the paper is honest about residual biases at the fifth bandpower and at l>50000 in the deltaFF case.\n\nThe soft spots are real but addressable. Section 7 states the mock tests assume perfect knowledge for source masking and do not directly emulate masking errors, so the central 'unbiased recovery' claim is demonstrated only for perfect masks. Separately, the shot-noise errors from the masking catalog on the COSMOS test set (Table 2) range from -5% to -22% in J and -2% to -16% in H, exceeding 10% at the deepest cuts. That is in direct tension with the abstract's claim that shot noise errors stay below 10% at all masking depths considered. The deepest cuts are not the fiducial science cuts, and the text does say 'slightly larger' for those cases, but the abstract is simply wrong as written. More importantly, those catalog-level masking errors are never propagated through the power-spectrum pipeline, even though the FF stacking estimator depends on the masks. Fix either by simulating masking errors or by softening the claims.\n\nThe citation pattern and framing are fair: MASTER lineage and Z14 handling are sound, and the ISL/DGL caveat on the common-spectrum assumption is stated rather than buried. No code or data release, which is a minor limitation for a methods paper.\n\nThis is for people building CIBER-2, SPHEREx, and LIBRAE analyses or PS-based intensity mapping pipelines. It deserves serious referee time; publish after the abstract is corrected and the masking-error caveat is made consistent with Table 2.","headline":"Careful FF-corrected pseudo-C_l formalism and a useful masking technique, but the abstract's <10% shot-noise claim contradicts Table 2 and the unbiased-recovery validation skips masking errors.","tokens_in":32310,"tokens_out":4205,"would_cite":true,"duration_ms":38950,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that CIBER can recover unbiased near-infrared sky fluctuation power spectra from single fields by correcting flat-field errors in the pseudo-power-spectrum domain and by masking point sources two magnitudes deeper than…","keywords":["cosmic infrared background","extragalactic background light","intensity mapping","angular power spectrum","flat field calibration","source masking","random forest regression","CIBER"],"falsifier":"Run the pipeline on mocks in which the off-field sky fluctuations are drawn from a different power spectrum than the target field, such as one field with suppressed large-scale diffuse galactic light, and check whether the recovered target power spectrum remains unbiased after the multiplicative correction; alternatively, inject the masking errors quantified in Section 6, such as 0.25-pixel astrometric scatter or the ten-to-twenty percent completeness gaps, directly into the mock validation and see whether the claimed unbiased recovery on scales $500<\\ell<2000$ survives.","tokens_in":31226,"feed_emoji":"🌌","tokens_out":5221,"duration_ms":47082,"temperature":0.7,"pith_summary":"This paper argues that the two leading systematics in CIBER's measurement of extragalactic background light fluctuations, imperfect flat-field calibration and limited source-masking depth, can be corrected well enough to recover unbiased sky power spectra from individual fields rather than from field differences. The authors develop a pseudo-power-spectrum formalism that splits flat-field errors into an additive noise bias and a multiplicative bias, and they correct both using Monte Carlo mode-mixing matrices. They also train random forest regressors on deep UKIDSS photometry to predict J- and H-band magnitudes from PanSTARRS and unWISE colors, yielding masking catalogs about two magnitudes deeper than 2MASS alone with shot-noise errors below ten percent. On one thousand mock realizations of the fourth flight, the pipeline recovers unbiased power spectra on all but the smallest angular scales, with residual flat-field errors inflating uncertainties by less than twenty percent on scales $500 < \\ell < 2000$. If correct, this removes the sensitivity penalty of field differencing and opens the same analysis path for future near-infrared intensity-mapping experiments.","feed_headline":"CIBER pipeline stays unbiased while masking two magnitudes deeper","feed_subtitle":"Flat-field errors are corrected on pseudo-power spectra, keeping residual uncertainty below 20 percent on arcminute scales.","key_machinery":"The load-bearing object is the extended pseudo-$C_\\ell$ mode-mixing matrix $M_{\\ell\\ell'}$, which in this paper combines the survey mask, the flat-field stacking estimator, and the image filter into one linear operation. Additive flat-field noise bias is subtracted through modified Monte Carlo noise realizations that include mean sky levels and flat-field stacking; the multiplicative flat-field bias, which scales with the ratio of mean sky brightnesses between target and off-fields, is corrected analytically in the unmasked limit and included in the matrix for the masked case. The source-masking component is a random forest regressor trained on UKIDSS UDS photometry that maps PanSTARRS and unWISE magnitudes to predicted J and H magnitudes, with mask radii set by iteratively suppressing extended PSF power.","core_discovery":"The central claim is that flat-field errors, which previously forced CIBER to analyze differences between fields, can instead be corrected directly in the pseudo-power-spectrum domain. The flat field is estimated by stacking per-field sky flats from the four other science fields, and the resulting errors are propagated into two biases: an additive noise bias from instrument noise and a multiplicative bias of order $1+\\sum_i (w_i I_j/I_i)^2$ from sky fluctuations. Both are folded into a mode-mixing matrix that is estimated with Monte Carlo tone realizations, so masked, filtered, flat-field-corrected maps recover the input sky power spectrum after inversion. The paper further claims that random forest regression on PanSTARRS and unWISE photometry predicts J- and H-band magnitudes with more than ninety percent completeness and purity relative to UKIDSS UDS validation, allowing masks two magnitudes deeper than 2MASS completeness while keeping fractional shot-noise errors below ten percent. Mock tests with injected laboratory flat fields demonstrate unbiased recovery for all but the smallest angular scales and quantify the residual flat-field penalty as less than twenty percent on $500<\\ell<2000$.","pith_inferences":["A natural next test is to inject realistic masking errors, such as position noise, magnitude scatter, and catalog incompleteness, into the mocks; the current validation assumes perfect mask knowledge, so the unbiased-recovery claim has not yet been stress-tested against the masking systematics the paper itself characterizes.","The multiplicative-bias formula suggests that in surveys with large field-to-field sky-brightness variation, the flat-field stacking estimator could be redesigned to down-weight bright fields, reducing the bias rather than correcting it after the fact.","The random forest magnitude predictions could be turned into a de-projection method that subtracts rather than masks bright sources, which would preserve more Fourier modes on small scales.","For future wide-area surveys, the common-spectrum assumption underlying the flat-field correction will likely need to be replaced by a forward model that marginalizes over variations in diffuse galactic light and integrated stellar light across fields."],"forward_implications":["Single-field power spectra become usable, so the effective mask is no longer the union of two field masks and masking can be more aggressive.","Residual flat-field error contributes less than twenty percent to the power-spectrum uncertainty on arcminute scales, so the fourth-flight dataset gains sensitivity without field differencing.","Masking two magnitudes deeper reduces Poisson shot noise from unmasked sources while keeping shot-noise errors below ten percent at all tested depths.","The pipeline yields field-averaged power spectra and covariances from mock ensembles that can be used to test field-to-field consistency in the real data.","The formalism extends directly to cross-power spectra, with an analogous multiplicative flat-field bias correction, and to future instruments with similar imaging characteristics."],"supporting_citations":[{"why":"Prior CIBER fluctuation analysis whose field-differencing approach this paper replaces; provides the baseline signal, masking depth, and method being improved.","marker":"Zemcov et al. 2014"},{"why":"MASTER formalism that supplies the mode-mixing matrix framework extended here to include flat-field errors and filtering.","marker":"Hivon et al. 2002"},{"why":"Semi-empirical luminosity functions used to generate IGL mock galaxy realizations and predicted source counts for masking tests.","marker":"Helgason et al. 2012"},{"why":"Zodiacal light model used to set mean sky brightness levels and gradients in the mock observations.","marker":"Kelsall et al. 1998"},{"why":"TRILEGAL stellar population model used to simulate integrated stellar light and to test the efficacy of source masking.","marker":"Girardi et al. 2005"},{"why":"Field-dependent PSF models used to inject sources into mocks and to compute beam transfer functions.","marker":"Cheng et al. 2021"},{"why":"2MASS catalog that defines the shallow masking baseline and supplies bright-end sources merged into the final masks.","marker":"Skrutskie et al. 2006"},{"why":"COSMOS 2015 catalog used as an independent test of distribution shift for the random forest magnitude predictions.","marker":"Laigle et al. 2016"}],"fun_headline_variants":["CIBER's pseudo-spectra fix flat fields, mask deeper","Deep masks, flat-field fixes: CIBER's unbiased sky power","CIBER beats 2MASS masking, cuts flat-field error","Pseudo-power pipeline corrects flat fields, masks deeper","CIBER: flat-field <20% uncertainty with deep masking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The multiplicative flat-field correction assumes that the sky fluctuations in the off-fields are drawn from the same underlying power spectrum as the target field and that foreground point sources are removed perfectly, so any field-to-field difference in the foreground spectrum or any masking error enters the final power spectrum uncorrected.","fun_headline_variants_meta":{"raw":{"variants":["CIBER's pseudo-spectra fix flat fields, mask deeper","Deep masks, flat-field fixes: CIBER's unbiased sky power","CIBER beats 2MASS masking, cuts flat-field error","Pseudo-power pipeline corrects flat fields, masks deeper","CIBER: flat-field <20% uncertainty with deep masking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000891,"raw_usage":{"total_tokens":3869,"prompt_tokens":997,"completion_tokens":2872,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":2784}},"tokens_in":613,"tokens_out":2872,"duration_ms":25728,"temperature":1.0,"reasoning_tokens":2784,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:31:05.790263+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on mocks in which the off-field sky fluctuations are drawn from a different power spectrum than the target field, such as one field with suppressed large-scale diffuse galactic light, and check whether the recovered target power spectrum remains unbiased after the multiplicative correction; alternatively, inject the masking errors quantified in Section 6, such as 0.25-pixel astrometric scatter or the ten-to-twenty percent completeness gaps, directly into the mock validation and see whether the claimed unbiased recovery on scales $500<\\ell<2000$ survives.","supporting_citations":[{"cited_title":"2021, ApJ, 919, 69, doi: 10.3847/1538-4357/ac0f5b","cited_arxiv_id":null,"evidence_quote":"Field-dependent PSF models used to inject sources into mocks and to compute beam transfer functions."}],"review_version":1}