{"id":"413c98c8-3175-4cdd-9700-9753794b7347","arxiv_id":"2501.04085","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The CEERS survey overview shows that coordinated parallel JWST observations in the EGS field work as designed and have generated a public legacy dataset that enabled extensive early-universe science.","lead":"This paper is the official overview of the CEERS program, a 77.2-hour JWST survey of the Extended Groth Strip combining NIRCam, MIRI, NIRSpec, and NIRCam grism observations. It documents the survey design, public data releases, measured sensitivities, and two years of science highlights, and reports that the CEERS data have produced over 170 papers and more than 7,500 citations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Imaging depth validation lacks an independent fake-source recovery check, so the aperture/PSF-corrected 5-sigma depths in Table 5 are the least secure link in the survey-validation claim.","rationale":"The reader identified the measured 5-sigma depths as the weakest assumption, and my reading agrees: the paper's central descriptive claim (that CEERS validated efficient parallel surveys and reached its design depths) depends on these numbers. I have narrowed the concern to a specific technical gap: for imaging, the depths are estimated from aperture noise plus a PSF curve-of-growth correction, with no independent fake-source recovery experiment. This is the load-bearing step because the correction is large enough to matter and because no cross-check is presented to show the correction is unbiased. The concern is not that the paper is internally inconsistent; the depth measurements are plausible and the paper is candid about known failures. Rather, the validation would be materially strengthened by an injection/recovery test, and the absence of such a test is the most concrete place where the central claim could fail. The spectroscopic mock-line injection procedure is more self-contained and is not the primary weak point. The unfinished citation in Appendix C and assorted typos are editorial issues that do not affect the central argument. Since the reader already recommends conditional acceptance and my concern does not overturn that recommendation, the verdict should remain unchanged.","tokens_in":51154,"tokens_out":4751,"duration_ms":53007,"concrete_test":"Run a fake-source recovery test on the released v1.0 NIRCam and MIRI mosaics: inject thousands of point sources with known fluxes bracketing the quoted 5-sigma depths (e.g., 28.5-29.5 AB in F277W, and roughly 25.5-26.5 AB in F770W) into source-free regions, then run the same photometric pipeline used to generate the catalogs and measure the recovered flux bias and 50% completeness. If the median recovered flux or the 50% completeness limit deviates by more than about 0.1 mag from the Table 5 point-source depths, the aperture/PSF correction is biased and the survey depth validation is not secure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central validation claim is that CEERS reached its design sensitivities (Section 4.1, Tables 5 and 6). For NIRCam and MIRI imaging, the quoted point-source 5-sigma depths are not measured by injecting and recovering sources; they are inferred from the noise in small fixed apertures (0.2-arcsec diameter for NIRCam, PSF-FWHM apertures for MIRI) and then corrected to total using a PSF curve-of-growth. This two-step estimator inherits any error in the PSF model or in the assumption that the target population is unresolved, and the correction factor directly scales the reported depth. The paper shows mock-line injection for spectroscopy but no analogous fake-point-source test for imaging. The internal consistency check in Figure 8 suggests the measured depth is sensitive to reduction details, and the MIRI point-source depths exceed the catalog-median depths by roughly 1.3-1.5 mag in F560W/F770W, so the correction is not negligible. If the encircled-energy fraction used for the correction is biased, the reported depths do not accurately represent survey sensitivity, and the downstream claims about design-goal validation are weakened.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is the overview paper for the CEERS ERS program. It describes the survey design and executed layout for coordinated NIRCam/MIRI imaging and NIRSpec/NIRCam WFSS spectroscopy in the Extended Groth Strip, documents the team's data releases, reports empirically measured depths for all observing modes, summarizes early science results from the team and the community, and quantifies the publication impact of the dataset. The paper's central claim is that CEERS demonstrates, tests, and validates efficient coordinated-parallel extragalactic survey operations with JWST and reaches or modestly exceeds its pre-launch design sensitivities.","tokens_in":51296,"tokens_out":12495,"duration_ms":116491,"significance":"If the validation claim holds, this is a valuable legacy paper for one of the most heavily used public JWST datasets. The strengths include empirical depth measurements derived from the released data, mock emission-line injection for the spectroscopic depths, transparent reporting of known failures (the MSA electrical short and the z~16 interloper), public data releases with reproducibility notebooks, and a quantitative publication-impact analysis. The depth measurements contain no fitted parameters and are compared with pre-launch predictions and external noise statistics, so the circularity risk is low. The main weakness is that the imaging depths rest on aperture-noise measurements plus PSF curve-of-growth corrections without an independent source-injection/recovery test, which leaves the absolute imaging depths as the least independently verified element of the survey-validation evidence.","major_comments":[{"comment":"The point-source imaging depths in Table 5 are not measured by injecting and recovering fake sources; they are inferred from the noise in fixed apertures (0.2-arcsec diameter for NIRCam, PSF-FWHM apertures for MIRI) and then corrected to total flux using a PSF curve-of-growth. This correction directly scales the reported depth, and the abstract and §6 use these depths to claim that the survey 'reaches' and 'validates' its design sensitivity. Because no analogous fake-source test is shown for imaging, unlike the mock-line test for spectroscopy in the same section, please add an injection/recovery test or, failing that, a quantitative uncertainty budget for the PSF total-flux correction, and adjust the validation wording accordingly. The MIRI point-source versus catalog-median differences in F560W/F770W are plausibly explained by the catalog sources being resolved, so I do not treat that comparison as evidence of bias; the missing recovery test is the substantive gap.","section":"§4.1, Table 5, and §6"},{"comment":"The NIRSpec continuum and emission-line depths in Table 6 and Figure 7 are based on DR0.7 products reduced with 'custom procedures' and 'custom aperture extractions and masking of detector artifacts,' which are described only as forthcoming in Arrabal Haro et al. (in prep). Since the spectroscopic depth measurements and the reproducibility of the data release are part of the survey-validation claim, please include a concise description of these procedures in this paper or point to a released, citable notebook or software version, rather than relying solely on an in-preparation reference.","section":"§4.2.7 and Table 6"}],"minor_comments":[{"comment":"Table 3 is inconsistent with the text: the text says MIRI pointings 1 and 2 include F770W with 1648 s of exposure, but the table shows F770W blank for those pointings and lists 1648.4 s under F1000W, and the F2100W value in the table (4811.9 s) differs from the text (4757 s). Please make the table and text consistent.","section":"Table 3 and §3.4"},{"comment":"The reference list contains two entries, Barro et al. 2024a and 2024b, with identical journal, volume, page, and DOI; one of these entries is likely incorrect and should be corrected.","section":"References"},{"comment":"Appendix C.1 contains the unresolved placeholder 'Gaia-EDRS cite cite cite'; this should be replaced with the proper reference.","section":"Appendix C.1"},{"comment":"Section 5.4 repeats 'the galaxies the galaxies in their sample'; please remove the duplication.","section":"§5.4"},{"comment":"Section 3.5 says 'these three pointings' when referring to NIRSpec pointings 11 and 12; if the DDT pointing is also meant, say so explicitly, otherwise change the phrase to 'these two pointings'.","section":"§3.5"},{"comment":"The Figure 10 caption lists the z=5.61 broad-line AGN as Kocevski et al. (2023a), while the text in §5.2 cites Kocevski et al. (2023b) for this result; please harmonize the citation.","section":"Figure 10 caption and §5.2"},{"comment":"The claim that the achieved NIRCam depths are '~0.3–0.5 mag deeper' than the pre-launch expectation of ~28.7 is not representative of all filters: F410M is 28.7, equal to the quoted expectation, while F277W is 29.5. Please make the comparison filter-specific or quote a range that actually describes the full filter set.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"For the editor: this is a strong and useful survey overview with a low circularity risk, and the requested changes are tractable. The main technical gap is the imaging-depth validation; adding a recovery test or an explicit uncertainty budget would make the survey-validation claim secure. I also recommend cleaning up the Table 3 inconsistencies and the duplicate Barro reference before acceptance. No citation-pattern concerns: the paper represents both team-led and community-led work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the definitive CEERS overview and the paper people will cite when they use the public data. What's genuinely new: the survey design rationale for coordinated parallel observations, the v1.0 reduction details (2D background subtraction, three-step 1/f correction, wisp templates, Snowblind persistence handling), the empirical depth measurements, and the publication-impact analysis. The team is candid about known failures (the MSA electrical short, the z~16 interloper) and ships notebooks with the data releases. That transparency is worth crediting.\n\nThe science highlights are summaries of prior papers, not new results, which is fine for this genre. The publication statistics are simple ADS counts; they're illustrative, not rigorous, and the paper does not overclaim them.\n\nThe stress-test note identifies a real but modest soft spot. The NIRCam and MIRI point-source depths in Table 5 are measured from noise in small fixed apertures and corrected to total using a PSF curve of growth, without an independent fake-source recovery check. The spectroscopic depths use mock-line injections, so the imaging validation is the less secure link. That said, this is a standard and generally reliable technique, and the gap between point-source and catalog-median depths is largely expected because the catalog sources are mostly resolved. I would not call it a load-bearing flaw; I would call it a caveat the authors should state explicitly in the depth section.\n\nThe more concrete problems are mechanical. Appendix C.1 contains an unfinished reference (\"Gaia-EDRS cite cite cite\"), and there are typos like \"official official\" and \"red paintings\". Those need to be fixed before this is archival. The reference list also has at least one apparent duplicate (Barro et al. 2024a/b point to the same DOI), which should be checked.\n\nOverall, the central validation claim—that CEERS reached its design goals and the released data support the reported science—holds up. This deserves a serious referee and, after minor revision, publication. For anyone planning a JWST survey, it's a useful template; for anyone using CEERS data, it's the citation.","headline":"A solid, citable CEERS overview with a real but minor gap in imaging depth validation; acceptable after fixing the unfinished citation.","tokens_in":52404,"tokens_out":2812,"would_cite":true,"duration_ms":28957,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CEERS shows that coordinated-parallel JWST observations reach design depths and seed broad extragalactic science.","keywords":["early universe","galaxy formation","galaxy evolution","JWST survey","parallel observations","NIRCam","NIRSpec","MIRI"],"falsifier":"Independently reduce a subset of the same raw CEERS data (for example, one NIRCam pointing and one NIRSpec grating) with a different pipeline, re-measure the $5\\sigma$ point-source depth in 0.2-arcsecond apertures and the recovered mock-line flux in the same spectra, and compare to Tables 5 and 6; a systematic discrepancy larger than the quoted uncertainties would falsify the claimed validation.","tokens_in":50936,"feed_emoji":"🔭","tokens_out":10398,"duration_ms":87601,"temperature":0.7,"pith_summary":"This paper presents the Cosmic Evolution Early Release Science Survey (CEERS), a 77.2-hour JWST program built as a demonstration, test, and validation of efficient extragalactic surveys that use coordinated parallel observations across four instrument modes: NIRCam imaging, MIRI imaging, NIRSpec multi-object spectroscopy, and NIRCam slitless grism spectroscopy. The authors aim to show that such parallel operations can deliver the imaging depth and spectral coverage needed for the two core JWST science drivers, 'First Light' and 'Galaxy Assembly,' and that the resulting public data releases support a broad community science program. They report measured $5\\sigma$ depths modestly deeper than pre-launch expectations, list data releases and reproducibility tooling, and summarize science highlights from the first two years, including $z>10$ galaxy candidates, deep spectra of more than a thousand galaxies, and the first barred spirals at $z>2$. If the validation holds, CEERS becomes a template for future JWST survey programs and a benchmark dataset for early-universe studies.","feed_headline":"Parallel JWST survey hits depth goals, sparks 174 papers","feed_subtitle":"A 77.2-hour Early Release Science program validates efficient multi-instrument observing and opens z>10 galaxy science.","key_machinery":"The load-bearing object is the coordinated-parallel observing layout itself: pairs of prime and parallel JWST observations designed so that NIRCam imaging fills the mosaic while NIRSpec MSA and MIRI observe the same or overlapping footprints, plus NIRCam grism spectroscopy with MIRI in parallel. The argument that the survey 'worked' is carried by the empirical depth-measurement method in Section 4.1, which places random circular apertures in source-free regions to derive point-source $5\\sigma$ limits (corrected to total via encircled-energy fractions) and injects mock emission lines into real spectra to derive $5\\sigma$ line fluxes; these measurements appear in Tables 5 and 6 and Figures 6 and 7 and are compared to pre-launch expectations. The v1.0 data-reduction pipeline, with custom wisp, $1/f$, and background subtractions and astrometric alignment, is what turns raw data into the mosaics and spectra whose quality is being validated.","core_discovery":"CEERS demonstrates, tests, and validates efficient extragalactic survey operations with JWST by executing coordinated, overlapping parallel observations with NIRCam and MIRI imaging and NIRSpec and NIRCam slitless spectroscopy. On the survey's own terms, the program reached its design goals: ten NIRCam pointings covering about 90 square arcminutes reach point-source $5\\sigma$ depths of roughly 29 to 29.5 AB magnitudes across 1 to 5 microns; MIRI reaches about 26th magnitude at wavelengths below 10 microns; NIRSpec medium-resolution gratings reach emission-line sensitivities of about $1\\times10^{-18}$ to $2\\times10^{-18}$ erg s$^{-1}$ cm$^{-2}$; and the achieved depths are modestly deeper than pre-launch predictions. The paper further claims that the public data releases, documented reductions, and notebooks support a wide range of extragalactic science, including the discovery and spectroscopic confirmation of galaxies at $z>10$, spectra of more than 1000 galaxies, resolved structure and morphology studies at $z>3$, and MIRI-based characterization of obscured star formation and supermassive black hole growth, yielding more than 170 papers and 7500 citations within two years.","pith_inferences":["If the depth-validation approach is sound, the same empirical aperture-noise and mock-line-injection recipe could be adopted as a standard for quoting JWST survey depths, making different surveys' sensitivity limits directly comparable.","The high publication yield from a single early-release field suggests that allocating early observing time to a few public legacy fields can accelerate an entire subfield; the same logic would apply to future missions.","The two-epoch, MSA-rescheduling experience implies that parallel-survey designs should build in redundancy for instrument anomalies, a lesson that generalizes beyond this program."],"forward_implications":["If the parallel-survey template is as efficient as reported, future JWST extragalactic surveys can expect comparable per-hour yield, making wide, multi-instrument programs a standard way to build legacy fields.","The published depths give the community reliable sensitivity priors for planning follow-up observations and for interpreting non-detections in the CEERS footprint.","The program's science highlights imply that JWST can both discover and spectroscopically confirm substantial samples of galaxies at $z>10$, sharpening constraints on the ultraviolet luminosity function at early times.","The combined NIRCam, MIRI, and NIRSpec dataset reduces degeneracies in stellar-population modeling, which the paper argues improves stellar mass and star-formation rate estimates for galaxies at $z\\sim4$ to 9.","The two-epoch scheduling and MSA rescheduling experience offers concrete lessons for how to design robust time-constrained parallel programs around observability windows."],"supporting_citations":[{"why":"Supplies the post-launch JWST performance baseline the survey compares its achieved depths against.","marker":"Rigby et al. 2023"},{"why":"Defines the CANDELS program and the EGS field that CEERS targets for its HST legacy data.","marker":"Grogin et al. 2011"},{"why":"Provides the CANDELS HST mosaics and astrometric reference used for CEERS field choice and multiwavelength matching.","marker":"Koekemoer et al. 2011"},{"why":"Describes the CEERS NIRCam imaging reduction and quality analysis that the v1.0 release and depth measurements build on.","marker":"Bagley et al. 2023"},{"why":"Describes the CEERS MIRI imaging reduction, point-source measurements, and catalogs used for the MIRI depths.","marker":"Yang et al. 2023a"},{"why":"Characterizes the NIRSpec MSA instrument and its spectroscopic modes, the subject of the survey's validation.","marker":"Jakobsen et al. 2022"},{"why":"Characterizes the NIRCam instrument and its imaging and grism modes used in the survey.","marker":"Rieke et al. 2023a"},{"why":"Characterizes the MIRI instrument and its imaging mode used in the survey.","marker":"Wright et al. 2023"},{"why":"Presents the CEERS key-paper sample of $z\\sim9$–14 galaxy candidates and UV luminosity functions, a headline science result of the survey.","marker":"Finkelstein et al. 2023"}],"fun_headline_variants":["CEERS survey validates multi-instrument JWST observing","CEERS: 77-hour survey proves parallel JWST strategy","Deep JWST survey opens z>10 galaxy science","Efficient JWST survey validates, sparks 170+ papers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the premise that the empirically measured $5\\sigma$ depths for the imaging and spectroscopy accurately represent the true survey sensitivity; if the aperture and PSF corrections or the mock-line recovery method are biased, the validation of the survey's capabilities is weakened.","fun_headline_variants_meta":{"raw":{"variants":["CEERS survey validates multi-instrument JWST observing","CEERS: 77-hour survey proves parallel JWST strategy","Deep JWST survey opens z>10 galaxy science","Efficient JWST survey validates, sparks 170+ papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000338,"raw_usage":{"total_tokens":1965,"prompt_tokens":1140,"completion_tokens":825,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":756,"completion_tokens_details":{"reasoning_tokens":758}},"tokens_in":756,"tokens_out":825,"duration_ms":7302,"temperature":1.0,"reasoning_tokens":758,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:40:52.639452+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently reduce a subset of the same raw CEERS data (for example, one NIRCam pointing and one NIRSpec grating) with a different pipeline, re-measure the $5\\sigma$ point-source depth in 0.2-arcsecond apertures and the recovered mock-line flux in the same spectra, and compare to Tables 5 and 6; a systematic discrepancy larger than the quoted uncertainties would falsify the claimed validation.","supporting_citations":[],"review_version":1}