{"id":"a9cd5148-c837-4c46-9800-e1b0f139c204","arxiv_id":"2608.08717","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 3.5-year, 19-star campaign delivers updated MJy/sr calibration factors for all NIRCam imaging filters, coronagraphic masks, weak lenses, and subarrays, with typical scatter below 2 percent.","lead":"JWST's NIRCam camera now has a new, more complete flux calibration that turns raw counts into physical brightness units for all imaging modes, including previously uncalibrated weak-lens and coronagraph settings. The result matters because every JWST near-infrared image's brightness measurements depend on this conversion, and the new factors change some earlier values by up to 30 percent.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Absolute scale rests on unverified CALSPEC2 models; excluded faint standards are the only checks on a common-mode systematic.","rationale":"The paper is a careful, transparent calibration delivery with a public code repository and detailed tables; the internal consistency checks (scatter, source-type agreement, repeatability, subarray offsets) are well executed. The single most load-bearing assumption is that the CALSPEC2 models for the 15 retained stars define the absolute zero point. This is the one link in the chain that has no independent verification within the paper. The exclusion of the four discrepant stars is statistically defensible but removes exactly the measurements that could have flagged a common-mode systematic in the photometric chain. The reader's weakest assumption identifies this same link; my reading agrees. A recomputation with and without the excluded stars is a feasible, decisive check because the code and data are public. If the final factors shift by more than their quoted errors, the delivered pmap 1490 values are not robust to a defensible alternative analysis choice and the paper's absolute accuracy claim should be softened to internal consistency at <2% plus CALSPEC2 model accuracy. The verdict CONDITIONAL remains appropriate; no change needed.","tokens_in":36452,"tokens_out":8548,"duration_ms":94502,"concrete_test":"Using the public nircam-fluxcal code, recompute the Table 6 / pmap 1490 PHOTMJSR values for all filters, detectors, and modes under three variants: (a) include LDS 749B and the three WDFS stars with equal weight, (b) only the four hot DA white dwarfs (G191-B2B, GD71, GD153, WD1057), (c) only A dwarfs and solar analogs. Compare each variant to the delivered values. If any PHOTMJSR shifts by more than its PHOTMJSR_ERR in a filter (especially F277W-F480M, where WDFS2317 is ~5% low), the calibration scale depends on the post-hoc exclusion and the <2% accuracy claim is not robust. If all shifts are <0.5%, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1.1 removes LDS 749B and the three WDFS white dwarfs from the averages because they deviate by 2-5% from the other 15 stars, and Section 2.3.5 states the final factors are combined after excluding stars outside 2.5 sigma. The delivered PHOTMJSR scale therefore rests entirely on the CALSPEC2 models of the 15 retained stars. The quoted scatter (<2%, and <1% for about half the imaging modes) measures agreement among those 15 after the exclusions; it cannot detect a common-mode error in the models or in the shared photometric chain (STPSF aperture corrections in Section 2.3.2, flat fields, subarray offsets in Section 2.3.3). The LMC/47 Tuc check (Section 3.1.2) only measures detector-to-detector offsets after calibration, and the comparison to previous deliveries (Section 4.2) measures stability, not absolute accuracy. The four excluded stars are the only faint, full-frame standards in the program; if WDFS2317's ~5% long-wavelength discrepancy (rather than its CALSPEC2 model) is caused by an aperture-correction or flat-field error that also affects the retained stars at lower amplitude, all long-wavelength calibration factors are biased. NIRISS independently found LDS 749B discrepant, which supports the model interpretation for that star, but no independent check exists for WDFS2317. Thus the central claim of an absolute scale with <2% accuracy is not independently verified; the paper's internal-consistency statements are being interpreted as absolute accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the updated absolute flux calibration of the JWST Near-Infrared Camera (NIRCam) imaging, time-series imaging, and coronagraphic modes. The calibration factor PHOTMJSR, converting DN/s/pixel to MJy/sr, is derived by comparing aperture photometry of 19 standard stars (A dwarfs, solar analogs, hot stars, and faint white dwarfs) with predicted fluxes from CALSPEC2 synthetic spectra, using STPSF aperture corrections and measured subarray-to-full-frame offsets. The authors deliver factors for all 29 imaging filters, 5 coronagraph masks, both weak lenses, and the coronagraph target-acquisition subarrays, with separate values by detector and subarray; the factors were installed in CRDS pmap 1490 in 2026 March. They report calibration scatter <2% for most imaging modes (<1% for about half), residual detector offsets typically <1% but up to 5%, subarray-dependent corrections up to ~1%, and slow responsivity declines <0.4%/yr. The paper provides the first on-sky calibration of WLP4/WLP8 and of dual-channel coronagraph configurations, and includes reproducibility checks against LMC and 47 Tuc fields.","tokens_in":36755,"tokens_out":6688,"duration_ms":72722,"significance":"If the delivered factors are accurate, this paper defines the reference flux scale for essentially all NIRCam imaging science from 2026 March onward, including time-series and coronagraphy. The analysis is a direct measurement (Eq. 1) rather than a fit to a desired outcome, uses 3.5 years and 19 stars, and ships the reduction code and an electronic table of factors, which is exemplary for an operational calibration paper. The independent LMC/47 Tuc comparison and the repeatability monitoring are valuable validation steps. The main limitation is that the quoted scatter measures internal agreement among stars retained after sigma-clipping, so it does not bound common-mode systematics in the CALSPEC2 model SEDs or in shared photometric chain elements; this is normal for JWST calibration, but it means the headline 'absolute accuracy' should be presented with an explicit model-error term, and the several single-star modes warrant prominent qualification.","major_comments":[{"comment":"The delivered PHOTMJSR_ERR and the paper's <2% scatter are internal statistical quantities, not absolute accuracy bounds. Four stars (LDS 749B and the three WDFS white dwarfs) are excluded from the averages because they deviate from the mean; this is a post-hoc selection, and while the NIRISS result supports the model interpretation for LDS 749B, no independent check is cited for WDFS2317, whose ~5% long-wavelength deviation could signal a flat-field or aperture-correction error common to the photometric chain. Please add to Section 3.1.1 and Table 6 an explicit statement and, where possible, a numeric systematic error floor representing CALSPEC2 model uncertainty, so that users do not mistake the reported scatter for total uncertainty.","section":"§2.3.5, §3.1.1"},{"comment":"The final WLP4/WLP8 calibration factors for Module A rest on a single star (P330E) after the SUB320 data are excluded, and the WLP8 Module B factors add only J1743045 and G 191-B2B. With N_stars=1, PHOTMJSR_STD and PHOTMJSR_SEM are zero by construction, so Table 6 gives no uncertainty for a mode the abstract lists as 'calibrated.' The text should state that these factors are provisional and carry a systematic uncertainty that cannot be estimated from this dataset, and should flag N_stars=1 entries in the electronic table.","section":"§3.2.1, Table 6"},{"comment":"Dual-channel coronagraphy is calibrated in far fewer configurations than the abstract suggests: the secondary-channel combinations are based on P330E alone (plus J1743045 for 210R), the LWB+SW filter combinations are not measured and are delivered using SWB factors substituted, and for FULL-frame coronagraphy the pipeline applies one factor averaged over the round or bar masks (up to ~3% deviation). These substitutions and averages are recorded in the text, but they do not appear in Table 6 in a way that users can identify. Please add a column or flag that marks substituted/averaged entries and include the substitution uncertainty in PHOTMJSR_ERR.","section":"§3.3, §4.1.1"},{"comment":"Detector-to-detector offsets are measured in only 8 filters and applied to those filters only, while offsets up to 0.055 mag (5.5%) are found for NRCA3 in F070W. For the remaining 21 imaging filters the delivered single per-detector factors contain unknown inter-detector errors; the text itself speculates that NRCA3 and NRCB4 may show the largest offsets in every filter. Please either derive correction factors for all filters or propagate an inter-detector uncertainty floor for the unmeasured filters in the delivered reference file.","section":"§3.1.2, Table 5"},{"comment":"Residual subarray offsets on NRCB1, especially between SUB160P and SUB160, reach ~3%, and because the final factors are averaged with SUB160 dominating, the paper warns that Time Series Imaging with standard filters on NRCB1 may be discrepant by up to ~2%. This residual is not included in the delivered error budget for those modes. Since Imaging Time Series is one of the modes being calibrated, please quantify the effect on the NRCB1 standard-imaging PHOTMJSR values and add it to the stated uncertainties, or correct for the residual before delivery.","section":"§3.4.1"}],"minor_comments":[{"comment":"In the Cycle 1 weak lens imaging row, the date '2002 Aug 20' should read '2022 Aug 20'.","section":"Table A2"},{"comment":"The note under Table 3 contains a typo: 'Ths inverse' should be 'The inverse'.","section":"Table 3"},{"comment":"Equation (1) defines N_ap as 'the flux density measured in a finite aperture (DN s^-1 pix^-1)'; this is a count rate, not a flux density, so please adjust the wording.","section":"§2.3.5"},{"comment":"The caption contains an erroneous line break in 'subar- rays'; please repair the hyphenation.","section":"Figure 5"},{"comment":"The equation for the Vega-Sirius zeropoint should state explicitly whether C is the full-frame or subarray-corrected calibration factor; clarifying this will avoid ambiguity for users applying Eq. (3).","section":"§4.3"},{"comment":"The limitation that the pipeline cannot distinguish coronagraph masks in FULL-frame observations is important for users; please add a prominent note in the abstract or Section 1 that mask-specific FULL-frame values must be applied manually.","section":"§4.1.2"}],"recommendation":"major_revision","confidential_remarks":"This is an operational calibration delivery of high value to the JWST community. The main issue is the gap between the abstract's unconditional claims (all modes calibrated, <2% scatter) and the body's many single-star and substituted entries, which the authors themselves document honestly. The revision should align the headline claims with the caveats and add explicit uncertainty floors for model systematics and unmeasured modes. This is fixable within the manuscript's scope; I do not see a basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Ben — here's my read of the NIRCam flux calibration paper (2608.08717). The genuinely new stuff is real: first on-sky weak-lens calibrations, first dual-channel coronagraphy, subarray-dependent factors, and detector offsets from overlapping LMC fields. The sample is 19 stars over 3.5 years, and the delivered table covers every imaging filter, mask, and weak-lens configuration. The authors are transparent about what is measured and what is substituted. For modes with three or more stars, the quoted scatter <2% holds up; about half the imaging modes are <1%. That is a solid piece of calibration work.\n\nThe soft spots are exactly where the stress-test note lands, but slightly less severe than the note implies. The four excluded stars (LDS 749B and three WDFS white dwarfs) are removed after the fact, and the retained 15 CALSPEC2 models carry the entire absolute scale. The quoted scatter measures agreement among those 15, not accuracy against an independent absolute reference. The LMC/47 Tuc check only measures detector-to-detector residuals after the fact, and the comparison to pmap 1126 measures stability, not ground truth. So the abstract's claim of 'scatter typically <2%' is precise; it would be wrong to read it as <2% absolute accuracy. The paper itself is careful with the word 'scatter,' though the summary bullet could mislead a casual reader.\n\nTwo minor things. First, several delivered modes rest on a single star: WLP4/WLP8 on Module A is P330E only, and most secondary coronagraphy channels are P330E with one exception. Second, the LWB+SW substitution and the ground-based WLP8+F150W2 are labeled clearly in Section 4.1.1, but those factors are placeholders, not measurements. Neither flaw is fatal; the paper says exactly what was measured and what was not.\n\nBottom line: this is the calibration paper that thousands of NIRCam papers will cite. The methodology is reproducible — code on GitHub, CRDS delivery, detailed appendix tables. The internal-consistency claims are well supported; the absolute-scale dependence on CALSPEC2 is openly stated. I'd send it to a serious referee, mainly to make sure the caveats stay prominent in the abstract and that Table 6 is machine-readable and complete. Yes, I'd cite it for the PHOTMJSR factors, and I'd bring it to reading group.","headline":"A thorough, honest calibration delivery whose <2% scatter is real but measures internal consistency, not absolute accuracy; the single-star and substituted modes are clearly labeled and the paper deserves a serious referee.","tokens_in":37323,"tokens_out":1984,"would_cite":true,"duration_ms":22205,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper delivers the updated absolute flux-calibration factors for every NIRCam imaging mode, converting DN/s to MJy/sr with scatter typically below 2%, and shows the instrument response drifts by less than 0.4% per year.","keywords":["flux calibration","JWST","NIRCam","PHOTMJSR","CALSPEC2","coronagraphy","weak lenses","subarray offsets"],"falsifier":"Obtain a high-signal NIRSpec spectrum of LDS 749B (or one of the excluded WDFS white dwarfs) and compare it directly to its CALSPEC2 model; if the model is the problem, the spectral mismatch will match the 2-5% photometric offset, while if the spectrum matches the model, the offset must originate in the NIRCam photometric chain. Alternatively, recompute the delivered factors using an independent stellar-atmosphere model grid and check whether the fifteen retained stars still agree within 1%.","tokens_in":36234,"feed_emoji":"🔭","tokens_out":9001,"duration_ms":80318,"temperature":0.7,"pith_summary":"The paper establishes the new absolute flux scale for JWST's NIRCam: a set of PHOTMJSR factors that convert measured counts (DN/s) into physical surface brightness (MJy/sr) for every imaging filter, coronagraphic mask, weak lens, and subarray. The factors come from 3.5 years of observations of 19 standard stars across three stellar types, and for most configurations the scatter is below 2%, with about half of the combinations below 1%. Because the factors were loaded into the JWST pipeline in March 2026, all NIRCam imaging science processed after that date inherits this scale. The paper also quantifies detector stability, finding count-rate changes under 0.4% per year that are still within the calibration uncertainties. A careful reader would care because this is the reference scale against which every NIRCam imaging measurement will be interpreted.","feed_headline":"New NIRCam calibration puts all imaging modes on a 2% flux scale","feed_subtitle":"The new PHOTMJSR values, in the pipeline since March 2026, are the reference scale for all NIRCam imaging science.","key_machinery":"The load-bearing object is the calibration factor C defined by C = Fν / (Nap Acor Scor Ωpix), the ratio of the stellar model flux density Fν to the measured count rate Nap in a finite aperture, corrected for aperture losses Acor, subarray offsets Scor, and the average pixel solid angle Ωpix. Measuring C for each filter, detector, mask, and subarray, then averaging across standard stars, is what turns every NIRCam image into physical surface-brightness units. The CALSPEC2 stellar models supply the flux reference, STPSF supplies the aperture corrections, and the full-frame-versus-subarray comparison supplies the subarray offsets; together these convert the raw DN/s measurements into the PHOTMJSR keyword values that the pipeline applies.","core_discovery":"The central discovery is a measured PHOTMJSR calibration factor for each of the 29 NIRCam imaging filters on all ten detectors, all five coronagraphic masks (including dual-channel pairings), both weak lenses (WLP4 and WLP8) used in time-series modes, and every science subarray, delivered as pmap 1490 in March 2026. The factors are derived by comparing CALSPEC2 model spectra to aperture photometry, corrected for aperture losses with STPSF simulations, for subarray-to-full-frame count-rate differences, and for detector-to-detector offsets measured from LMC and 47 Tuc mosaics. The scatter in the factor is below 2% in most filter+detector[+mask] combinations and below 1% in about half; no trends appear with count rate or well depth. Four white dwarfs (LDS 749B and three WDFS stars) are excluded from the averages because their model predictions disagree at the 2-5% level, and the paper attributes these disagreements to the CALSPEC2 models rather than to the photometric chain. This is the first on-sky calibration for the weak lenses and for secondary coronagraphy channels, and the delivered factors change by less than 4% from the previous delivery for most combinations.","pith_inferences":["If the delivered factors are right at the percent level, the 2-5% discrepancies seen for LDS 749B and WDFS2317 are model-atmosphere failures rather than instrument problems; an independent spectrum of those stars would settle which side is wrong.","The detector-offset pattern, with NRCA3 and NRCB4 showing the largest offsets in the four measured SW filters, suggests those two detectors may be the main source of residual scatter in the unmeasured filters; a 47 Tuc observation in an unmeasured filter would test this directly.","Because the faint white dwarfs were added specifically for full-frame calibration, a future model-grid update could convert those excluded stars into independent cross-checks without new observations, simply by recomputing the factors with corrected models.","The weak-lens flat fields are still ground-based and known to differ from sky flats by 6-8%; once sky-based weak-lens flats are delivered, the WLP4/WLP8 factors may shift by more than their current scatter, so those factors should be re-derived before relying on them for time-series science."],"forward_implications":["All NIRCam imaging, time-series imaging, and coronagraphic data processed with pmap 1490 share one absolute flux scale; calibrated images carry the PHOTMJSR keyword that converts DN/s to MJy/sr, and the delivered table gives the associated zeropoints.","Subarray-dependent factors remove count-rate offsets up to about 1%, so observations taken in different subarrays or with different readout patterns can be compared directly.","Weak-lens and dual-channel coronagraph modes now have an on-sky calibration for the first time, with scatter below 1.5% for setups measured with at least three stars.","NIRCam's response drift (under 0.4% per year in the short-wavelength channel and under 0.1% in the long-wavelength channel) is smaller than the calibration uncertainty, so no time-dependent correction is delivered yet, but future deliveries will add one as more epochs accumulate.","Detector-to-detector offsets measured in eight filters are folded into the reference files; residual offsets from LMC and 47 Tuc images are typically below 1% but reach 4-5% in a few cases, and a Cycle 5 program will measure offsets in all remaining filters."],"supporting_citations":[{"why":"Defines the absolute flux calibration program design, target selection, and the equation used to compute C.","marker":"K. D. Gordon et al. (2022)"},{"why":"Supplies the CALSPEC2 stellar model spectra that serve as the flux reference for every calibration factor.","marker":"R. C. Bohlin et al. 2014, 2022"},{"why":"Provides the STPSF simulations used to compute all aperture corrections.","marker":"M. D. Perrin et al. (2014)"},{"why":"Supplies the faint white dwarf standards used for full-frame NIRCam observations.","marker":"T. Axelrod et al. (2023)"},{"why":"Documents the 6-8% ground-versus-sky weak-lens flat-field variations that justify excluding off-reference-point weak-lens data.","marker":"B. Sunnquist et al. (2022)"},{"why":"Independent NIRISS analysis that also finds LDS 749B discrepant, supporting the decision to exclude it from the NIRCam averages.","marker":"K. Volk & P. Goudfrooij (2025)"},{"why":"Defines the Vega-Sirius zeropoint convention used for the delivered magnitude zeropoints.","marker":"G. H. Rieke et al. (2022)"},{"why":"Supplies the Photutils routines used for centroiding and aperture photometry.","marker":"L. Bradley et al. (2023)"}],"fun_headline_variants":["NIRCam's full imaging suite now calibrated to better than 2%","First on-sky calibration for JWST weak lenses and cross-pair coronagraphy","NIRCam flux scale updated: every mode, every filter, calibrated to <2%","JWST NIRCam calibration now covers all imaging modes with <2% scatter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire delivered scale inherits the accuracy of the CALSPEC2 model spectra for the fifteen stars kept in the averages; if those models carry a common bias, every PHOTMJSR value is biased by the same amount and the sub-2% scatter measures internal consistency rather than absolute accuracy.","fun_headline_variants_meta":{"raw":{"variants":["NIRCam's full imaging suite now calibrated to better than 2%","First on-sky calibration for JWST weak lenses and cross-pair coronagraphy","NIRCam flux scale updated: every mode, every filter, calibrated to <2%","JWST NIRCam calibration now covers all imaging modes with <2% scatter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3624,"prompt_tokens":1099,"completion_tokens":2525,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":715,"completion_tokens_details":{"reasoning_tokens":2436}},"tokens_in":715,"tokens_out":2525,"duration_ms":20301,"temperature":1.0,"reasoning_tokens":2436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:25:39.426073+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain a high-signal NIRSpec spectrum of LDS 749B (or one of the excluded WDFS white dwarfs) and compare it directly to its CALSPEC2 model; if the model is the problem, the spectral mismatch will match the 2-5% photometric offset, while if the spectrum matches the model, the offset must originate in the NIRCam photometric chain. Alternatively, recompute the delivered factors using an independent stellar-atmosphere model grid and check whether the fifteen retained stars still agree within 1%.","supporting_citations":[{"cited_title":"2022, title NIRCam Commissioning Results NRC-10-Flat Fields, Scattered Light, and Backgrounds , , Tech","cited_arxiv_id":null,"evidence_quote":"Documents the 6-8% ground-versus-sky weak-lens flat-field variations that justify excluding off-reference-point weak-lens data."}],"review_version":1}