{"id":"b561e12b-3e1d-4f55-a1ef-3a2d5dedd5af","arxiv_id":"2608.10017","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 14-night, 101-star campaign on one field shows that STDWeb formal errors are accurate within a night, and the multi-night excess scatter is almost entirely removed by applying the pipeline's per-epoch colour term.","lead":"STDWeb's reported magnitude errors match the true scatter within a single night, but across nights they understate the real uncertainty by roughly half unless the pipeline's time-varying colour term is applied. This empirical error budget, from 14 nights on one field, gives pro-am networks a concrete model for when formal errors can be trusted.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing question is whether c_t is a physical transformation or an in-sample nuisance fitted to the same ensemble; a split-star or second-pipeline test would settle it.","rationale":"The reader's weakest assumption identifies c_t correctness and possible self-fulfilment, and I agree that this is the pivotal point. However, the reader frames it largely as a suspicion that the correction may be 'partially self-fulfilling.' I have sharpened this into a concrete, testable concern: because c_t is fitted on the same ensemble and frames used for validation, the residual collapse is guaranteed for any per-epoch linear colour trend, physical or not. The split-star test directly distinguishes the two cases using data already in hand, without external catalogues or new observations. I do not see a stronger objection: the variance budget closes, the Gaia-colour validation is genuinely independent of the colour-term fit in the sense that the light-curve-fitted colours reproduce Gaia, and the caveats in Sect. 5.5 are honest about the single-field, single-season limitation. The remaining 3.4 mmag night term is acknowledged and does not undermine the central diagnosis. The paper's main weaknesses, as the reader notes, are the lack of public data/code and the reliance on a private communication for Eq. (1); these are reproducibility issues rather than internal inconsistencies. The conditional verdict is therefore appropriate, and my concern does not change it, but the out-of-sample test proposed here would materially raise confidence in the 'entire excess' claim.","tokens_in":7500,"tokens_out":5628,"duration_ms":67902,"concrete_test":"Split the 101 ensemble stars into two disjoint sets (e.g., 50 and 51 stars). For every epoch, re-fit the colour-term coefficient c_t using only set A, then compute m_sys for set B using these held-out coefficients. Re-evaluate σ_night, σ_camp, and χ_camp for set B alone. If χ_camp(B) remains ~1.0-1.1 and σ_night stays ≈3.5 mmag, the colour-term explanation is not an artifact of fitting and validating on the same stars. If χ_camp(B) rises significantly, c_t is partly absorbing ensemble-specific noise, and the 'entire excess' claim must be weakened. A complementary check would be to reduce a subset of frames with an independent pipeline (e.g., Muphoten or a custom forced-aperture reduction with its own per-frame colour solution) and compare the resulting corrected budget.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the epoch-dependent colour term c_t, as emitted by STDWeb, correctly describes the instrumental-to-catalogue transformation, so that m_sys = mag_calib + c_t (BP-RP) removes the true chromatic systematics and the 'entire excess' disappears (Eq. 1, Tables 1-2). The weakest link is that c_t is fitted per frame using the same 101 ensemble stars and the same frames on which the residuals are then evaluated. Any per-frame linear dependence of the photometric solution on (BP-RP) is therefore removed from those residuals by construction, whether its origin is a real atmospheric/instrumental colour term or a flexible nuisance absorbing colour-dependent flat-field, illumination, or PSF systematics. The paper's internal evidence (χ_camp = 1.0-1.1, vanishing correlations, data-fitted colours matching Gaia to 0.027 mag) strongly supports a physical interpretation, but it does not by itself rule out an in-sample component, because the validation and the calibration share the same stars and epochs. A fully convincing demonstration of 'entire excess' requires out-of-sample confirmation that c_t generalizes to stars and reductions not used to fit it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents an empirical multi-night photometric error budget for the STDWeb pipeline, based on 157 G-band exposures of a single field over 14 nights, analysed through 101 constant field stars. The authors find that within-night single-epoch scatter matches the reported formal errors (chi_within ~ 1), while the naive campaign-level scatter is 11-12 mmag with chi_camp ~ 1.5-1.7. They attribute the entire excess to an epoch-dependent colour term: computing m_sys = mag_calib + c_t (BP-RP) with the pipeline's per-frame colour-term coefficient reduces the night term to 3.4 mmag, restores chi_camp ~ 1.0-1.1, and improves Gaia-anchored per-star repeatability from 43 to 9.5 mmag. Colours fitted from the light curves themselves reproduce Gaia BP-RP to 0.027 mag and yield the same budget. The paper includes a closed variance decomposition, practical recommendations, and a cautionary note on re-deriving zero points outside the pipeline.","tokens_in":7780,"tokens_out":9567,"duration_ms":91342,"significance":"The paper addresses a real need for pro-am networks: a validated, closed error budget for a widely used web pipeline. If the central claim holds, the practical message is strong and directly actionable: formal errors are trustworthy within a night, and applying the provided colour-term correction makes them trustworthy across nights. The variance budget closes internally, and the data-fitted colours are validated against Gaia (r=0.98, 0.027 mag). The paper is honest about its caveats (single field, single instrument, unbalanced exposure mix) and includes a cautionary tale about re-deriving zero points. However, the central 'entire excess' claim rests on a correction that is fitted to the same ensemble on which it is evaluated, so the significance of the physical interpretation is conditional on an out-of-sample demonstration.","major_comments":[{"comment":"The decisive test is in-sample: STDWeb's per-frame colour-term coefficient c_t is fitted on the same 101 ensemble stars and the same frames whose residuals are then used to compute the night term and chi_camp. Because the per-frame photometric solution includes a colour term, the least-squares fit forces the ensemble residuals in m_sys to have zero mean and zero linear colour dependence within each frame. The observed collapse of sigma_night from 8.6 to 3.4 mmag and chi_camp from 1.69 to 1.03 is therefore partly a property of the fitting procedure. I request an out-of-sample test: split the 101 stars into two disjoint sets, fit c_t on one set and compute the budget on the other, or repeat the analysis with an independent reduction pipeline. Without such a test, the claim that the entire excess is the epoch-dependent colour term is not fully established.","section":"Sect. 2.3, Eq. (1), Sect. 4.3"},{"comment":"The data-driven colour fitting is partially circular. Fitting each star's colour as the slope of mag_calib versus c_t removes the linear c_t dependence from the residuals by construction, so the identical budget in the last column of Table 2 is not an independent confirmation of the correction. The Gaia-colour column is the meaningful external validation; the fitted-colour column demonstrates only that the method is self-contained. The text should clearly state that the Gaia-colour test is the one that validates the physical origin, and that the fitted-colour test is a practical alternative, not an independent check.","section":"Sect. 3.3, Table 2"},{"comment":"chi_camp is computed as the ratio of the empirical scatter on m_sys to sigma_form, the formal error reported on mag_calib. After applying Eq. (1), the uncertainty in m_sys should include contributions from the uncertainty in c_t and in (BP-RP). The paper does not propagate these terms or argue that they are negligible. Because the central claim is that formal errors are valid across nights on m_sys, the authors should either propagate these uncertainties into the formal error or justify that they are small relative to the 8.1 mmag median formal error.","section":"Sect. 2.3, Sect. 5.2, Table 2"}],"minor_comments":[{"comment":"The symbol c_k is used for the per-epoch ensemble offset while c_t denotes the colour-term coefficient; this is confusing. Suggest renaming the epoch offset to z_k or d_k.","section":"Sect. 3.2, Eq. (2)"},{"comment":"The field coordinates or a reference to a public chart would aid reproducibility.","section":"Sect. 2.2"},{"comment":"Please specify the bin width or smoothing scale of the running median of the formal error.","section":"Fig. 1"},{"comment":"The caption should state that the 16-84% range is the central-quantile range over stars and clarify how it is computed.","section":"Table 1"},{"comment":"The residual night term of ~3 mmag is attributed to 'second-order chromatic effects' among other causes; a simple quadratic colour-term test in the residuals would help quantify this.","section":"Sect. 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and likely useful for the pro-am community. The main technical issue is the in-sample nature of the colour-term validation; if the authors can provide a split-star out-of-sample test, the central claim would be much stronger. I do not see any fundamental flaw that would require rejection, but the 'entire excess' phrasing should be softened or supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know about this one: it's a clean empirical error budget for STDWeb, built from 157 exposures of one field over 14 nights. The headline result is that the pipeline's formal errors are accurate within a night (chi_within ~1) but the naive multi-night scatter is dominated by a time-varying color-term coefficient, so converting to the catalogue-system magnitude m_sys = mag_calib + c_t (BP-RP) restores chi_camp ~1. That is new for STDWeb and practically important for the pro-am community.\n\nWhat the paper does well: the variance budget closes (sqrt(6.8^2+8.6^2)=11.0 vs 11.3 mmag measured), the airmass and colour correlations vanish after correction, and the data-fitted colours match Gaia BP-RP to 0.027 mag with r=0.98. That Gaia validation is the strongest evidence that c_t is physical rather than an in-sample nuisance. The paper also gives a clear practical warning against re-deriving zero points outside the pipeline, and documents its own caveats honestly (one field, one instrument, one season, unbalanced exposures).\n\nThe soft spots are real but not fatal. The key correction formula rests on a private communication (ref [3]), which is a reproducibility problem; no public data or code means the central numbers can't be independently checked. The claim that the 'entire excess' is explained is too strong: a 3.4 mmag night term remains after correction, and the analysis covers only one field/season. The stress-test worry about c_t being fitted from the same ensemble stars and frames is legitimate in principle—any per-frame linear colour dependence is removed from those residuals by construction. But the external Gaia cross-check and the vanishing correlations make the in-sample interpretation unlikely to be the whole story. I'd want a split-star or second-pipeline test before calling it fully closed, but the diagnosis is solid enough for practical use.\n\nWho is this for: anyone using STDWeb for campaign photometry, or building error models for pro-am networks. It deserves a serious referee, but the author should be pushed to release source tables or a reproducible reduction, provide a public derivation of Eq. (1), and soften 'entire'. I'd take it as a conditional accept.\n\nRecommendation: send it to review, with those requests.","headline":"A genuinely useful, carefully quantified error budget for STDWeb whose central color-term diagnosis holds up, modulo the 'entire excess' overclaim and an unresolved in-sample fitting question.","tokens_in":8225,"tokens_out":1918,"would_cite":true,"duration_ms":16803,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A drifting colour term, not bad errors, explains STDWeb's multi-night scatter.","keywords":["photometric repeatability","colour term","time-series photometry","error budget","STDWeb","differential photometry","pro-am astronomy","reference-catalogue photometry"],"falsifier":"Redo the multi-night budget with the per-epoch $c_t$ values randomly permuted across epochs while keeping each star's colour fixed: the explanation predicts that the night term and the $\\chi_{\\rm camp}$ excess must return, because the colour-term correction has been destroyed. If the improved $\\chi_{\\rm camp}$ persists under permutation, the colour term is not the actual cause and the collapse seen with the true $c_t$ is a fitting artifact.","tokens_in":7222,"feed_emoji":"🔭","tokens_out":7505,"duration_ms":75017,"temperature":0.7,"pith_summary":"This paper measures how well STDWeb's reported photometry repeats on the same field night after night, using 157 exposures taken on 14 nights and the light curves of 101 constant field stars. It finds that the pipeline's formal errors are accurate within a single night, but that naive use of its calibrated magnitude across nights gives a true single-epoch scatter roughly 1.5 to 1.7 times the formal error. The paper shows that this entire excess is a time-varying colour term: transforming magnitudes into the catalogue system with $m_{\\rm sys} = {\\rm mag\\_calib} + c_t ({\\rm BP}-{\\rm RP})$ collapses the night-level offset to 3.4 mmag and restores $\\chi_{\\rm camp}\\simeq1$. The correction improves absolute per-star repeatability from 43 to 9.5 mmag, and it works without an external colour catalogue because colours fitted from the light curves themselves give the same budget.","feed_headline":"STDWeb's multi-night scatter is a drifting colour term","feed_subtitle":"Formal errors are accurate within a night; a drifting colour term explains the rest, cutting scatter from 43 to 9.5 mmag.","key_machinery":"The load-bearing object is the catalogue-system magnitude $m_{\\rm sys}={\\rm mag\\_calib}+c_t({\\rm BP}-{\\rm RP})$, where $c_t$ is the per-frame colour-term coefficient that STDWeb already fits as part of its photometric solution and ${\\rm BP}-{\\rm RP}$ is the star's colour. The paper's argument is that ${\\rm mag\\_calib}$ is expressed in a time-varying instrumental pseudo-band, so making the transformation to the catalogue system explicit removes the otherwise unexplained night-level offsets. The supporting machinery is a closed variance decomposition $\\sigma_{\\rm camp}^2\\simeq\\sigma_{\\rm within}^2+\\sigma_{\\rm night}^2$, which lets the paper isolate what the colour term contributes, plus a data-driven colour fit that allows the correction to run when no external colour catalogue is available.","core_discovery":"The paper's central claim is that the excess scatter of STDWeb magnitudes across nights—the factor by which the true single-epoch uncertainty exceeds the reported formal error—is the epoch-dependent colour term. STDWeb's calibrated magnitude ${\\rm mag\\_calib}$ lives in an instrumental pseudo-band whose relation to the catalogue system changes from epoch to epoch, as the fitted colour-term coefficient $c_t$ wanders between $-0.14$ and $-0.28$. Defining $m_{\\rm sys}={\\rm mag\\_calib}+c_t({\\rm BP}-{\\rm RP})$ collapses the night-level offset from 8.6 to 3.4 mmag for the bright half of the ensemble, restores $\\chi_{\\rm camp}=1.0$–$1.1$, and improves the absolute per-star repeatability anchored to the reference catalogue from 43 to 9.5 mmag, while correlations between residuals and airmass or colour vanish. Repeating the budget with stellar colours fitted from the light curves as the slope of ${\\rm mag\\_calib}$ against $c_t$ gives an indistinguishable result, so the correction needs no external colour catalogue.","pith_inferences":["Editorial inference: the same qualitative mechanism—a time-varying colour term inflating multi-night scatter—should appear in any pipeline that calibrates against a broadband survey catalogue with a fitted colour term, so the pattern is likely to generalise even if the exact numbers are station-specific.","Editorial inference: because the colour correction is self-contained and matches catalogue colours so closely, the method could be used to inter-calibrate heterogeneous cameras within a network, effectively building a common colour system from photometry alone.","Editorial inference: a direct testable extension would be to repeat this campaign with a filter that more closely matches the catalogue passband; the paper's mechanism predicts the night term should shrink roughly in proportion to the reduction in colour-term variance.","Editorial inference: the cautionary result on re-deriving zero points suggests that archived light curves from other pipelines may contain similar time-varying pseudo-band terms, and re-analysing them with a colour-term transformation could reduce systematic scatter in existing transient and variability surveys."],"forward_implications":["Within a single night, STDWeb's formal errors can be used as-is: the measured within-night scatter matches the reported formal error ($\\chi_{\\rm within}\\simeq1$).","Across nights, users should analyse $m_{\\rm sys}={\\rm mag\\_calib}+c_t({\\rm BP}-{\\rm RP})$ rather than the raw pipeline magnitude; doing so restores $\\chi_{\\rm camp}=1.0$–$1.1$ and makes formal errors usable across nights.","Targets with no colour information and too few epochs to fit one should have their formal errors inflated by a factor of roughly 1.5–2, or an 8 mmag night term added in quadrature.","A low-cost colour-CMOS system can deliver absolute photometry that repeats at 9.5 mmag per star across 14 nights, with a nightly zero point stable to 2.3 mmag, when the colour-term correction is applied.","Campaign-level meta-analyses of STDWeb products should consume ${\\rm mag\\_calib}$ and $c_t$ directly; re-deriving zero points outside the pipeline is an avoidable failure mode that produced spurious excursions of up to 0.7 mag in this campaign.","The error budget is self-contained: stellar colours fitted from the light curves reproduce catalogue ${\\rm BP}-{\\rm RP}$ colours to 0.027 mag and produce an identical budget, so the correction works without an external colour catalogue."],"supporting_citations":[{"why":"Supplies the STDWeb pipeline whose formal errors and per-frame colour-term coefficients are the objects under study.","marker":"[1]"},{"why":"Provides the underlying reduction library that implements the photometric solution, astrometric calibration and source extraction behind STDWeb.","marker":"[2]"},{"why":"Introduces the key idea that the catalogue-system magnitude, not the raw calibrated magnitude, is the quantity to analyse, and suggests the data-driven colour-fitting method.","marker":"[3]"},{"why":"Supplies the reference-catalogue colours (${\\rm BP}-{\\rm RP}$) and absolute $G$ magnitudes used in the correction and in the anchored repeatability metric.","marker":"[4]"},{"why":"Provides the SExtractor source extraction step that produces the per-frame detections and measurements used throughout the STDWeb reduction.","marker":"[13]"}],"fun_headline_variants":["Drifting colour term explains STDWeb's night-to-night scatter","STDWeb's extra scatter traced to epoch-dependent colour term","Self-calibrated colour term cuts STDWeb scatter 4.5x","STDWeb's true error 1.5-1.7x formal: colour term fix","Colour term drift, not noise, inflates STDWeb's multi-night errors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands or falls on the assumption that STDWeb's per-frame colour-term coefficient is a faithful description of the real, time-varying transformation between the instrumental band and the catalogue system, so that a linear colour correction removes the true chromatic systematics and does not merely re-fit the scatter.","fun_headline_variants_meta":{"raw":{"variants":["Drifting colour term explains STDWeb's night-to-night scatter","STDWeb's extra scatter traced to epoch-dependent colour term","Self-calibrated colour term cuts STDWeb scatter 4.5x","STDWeb's true error 1.5-1.7x formal: colour term fix","Colour term drift, not noise, inflates STDWeb's multi-night errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000894,"raw_usage":{"total_tokens":3990,"prompt_tokens":1220,"completion_tokens":2770,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":836,"completion_tokens_details":{"reasoning_tokens":2669}},"tokens_in":836,"tokens_out":2770,"duration_ms":19967,"temperature":1.0,"reasoning_tokens":2669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:28:15.244761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Redo the multi-night budget with the per-epoch $c_t$ values randomly permuted across epochs while keeping each star's colour fixed: the explanation predicts that the night term and the $\\chi_{\\rm camp}$ excess must return, because the colour-term correction has been destroyed. If the improved $\\chi_{\\rm camp}$ persists under permutation, the colour term is not the actual cause and the collapse seen with the true $c_t$ is a fitting artifact.","supporting_citations":[{"cited_title":"2021,STDPipe: Simple Transient Detection Pipeline, Astrophysics Source Code Li- brary, ascl:2112.006","cited_arxiv_id":null,"evidence_quote":"Provides the underlying reduction library that implements the photometric solution, astrometric calibration and source extraction behind STDWeb."},{"cited_title":"2026, private communication","cited_arxiv_id":null,"evidence_quote":"Introduces the key idea that the catalogue-system magnitude, not the raw calibrated magnitude, is the quantity to analyse, and suggests the data-driven colour-fitting method."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reference-catalogue colours (${\\rm BP}-{\\rm RP}$) and absolute $G$ magnitudes used in the correction and in the anchored repeatability metric."},{"cited_title":"1996,SExtractor: Soft- ware for source extraction, A&AS, 117, 393","cited_arxiv_id":null,"evidence_quote":"Provides the SExtractor source extraction step that produces the per-frame detections and measurements used throughout the STDWeb reduction."}],"review_version":1}