{"id":"55d1b9a9-4a37-4a13-885a-8d9eb3fee62a","arxiv_id":"2505.00553","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A leave-one-source-out Jackknife cross-validation distinguishes accurate from overfitted strong lens mass models in cases where the reduced chi-square cannot.","lead":"The authors propose a new way to check whether a computer model of a galaxy cluster's mass distribution is trustworthy: remove one background galaxy's images from the model fitting, then try to predict where those images should be. This 'Jackknife' test can flag overfitted lens models that look good in the usual chi-square test but predict poorly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Sect. 2.2 claim that correct models give N(0,1) Jackknife residuals is un-derived and setup-dependent (the paper's own 5-source correct model gives 1.24), so the 'much larger than 1' overfitting flag has no calibrated null threshold.","rationale":"The reader's weakest assumption identifies exactly the load-bearing point: the N(0,1) reference distribution is asserted, not derived, and the paper's own correct-model simulation already deviates from it. My stress-test agrees and sharpens it: the correct-model width is setup-dependent (1.24 vs 1.02 for 5 vs 20 sources), which is the expected signature of leverage in leave-one-out prediction, and without error bars on the quoted widths the headline 1.24 vs 2.28 separation is not quantitatively established. This is a correctness risk, not a question of external consensus; it concerns internal support for the method's threshold. I credit the paper for a clear mock experiment, multiple source-number configurations, and an honest statement in Sect. 5.1 that the calibration trend needs more analysis. Those features make the paper a useful proposal, but they do not supply the missing null distribution. The verdict should remain CONDITIONAL, unchanged from the reader: accept pending error bars plus either a derivation or a calibration table of Jackknife widths for correct models.","tokens_in":16869,"tokens_out":7824,"duration_ms":87829,"concrete_test":"Re-run the 5-source and 20-source mock experiments and report the per-realization Jackknife standard deviations for each of the 100 noise realizations (or the standard error of the pooled width), then construct 16-84 percentile ranges for correct and incorrect models. If the correct-model range overlaps the incorrect-model range, the Sect. 3.2 separation is not statistically established. In parallel, derive the expected leave-one-out residual variance for the linearized lens model, Var(Delta x / sigma) = 1 + h_ii, and check whether the 5-source correct-model value 1.24 is consistent with the leverage h_ii for that design; if the null width is not 1, produce a calibration table of null widths versus number of sources, images per source, and model parameters.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Sect. 3.2 is that the Jackknife width separates a correct model (standard deviation 1.24) from an incorrect, overfitted model (2.28) even when both have reduced chi-square near 1. This interpretation rests on the Sect. 2.2 assertion that for a correct model, Delta x/sigma and Delta y/sigma follow N(0,1). That assertion is not derived, and it is not what leave-one-out residuals are expected to do: for a nonlinear model with estimated parameters, the predictive variance is sigma^2 times (1 + leverage), where leverage depends on the number of constraints, the number of images per held-out source, and parameter degeneracies. The paper's own numbers illustrate the problem: the correct-model width is 1.24 with 5 sources but 1.02 with 20 sources, so the null value is evidently setup-dependent. No uncertainties are quoted on these widths or on the 2.28 vs 1.14 values, so the separation significance is unquantified. The MACS0647 demonstration (HST 4.03 vs JWST 1.72) is presented as correct vs overfitted, but a width of 1.72 could equally be a correct model with limited constraints. Section 5.1 and Fig. 6-7 show the proposed calibration relation is only positive at about 2 sigma with large scatter, and the text itself states more analysis is needed. Until a derivation or calibration table provides the null distribution of Jackknife widths, the method cannot cleanly separate 'correct but few constraints' from 'incorrect and overfitted' in real clusters.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a jackknife validation method for cluster-scale strong lens mass models. The procedure removes all multiple images of one source, refits the model to the remaining sources, predicts the removed images, and compares observed and predicted positions through standardized residuals Delta_x/sigma, Delta_y/sigma. The central claim, stated in Section 3.2, is that this jackknife distribution separates a correct model (standard deviation 1.24) from an incorrect, overfitted model (2.28) in a 5-source mock experiment, even though both models have reduced chi-square near 1. The paper also applies the method to MACS0647 using HST and JWST data and explores whether the jackknife width can calibrate MCMC error estimates for physical quantities.","tokens_in":17207,"tokens_out":3222,"duration_ms":34757,"significance":"If the central claim holds, the jackknife method would provide a quantitative, hold-out-based diagnostic for overfitting in strong lens mass modeling, addressing a real problem: the reduced chi-square is often made acceptable by inflating positional errors. The controlled mock experiment is a genuine strength: it uses a known input model, genuine hold-out predictions, and a fair construction in which the incorrect model is rescaled to chi-square/DoF ~ 1. The demonstration with real observations and the discussion of MCMC error validation are useful first steps. However, the interpretation depends on an uncalibrated and setup-dependent null distribution, and the quoted widths lack uncertainty estimates, so the method as currently stated cannot cleanly classify models in practice.","major_comments":[{"comment":"The assertion that for a correct model the jackknife residuals Delta_x/sigma, Delta_y/sigma follow a Gaussian distribution with standard deviation 1 is not derived, and it is contradicted by the paper's own simulations: the correct model gives a jackknife width of 1.24 with 5 sources and 1.02 with 20 sources. Leave-one-out predictive residuals for models with estimated parameters have variance sigma^2 times (1 + leverage), where the leverage depends on the number of constraints, the number of images per held-out source, and parameter degeneracies. Without a derivation or a calibration table, the reference value 1 cannot serve as a universal null threshold, and the 'much larger than 1' criterion for overfitting is not quantitatively grounded. This is the main load-bearing issue for the paper's central claim.","section":"Section 2.2 (with Figs. 2 and 3)"},{"comment":"The quoted jackknife widths (1.24 vs 2.28 for 5 sources; 1.02 vs 1.14 for 20 sources) are presented without uncertainties or effective sample sizes. Residuals from multiple images of the same removed source are correlated, so the number of independent jackknife residuals is closer to the number of sources R (5 or 20) than to the number of images (15 or 60). The sampling uncertainty on these widths is therefore non-negligible, and the separation between 1.24 and 2.28 may be less significant than it appears. The paper should report confidence intervals, bootstrap errors, or the distribution of widths over the 100 mock realizations to support the claimed discrimination.","section":"Section 3.2 (Figs. 2 and 3)"},{"comment":"The MACS0647 demonstration interprets the JWST jackknife width of 1.72 as indicating a relatively accurate model and the HST width of 4.03 as indicating overfitting. Given the uncalibrated null distribution, a width of 1.72 is also consistent with the 1.24 width of the correct 5-source mock model, so the observation alone cannot distinguish 'correct but limited constraints' from 'overfitted'. The application should be framed strictly as a feasibility demonstration, or it must be accompanied by a calibrated threshold and confidence intervals on the widths.","section":"Section 4 (Fig. 4)"}],"minor_comments":[{"comment":"The calibration relation between sigma_Jackknife/sigma and sigma_realization/sigma_MCMC is described as positive only at about the 2-sigma level with large scatter, and the text states more analysis is needed. This section should be explicitly labeled as exploratory; the wording 'we can validate the statistical error' overstates what is currently established.","section":"Section 5.1 (Figs. 6 and 7)"},{"comment":"The caption notes that source-plane fitting can produce a different number of predicted images and that the incorrect model predicts extra images near halo centers. The paper should clarify how observed and predicted images are matched when the predicted multiplicity differs from the observed multiplicity, since this affects the computed Delta_x, Delta_y residuals.","section":"Figure 2 caption"},{"comment":"The table uses 'DoF' as both a column name and a concept; for readability, consider defining 'DoF = Constraint - Parameter' explicitly in the table caption, as is done in the text.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the journal's scope and the controlled mock experiment is a solid foundation. The main obstacle is the uncalibrated null distribution for the jackknife width; a derivation or a calibration study would be needed before the method can claim to distinguish correct from incorrect models in real clusters. I recommend major revision rather than rejection because the gap appears fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nWhat you should know: this paper proposes a source-level jackknife (leave-one-source-out) as a validation test for cluster-scale strong lens mass models, where reduced chi-square is known to be uninformative because positional errors are inflated to absorb substructure. The basic idea is right, and the 5-source mock experiment is a clean demonstration that a deliberately overfitted model can have chi-square near 1 while its hold-out predictions are systematically off. But the paper's key claim—that a correct model gives N(0,1) jackknife residuals—is asserted, not derived, and the paper's own numbers show the null width is setup-dependent: 1.24 with 5 sources, 1.02 with 20. That undermines the \"much larger than 1\" flag until a calibration table or derivation exists.\n\nWhat's new and good. Leave-one-source-out prediction is a natural extension of earlier hold-out work (Remolina González et al. 2018; Raney et al. 2021; Perera et al. 2025), and the paper cites those fairly. The controlled simulation is well posed: both models have reduced chi-square near 1, yet the jackknife widths separate (1.24 vs 2.28). The MACS0647 demonstration is appropriately presented as illustrative, not definitive. The authors are also honest in Section 5.1: the MCMC-error calibration slope is only positive at about 2 sigma and they say more work is needed.\n\nSoft spots. The N(0,1) reference is the load-bearing issue. For a nonlinear model with estimated parameters, leave-one-out residuals have variance sigma^2 times (1 + leverage), and leverage depends on the number of constraints, images per source, and degeneracies. So the null cannot be universal. No uncertainties are quoted for the widths, so we don't know whether 2.28 vs 1.24 is a 3-sigma separation or a noisy coincidence. The JWST width of 1.72 could just as easily be a correct model with limited constraints, not an overfitted model. The 20-source case shows almost no separation, consistent with the authors' story but limiting the method to clusters where overfitting is actually happening.\n\nWho it's for. Cluster lens modelers who want a data-driven check, and method people working on validation statistics. It's a legitimate first step, not a finished tool.\n\nI'd send it to peer review rather than desk reject. The referee should ask for a derivation or calibration of the null distribution, error bars on the widths, and a practical threshold.\n\nBest","headline":"A promising but uncalibrated jackknife validation test for cluster lens models; the core simulation works, but the N(0,1) null is setup-dependent and no error bars are quoted.","tokens_in":17780,"tokens_out":2850,"would_cite":false,"duration_ms":27607,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A leave-one-source-out Jackknife test separates correct from overfitted strong-lens mass models even when both have reduced chi-square near unity.","keywords":["strong gravitational lensing","galaxy clusters","mass model validation","jackknife method","overfitting","predictive power","MACS0647","chi-square"],"falsifier":"Run the same leave-one-source-out analysis on simulated clusters with a known true model but deliberately fewer constraints or strong parameter degeneracies, keeping chi-square per degree of freedom near 1; if the correct model's Jackknife standard deviation rises well above 1, the fixed N(0, 1) reference would mislabel correct models as overfitted. Equivalently, an analytic calculation of the expected leave-one-out residual variance for a linearized lens model with 2P constraints and Q parameters would settle whether the reference value should depend on the degrees of freedom.","tokens_in":16618,"feed_emoji":"🔭","tokens_out":6504,"duration_ms":62182,"temperature":0.7,"pith_summary":"Strong-lens mass models of galaxy clusters are usually judged by reduced chi-square, but the unknown positional scatter from dark-matter substructure makes that test unreliable. This paper proposes a hold-out test: remove all images of one source, refit the model on the remaining sources, then predict the removed source's image positions and compare residuals with the assumed positional error. In mock data where a correct model and an intentionally overfitted model both have chi-square per degree of freedom near 1, the Jackknife residuals of the correct model scatter with standard deviation 1.24, close to the expected 1, while the overfitted model's residuals scatter with standard deviation 2.28. The same test applied to the cluster MACS0647 flags the 31-image HST model as overfitted (standard deviation 4.03) and the 86-image JWST model as much better behaved (1.72). If the method holds up, it gives cluster lens modelers a quantitative check on predictive power and a possible calibration of MCMC error estimates.","feed_headline":"Overfit lens models fail a hold-out test chi-square can't catch","feed_subtitle":"A correct model's held-out residuals track the assumed error; an overfitted model's are twice as wide.","key_machinery":"The central object is the leave-one-source-out Jackknife residual distribution. For each source, the procedure removes all of that source's multiple images, refits the mass model using the remaining sources, predicts the removed source's image positions, and records the positional differences ($\\Delta$ x, $\\Delta$ y); these are divided by the assumed positional error $\\sigma$ to form standardized residuals. The diagnostic is the standard deviation of the pooled standardized residuals, compared with the reference value 1 from a Gaussian N(0, 1) distribution. The method works by exploiting the difference between precision and accuracy: an overfitted model reproduces its training images well but predicts held-out images poorly, so its Jackknife residuals are wider than the assumed error, while a genuinely predictive model produces residuals that track the assumed error.","core_discovery":"The paper's central claim is that the Jackknife method can reveal whether a strong-lens mass model is overfitted, even when the usual reduced chi-square diagnostic cannot. In the 5-source simulation, the incorrect model is deliberately given extra multipole perturbations and is fit with an inflated positional error so that its reduced chi-square is also about 1; nevertheless, the standard deviation of the standardized Jackknife residuals is 2.28 for the incorrect model versus 1.24 for the correct model. The paper interprets this as the incorrect model failing to predict held-out image positions by roughly twice the assumed positional error, while the correct model predicts them at about the assumed error. In the MACS0647 demonstration, the HST-based model gives a Jackknife standard deviation of 4.03, suggesting overfitting, whereas the JWST-based model gives 1.72, suggesting a more reliable model. The paper further argues that the ratio sigma_Jackknife/sigma correlates with the ratio of true to MCMC-estimated errors on physical quantities like magnifications and time delays, so the method might eventually correct underestimated statistical errors.","pith_inferences":["An implication the paper leaves implicit is that the expected Jackknife width for a correct model is probably not exactly 1 for all designs; it should depend on the number of constraints and parameter degeneracies, so a calibration grid covering correct models with varying source counts would allow model-specific thresholds rather than a fixed Gaussian reference.","The same leave-one-source-out logic could be applied directly to predicted magnifications or time delays, not just image positions, turning the method into a per-quantity predictive check that would speak more directly to the science outputs that matter.","A natural next test is to apply the Jackknife to mock clusters with simulated dark-matter substructure; if substructure inflates the Jackknife width of a true model, the method may need to distinguish genuine complexity from a wrong model.","Because the 20-source case shows no separation, the method's power is highest in the low-constraint regime where overfitting is most dangerous, which is also the regime where the reduced chi-square is least informative."],"forward_implications":["A model with reduced chi-square near 1 can still be overfitted, so the Jackknife test provides a complementary diagnostic that does not require knowing the true substructure error.","The 5-source mock result quantifies the separation: a Jackknife standard deviation of 1.24 for the correct model versus 2.28 for the incorrect model suggests that values well above 1 indicate low predictive power.","In the MACS0647 demonstration, the larger JWST image sample (86 images) yields a Jackknife width much closer to 1 than the smaller HST sample (31 images), supporting the idea that more constraints curb overfitting.","If the sigma_Jackknife/sigma versus sigma_realization/sigma_MCMC trend is established, one could correct underestimated MCMC errors for magnifications and time delays using observed data alone.","In the 20-source simulation, the incorrect model is no longer overfitted, so the Jackknife distributions agree; the method detects overfitting rather than model misspecification in general."],"supporting_citations":[{"why":"Supplies the lens-modeling code and fitting machinery used to generate the mock strong-lens data and to construct the MACS0647 mass models.","marker":"Oguri 2010, 2021"},{"why":"Ray-tracing simulations used to argue that dark-matter substructure adds a positional uncertainty that makes reduced chi-square unreliable.","marker":"Meneghetti et al. 2017"},{"why":"Documents typical image-plane RMS values of roughly 0.4-0.5 arcsec, used to set the effective positional error in the simulations.","marker":"Cha and Jee 2023"},{"why":"Provides the member-galaxy scaling-relation parameterization adopted in the realistic simulation and in the MACS0647 mass models.","marker":"Kawamata et al. 2016"},{"why":"HST multiple-image identifications, 31 images from 11 sources, that define the HST mass model of MACS0647.","marker":"Zitrin et al. 2015"},{"why":"The HST mass model of MACS0647 with chi-square per degree of freedom 24.3/20 used as the HST demonstration case.","marker":"Okabe et al. 2020"},{"why":"JWST multiple-image identifications used to build the new JWST mass model of MACS0647.","marker":"Meena et al. 2023"},{"why":"The NFW density profile used for the dark-matter haloes in the input and fitting models.","marker":"Navarro et al. 1997"}],"fun_headline_variants":["Jackknife exposes lens overfitting that chi-square misses","Hold-out images reveal when lens models are overfit","Jackknife residual test beats chi-square for lens models","Jackknife catches overfit lens models chi-square can't"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a correct model's Jackknife residuals should scatter with standard deviation 1, but the paper states this expectation without deriving it; its own correct-model simulation gives 1.24, so the threshold cannot cleanly separate correct-but-sparsely-constrained models from incorrect ones without calibration.","fun_headline_variants_meta":{"raw":{"variants":["Jackknife exposes lens overfitting that chi-square misses","Hold-out images reveal when lens models are overfit","Jackknife residual test beats chi-square for lens models","Jackknife catches overfit lens models chi-square can't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1269,"prompt_tokens":975,"completion_tokens":294,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":591,"completion_tokens_details":{"reasoning_tokens":227}},"tokens_in":591,"tokens_out":294,"duration_ms":3131,"temperature":1.0,"reasoning_tokens":227,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:39:59.082488+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same leave-one-source-out analysis on simulated clusters with a known true model but deliberately fewer constraints or strong parameter degeneracies, keeping chi-square per degree of freedom near 1; if the correct model's Jackknife standard deviation rises well above 1, the fixed N(0, 1) reference would mislabel correct models as overfitted. Equivalently, an analytic calculation of the expected leave-one-out residual variance for a linearized lens model with 2P constraints and Q parameters would settle whether the reference value should depend on the degrees of freedom.","supporting_citations":[],"review_version":1}