{"id":"8f5accf7-bcb6-4265-bc1f-56bbad765789","arxiv_id":"2602.03934","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Dynamical systematics in eight time-delay lenses can shift H0 by more than the 2% target, so velocity dispersions must be controlled to about 1% accuracy.","lead":"This paper tests many small measurement and modeling errors in time-delay lens galaxies and finds they can shift the inferred Hubble constant by more than the 2% precision goal. It warns that combining more lenses will not erase errors that affect all early-type galaxies in the same way.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative amplitudes rest on spherical SIS/Hernquist models and are not propagated to H0; broad >2% conclusion likely robust, magnitudes unverified.","rationale":"The paper is a careful, transparent systematics study, and its most novel element—the population-level marginalization argument in §5—is analytically correct under the stated homogeneity premise. The reader's CONDITIONAL verdict is appropriate. My read does not move the verdict: the main residual risk is that Table 2's percentages are model-dependent. Because all calculations share the SIS/Hernquist assumption, and because the paper itself flags nonphysical SIS behavior and does not run an end-to-end H0 inference, the specific amplitudes should be treated as illustrative rather than definitive. A single targeted recomputation with a realistic mass model and an H0 refit would settle whether the >2% headline is quantitatively robust. I therefore recommend UNCHANGED (still CONDITIONAL), and I partially agree with the reader: the weakest assumption is indeed the spherical SIS/Hernquist setup, but the more precise concern is the combination of that setup with the lack of propagation into an actual H0 posterior.","tokens_in":33003,"tokens_out":8750,"duration_ms":98260,"concrete_test":"Choose RX J1131−1231 and rerun its TDCOSMO-style lens+dynamics model replacing the SIS/Hernquist dynamical model with the actual best-fit power-law/composite mass model and double-Sersic light profile. Apply the Table 2 Cuddeford β0=0.5 and Sersic n=2 systematic shifts to the kinematic likelihood and refit H0. If the largest resulting H0 shifts remain above 2%, the central claim is confirmed; if they fall below 2%, the quantitative claim needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 2, the quantitative core of the paper's claim that most systematics bias H0 by >2%, is computed throughout with a spherical SIS total-mass potential and a spherical Hernquist stellar tracer. §2.2 explicitly notes the SIS potential creates 'nonphysical behavior at the center' that affects the NIRSpec comparisons, and the photometric-model rows use Hernquist/Jaffe/Sersic tracers in the same SIS potential. The actual TDCOSMO mass models are power-law or composite and often triaxial; the paper does not propagate the Table 2 Δσ^2 shifts through an H0 posterior, instead using the simple SIS-based sensitivity ψ (§1). If the true mass models differ, the amplitudes—especially the Sersic n=2 rows (up to 40%) and Cuddeford rows (up to 18%)—could change materially, and the mapping to H0 could be diluted. This is a quantitative-accuracy concern, not a refutation: several independent effects (velocity DF 3–8%, Cuddeford 5–18%, Sersic n=3 2–8%) already exceed the 2% budget under the stated assumptions, so the broad conclusion likely survives.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper assembles a systematic-error budget for the stellar dynamical constraints used in time-delay cosmography, applied to the eight lenses in the TDCOSMO/H0LiCOW joint analyses. Working mostly with spherical SIS total-mass models and Hernquist stellar tracers, the authors compute fractional biases in the squared velocity dispersion from PSF width and shape, aperture miscentering, the difference between the measured Gaussian dispersion and the Jeans rms velocity, the choice of orbital anisotropy, and the assumed stellar light profile. They conclude that most of these effects exceed the 2% σ² budget required for a 2% H0 measurement, and they argue that because early-type galaxies form a fairly homogeneous population, nuisance parameters such as the anisotropy radius must be marginalized globally rather than independently per lens. The paper also derives a population-weighted Jeans equation for mixed stellar populations and uses it to argue that the photometric profile entering dynamical models should be weighted by line equivalent width, not by broad-band flux.","tokens_in":33345,"tokens_out":7249,"duration_ms":85575,"significance":"If the quantitative conclusions hold, this is a useful and timely cautionary analysis for the TDCOSMO program. The paper is transparent about its approximations, provides analytic Jeans solutions and Monte Carlo line-of-sight velocity distributions, and makes a falsifiable point that Osipkov-Merritt models fail to reproduce the h4 values observed in early-type galaxies. The strongest contribution is the systematic side-by-side comparison in Tables 2 and 5, which gives the community a concrete checklist of where the 2% budget is violated. The population-level homogeneity argument in §5, though in need of empirical calibration, is conceptually important and could affect how future lens samples are averaged. The main weakness is that the quantitative amplitudes are computed for idealized spherical SIS/Hernquist models and are not propagated through an H0 posterior, so the mapping from Δσ²/σ² to ΔH0/H0 remains schematic.","major_comments":[{"comment":"The quantitative core of the paper, Table 2, is computed for a spherical SIS total mass distribution and a Hernquist tracer with s=0.55 R_e q^{1/2}. The authors state in §2.1 that this is for simplicity and due to unavailability of the TDCOSMO double-Sérsic parameters, and §2.2 notes that the SIS potential produces 'nonphysical behavior at the center' that affects the NIRSpec comparisons. Since the actual TDCOSMO mass models are power-law or composite and often triaxial, the amplitudes in Table 2, especially the Sérsic n=2 rows (up to +40%) and Cuddeford rows (up to +18%), are model-dependent. The mapping to H0 is made only through the SIS-based sensitivity ψ of Eq. (1), not through an H0 posterior. The broad claim that systematics exceed 2% is likely robust because several independent effects already exceed 2% under the stated assumptions, but the quantitative amplitudes are not yet ful","section":"§2.1–§2.3, Table 2, Eq. (1)"},{"comment":"The double-Moffat entry for RX J1131 observed with NIRSpec has a positive sign only because, as the authors write, the changes in the numerator and denominator are affected by 'the nonphysical behavior at the center from using an SIS potential.' This is an artifact of the model, not a physical PSF effect. Reporting it in Table 2 without a caveat overstates the certainty of that particular systematic. The same inverted trends for other NIRSpec systems are flagged in the text, so the table should either recompute these entries with a cored or power-law potential or mark them as model-dependent.","section":"§2.2, Fig. 4, Table 2"},{"comment":"The formal argument that a dynamically homogeneous population prevents the anisotropy radius uncertainty from averaging down with N is mathematically correct for a single shared r_a/s with a fixed Gaussian prior. However, the paper does not provide an empirical estimate of the correlation scale for r_a/s, stellar-template offsets, or light-profile choice across the lens sample. Early-type galaxies have mass and redshift trends, which the paper acknowledges, so the shared-value case is an upper limit. A partially correlated model, e.g. a correlation coefficient between lens-to-lens values, would interpolate between Eq. (17) and Eq. (19). The phrase 'must be marginalized over the lens sample as a whole' overstates the case as written; the homogeneous limit should be framed as one end of a correlation model, ideally quantified with simple test cases.","section":"§5, Eqs. (15)–(19)"}],"minor_comments":[{"comment":"The Gaussian prior is written as P(r_a) ∝ exp(−r_a/2σ_a²), which is an exponential, not a Gaussian. It should read exp(−r_a²/(2σ_a²)). The analytic marginalizations in Eqs. (17) and (19) rely on a true Gaussian prior, so this typo should be corrected.","section":"§5, text near Eq. (15)"},{"comment":"The RX J1131−1231 row is ambiguous: the seeing and aperture entries '0.15/0.96π×0.955 2' need a clearer layout specifying which values apply to the NIRSpec and KCWI observations. Also state explicitly that all FWHM and aperture values are in arcseconds.","section":"Table 1"},{"comment":"The Monte Carlo distributions use 10^9 particles and show residual shot noise. Since some reported differences are only 0.7–1.6%, a quantitative estimate of the Monte Carlo noise floor would help the reader judge which entries in the velocity-DF rows of Table 2 are significant.","section":"§2.3, Figs. 5–6"},{"comment":"The paper says LSF-related systematic errors were considered but are negligible compared with other sources, but no calculation is shown. Given that template/LSF issues are known to be non-negligible in some dispersion measurements, one sentence summarizing the magnitude or a pointer to the relevant check would be useful.","section":"§2.2"},{"comment":"The caption says the weighted mean temperature is 'B-band luminosity weighted,' but the text says this is roughly the rest wavelength range usually modeled. Clarify whether the weight is a B-band filter or the actual spectral window used in the dispersion fits, since the two need not be identical.","section":"Fig. 9 caption"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the SIS/Hernquist model dependence lands: Table 2 is the quantitative core, and its amplitudes are not propagated to H0. The broad >2% conclusion is probably correct, but the manuscript would be much stronger with one explicit robustness calculation using a power-law or composite mass model and a mapping to H0. The self-citations to Kochanek are natural given the prior work in this specific line of research; I do not see a citation-pattern concern. Overall, this is a solid systematic-effects paper that should be published after the model-dependence issue is addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Broad picture first: this is a genuine systematics study, not another H0 measurement, and it makes the point convincingly that the 2% sigma^2 requirement is threatened by a string of modeling choices. The strongest new content: (1) Cuddeford models expand the range of allowed sigma^2 relative to constant-beta and Osipkov-Merritt; (2) the equivalent-width weighting argument for which light profile to use in Jeans modeling; (3) the analytic demonstration that if the lens population is dynamically homogeneous, shared nuisance parameters like r_a/s do not get averaged down with N. The last point is the most novel and I think it is correct under the stated premise. The paper also does a real service by tabulating maximal shifts for the eight TDCOSMO lenses (Table 2) and by revisiting the older ground-based measurements in the appendix; that comparison is useful for anyone aggregating lens H0 results.\n\nWhat it does well: the Jeans calculations are transparent and internally consistent; the Monte Carlo LOSVD work is careful (10^9 particles); the authors flag their own limitations, including the unphysical SIS central behavior affecting NIRSpec comparisons and the lack of a verified DF for the Mamon-Lokas profile. The h4 argument—that Osipkov-Merritt DFs underproduce the observed fourth moments—is a good, physical reason to worry about that model family.\n\nSoft spots: nearly all of the quantitative amplitudes in Table 2 are computed with one fixed spherical SIS potential and a spherical Hernquist tracer (or Sersic/Jaffe in the same potential). The actual TDCOSMO mass models are power-law or composite and often triaxial. The authors acknowledge this, but it means the 5–40% numbers should not be read as exact corrections; they are an upper-limit family under one mass model. Also, the systematics are not propagated through an H0 posterior—the psi sensitivity is a heuristic. No code or data are provided, which limits immediate checking. These are proportionally moderate concerns: none of them sinks the central claim, because several independent effects (velocity DF 3–8%, Cuddeford 5–18%, Sersic n=3 2–8%) already exceed the 2% budget. The homogeneous-population argument also assumes strong homogeneity; real ETGs have scatter, and the authors note that the question is the degree of correlation. The equivalent-width argument is a nice proof-of-concept but wasn't applied to actual spectra.\n\nWho should read it: anyone working on time-delay lens cosmography, dynamical modeling of lens galaxies, or planning H0 constraints with TDCOSMO. It deserves a serious referee; it will be a reference for systematics budget discussions even if the exact magnitudes get revised with more realistic mass models.","headline":"A careful, transparent systematics study that makes a convincing broad case that kinematic modeling choices can break the 2% H0 budget; treat Table 2 as an illustrative warning, not exact corrections.","tokens_in":33817,"tokens_out":2727,"would_cite":true,"duration_ms":29569,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the systematic uncertainties in measuring lens-galaxy velocity dispersions, not statistical noise, set the real floor on time-delay measurements of the Hubble constant, with biases in H0 typically exceeding the 2% pre","keywords":["gravitational lensing","time-delay cosmography","Hubble constant","stellar velocity dispersion","orbital anisotropy","point spread function systematics","early-type galaxies","Jeans equations"],"falsifier":"Measure the fourth Gauss–Hermite moment h4 of the eight lens galaxies; if early-type lenses have h4 near the typical 0.01–0.04 seen in local ellipticals, the claim that Osipkov-Merritt models underestimate anisotropy systematics is supported, and if h4 is near zero the up-to-18% anisotropy shifts are overestimated. Alternatively, compare measured dispersions of the same lenses fitted with Gaussian versus double-Moffat PSFs: the predicted 2–6% shifts, and the larger 5–40% shifts from light-profile choices, should appear in the data.","tokens_in":32895,"feed_emoji":"🔭","tokens_out":6136,"duration_ms":61377,"temperature":0.7,"pith_summary":"The paper argues that stellar-dynamical systematics, not statistical noise, are the real limit on measuring the Hubble constant from time-delay lenses. Because H0 errors scale directly with errors in the squared velocity dispersion (ΔH/H0 ∝ Δσ²/σ²), biases of 2–40% in σ² translate into biases of several percent in H0, far above the 2% target. The paper identifies five sources: PSF modeling, the difference between the measured dispersion and the true mean-square velocity, choices of orbital-anisotropy models, mismatches between the true light profile and the assumed Hernquist profile, and population-level correlations. It also argues that early-type lens galaxies are dynamically homogeneous, so shared parameters like the anisotropy radius must be marginalized once for the whole sample; their uncertainties then do not shrink with more lenses. A sympathetic reader would take away that the current error budget for time-delay cosmography is dominated by systematics that averaging over lenses cannot remove.","feed_headline":"Most lens kinematic errors push H0 estimates past 2 percent","feed_subtitle":"PSF, anisotropy, and light-profile choices can bias lens velocity dispersions by up to 40%.","key_machinery":"The driving mechanism is the sensitivity relation ΔH/H0 ∝ Δσ²/σ², which converts every fractional kinematic error into a proportional error in H0. The population-level machinery is the marginalization identity for a shared anisotropy radius: with N independent lenses the systematic variance shrinks as (σ_H² + α²σ_a²)/N, but with a homogeneous population it becomes σ_H²/N + α²σ_a², so the anisotropy contribution never averages away. The numerical work uses spherical Jeans models with a Hernquist stellar tracer inside a singular-isothermal-sphere mass distribution to compute how each systematic shifts the predicted velocity dispersion.","core_discovery":"The central claim is that most systematic errors in the kinematic measurements of the eight time-delay lenses can bias H0 by more than 2%, with some per-lens fractional changes in σ² reaching 40%. Specifically: a 10% error in PSF FWHM causes 0.2–2.6% changes; replacing a Gaussian PSF with Moffat wings causes up to 2–6%; the difference between the measured dispersion and the true mean-square velocity can be 2–8%; anisotropy model choice causes up to 5–18% shifts, with the commonly used Osipkov-Merritt model failing to produce h4 values typical of early-type galaxies; and changing the assumed stellar light profile from Hernquist to a Sérsic n=2 profile raises the predicted dispersion by 5–40%.","pith_inferences":["A testable prediction of the homogeneity argument is that velocity-dispersion offsets between different template stars should be systematically correlated across early-type lenses of similar redshift; re-fitting archival spectra with matched templates and checking for lens-to-lens correlation would test this directly.","The numerical magnitudes in Table 2 rest on SIS plus spherical Hernquist models; applying the same systematics to the newer double-Sérsic, power-law, or composite mass models would likely change the numbers but not the qualitative conclusion that systematics exceed the 2% budget.","The population-level marginalization logic also applies to template-star choice and light-profile choice: posterior distributions should be computed for each global model assumption and then marginalized over assumptions, rather than averaging per-lens shifts.","The same homogeneous-population reasoning used here has been applied in cluster cosmology to model mass–richness relations; borrowing those correlation-marginalization tools could handle partially correlated lens populations."],"forward_implications":["A 1% bias in the squared velocity dispersion produces roughly a 1% bias in H0 for a typical lens, so kinematics must be known to better than 1% in σ (2% in σ²) to reach the 2% H0 target.","PSF FWHM must be measured to better than 10% accuracy; otherwise seeing errors alone can exceed the 2% budget.","If early-type galaxies are dynamically homogeneous, the uncertainty from the anisotropy radius does not decrease with sample size, so combining more lenses cannot cure this systematic.","Measured dispersions must be corrected to true mean-square velocities using h4 or full velocity distributions; for radially anisotropic galaxies this raises v_rms relative to σ*, typically shifting inferred H0 upward.","Using the real photometric profile (e.g., Sérsic n≈2) instead of the Hernquist profile changes predicted dispersions by up to 40% for some lenses, biasing H0 in either direction depending on the true profile."],"fun_headline_variants":["Lens kinematics biases can push H0 error past 2%","Systematics in lens velocity measurements threaten H0 precision","Up to 40% bias in lens dispersions impacts H0","Time-delay lens kinematic errors exceed H0 precision budget","Hidden systematics inflate lens σ², skewing Hubble constant"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The numerical amplitudes are computed with a spherical singular-isothermal-sphere mass distribution and a spherical Hernquist stellar tracer; if real lens mass distributions are power-law, composite, or triaxial, the percentages in Table 2 could change materially, even though the qualitative conclusion would likely survive — as the paper itself notes when discussing the nonphysical behavior of the SIS potential at the center.","fun_headline_variants_meta":{"raw":{"variants":["Lens kinematics biases can push H0 error past 2%","Systematics in lens velocity measurements threaten H0 precision","Up to 40% bias in lens dispersions impacts H0","Time-delay lens kinematic errors exceed H0 precision budget","Hidden systematics inflate lens σ², skewing Hubble constant"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1191,"prompt_tokens":918,"completion_tokens":273,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":197}},"tokens_in":662,"tokens_out":273,"duration_ms":3052,"temperature":1.0,"reasoning_tokens":197,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T04:47:56.358083+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the fourth Gauss–Hermite moment h4 of the eight lens galaxies; if early-type lenses have h4 near the typical 0.01–0.04 seen in local ellipticals, the claim that Osipkov-Merritt models underestimate anisotropy systematics is supported, and if h4 is near zero the up-to-18% anisotropy shifts are overestimated. Alternatively, compare measured dispersions of the same lenses fitted with Gaussian versus double-Moffat PSFs: the predicted 2–6% shifts, and the larger 5–40% shifts from light-profile choices, should appear in the data.","supporting_citations":[],"review_version":1}