{"id":"6a4377d1-39d8-43d0-913d-b0e6f91219d9","arxiv_id":"2506.18708","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Open clusters in the outer Milky Way are more likely to be old than young even after correcting for Gaia detection biases, implying young massive clusters are rare in the outer disc.","lead":"Astronomers injected hundreds of thousands of fake star clusters into Gaia data to measure which real clusters the census misses. After correcting for these detection biases, old clusters are about three times more common in the outer Milky Way than near the Sun, meaning the shortage of young clusters is a physical fact, not an observational artifact.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 2.97 ratio is computed by dividing observed counts by detection probabilities conditioned on true input parameters; recovered mock clusters are never re-fitted, so biased HR24 masses/ages for old distant clusters could inflate the corrected old-cluster excess.","rationale":"The paper is a substantial, carefully constructed injection-recovery experiment, and the empirical selection function is a real advance. The reader's CONDITIONAL verdict is appropriate: the statistical machinery is strong, but the headline ratio lacks systematic-error accounting. My stress test agrees with the reader's weakest_assumption in broad terms, but sharpens it: the most load-bearing issue is not primarily the assumed warm kinematic model (which the authors partially stress-test with their 'hot' model), but the unmodeled mismatch between the true parameters used to train f_detected and the noisy, possibly biased HR24 parameters used to bin the real clusters. The correction divides observed counts by detection probabilities conditional on true parameters; if the observed parameters are systematically offset for the old distant clusters that dominate the signal, the 2.97 ratio can be inflated without any change in the underlying population. This is a concrete, testable gap rather than a speculative objection, and the authors themselves flag parameter quality as a limitation in Sect. 6. Because the concern is real but does not, on the evidence available, force a rejection of the qualitative conclusion, the appropriate verdict remains CONDITIONAL; hence UNCHANGED relative to the reader. The proposed test (re-fitting recovered mock clusters and propagating the error/confusion matrix) would settle whether the physical-excess conclusion is robust. I do not see grounds for REJECT or UNVERDICTED, and I credit the paper for its internal validation on real clusters (RMSE 3.75 vs 3.17) and for releasing a detailed methodology, even though the selection-function products are not yet public.","tokens_in":28292,"tokens_out":7607,"duration_ms":88496,"concrete_test":"Run the HR24 and Cavallo et al. (2024) parameter-estimation pipelines (or a calibrated fast surrogate) on the recovered member lists of the 147,639 detected mock clusters, and build the joint confusion matrix between estimated and true (mass, logt, distance). Then recompute the Sect. 5.1 corrected young-cluster fraction and the RGC=13 kpc ratio after convolving the observed HR24 bin counts with this matrix. If the 2.97 ratio shifts by more than roughly 30% or becomes consistent with unity, the physical-excess conclusion is not yet supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline ratio in Sect. 5.1 is obtained by binning HR24 clusters by measured mass, distance, and age, then weighting each bin by 1/f_detected, where f_detected comes from the XGBoost model trained on the true input parameters of injected mock clusters (Sect. 4.2). This correction is unbiased only if the measured cluster parameters used for binning are close to the true parameters used to train the model. The paper's own caveat in Sect. 6 admits that the selection function is limited by the quality of cluster parameters, but the size and direction of the resulting bias are never quantified. The danger is concrete: old, distant clusters in Gaia DR3 are often detected through only a handful of bright giants, so the HR24 Jacobi masses and Cavallo et al. (2024) ages inferred from the recovered members may be systematically offset. If the measured masses of old anticentre clusters are biased low relative to their true masses, those clusters fall into lower-mass bins with smaller f_detected, so the inverse correction overestimates their true counts. Because this effect is strongest for exactly the old, distant population that drives the 2.97 excess, the central claim that the old-cluster excess is physical could be at least partly an artifact of parameter-recovery bias (a Malmquist/Eddington-type effect). The injection-recovery experiment measures detection probability conditional on true parameters, but it never re-derives parameters for the recovered mock clusters, leaving this contamination completely unmeasured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constructs an empirical selection function for the open cluster census in the Galactic anticentre (140° ≤ ℓ ≤ 240°, |b| ≤ 10°, d ≥ 2 kpc), based on the HR23/HR24 HDBSCAN cluster search in Gaia DR3. The authors generate 194,752 realistic mock clusters spanning broad ranges in mass, age, distance, position, proper motion, and internal structure; inject them into Gaia DR3 data; and attempt blind recovery with the same HDBSCAN pipeline used for HR23. A gradient-boosting (XGBoost) model is trained on the resulting 147,639 recovered clusters to predict detection probability as a function of cluster parameters, yielding a smooth selection function. The main detectability drivers are found to be mass, extinction, distance, and age, with secondary effects from proper motion. Applying this selection function to HR24 clusters through a forward-modelled 'warm' kinematic correction, the paper reports that old clusters (log t > 8.5) are 2.97 ± 0.11 times more common at Galactocentric radius 13 kpc than in the solar neighbourhood, and interprets this as a physical property of the Milky Way rather than an observational bias. Additional applications address the Galactic warp in the old cluster population and the detectability of the distant clusters Berkeley 29 and Saurer 1.","tokens_in":28609,"tokens_out":5675,"duration_ms":62967,"significance":"If the central claim holds, the paper provides a strong, quantitative demonstration that the old-to-young open cluster ratio increases toward the outer Milky Way, with consequences for cluster formation thresholds, radial migration, and cluster destruction rates. The methodological contribution — an empirical, injection-based selection function for a full Gaia cluster catalogue — is important and likely to be reused for other surveys and regions. The paper has notable strengths: the injection-recovery experiment is large and closely replicates the original detection pipeline; the selection function is derived independently of the observed catalogue, so the central conclusion is not circular; and the CST predictor is validated against real clusters with an RMSE of 3.75, comparable to its simulated validation RMSE of 3.17. The principal weakness is that the size and direction of systematic errors in the headline 2.97 ratio, particularly those arising from the use of measured (rather than true) cluster parameters in the correction, are not quantified.","major_comments":[{"comment":"","section":"§5.1, Fig. 8; §4.2"},{"comment":"","section":"§5.1"},{"comment":"","section":"§4.3"}],"minor_comments":[{"comment":"The abstract states the 2.97 ± 0.11 ratio without noting that the analysis in Fig. 8 is restricted to clusters with masses between 250 and 2000 M☉ and to the adopted anticentre region and binning; please state this restriction in the abstract or results summary.","section":"Abstract and §5.1"},{"comment":"The caption says 'clusters in HR23 are shown by the blue points', while the text in §4.3 refers to HR24 clusters; please check which catalogue is plotted and make the caption consistent.","section":"Fig. 7 caption"},{"comment":"When the below-disk cluster population is flipped in Z to estimate detection probabilities above the disk, the text should clarify how extinction and line-of-sight quantities are handled, since the selection function depends on l, b, d and therefore implicitly on the dust column at the new position.","section":"§5.2"},{"comment":"The assumption of virial equilibrium (η = 10) is stated to have negligible impact at the distances considered, but the text could note the expected maximum effect on the internal velocity dispersion for the nearest simulated clusters, to make this assertion easier to verify.","section":"§3.1.2"}],"recommendation":"major_revision","confidential_remarks":"The central concern is parameter-recovery bias in the Sect. 5.1 correction. If the authors can add a re-fitting or perturbation test demonstrating that the 2.97 ratio is robust to realistic HR24 and Cavallo et al. (2024) parameter errors, I would support publication. Without such a test, the headline claim is not fully quantified, although the direction of the effect is well supported. The paper is otherwise thorough and well within the scope of A&A."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers the first empirical selection function for a Gaia DR3 open cluster catalogue, built by injecting 192k mock clusters into real Gaia data and re-running the original blind search. That alone is a substantial step forward. The main scientific claim – old clusters are ~3x more common at 13 kpc than locally, after correcting for selection – is plausible and probably correct. The mock cluster generation is careful, the recovery experiment is large and mostly blind, and the validation on real clusters (CST predictor RMSE 3.75 vs 3.17 on simulated data) gives some genuine confidence that the simulations capture the right physics.\n\nThe paper is also honest about its limits. They clearly state the selection function depends on the quality of cluster parameters, and they flag the moving-group classification and CMD classifier issues. The warp analysis and the Berkeley 29 / Saurer 1 discussion are sensible uses of the selection function, not overplayed.\n\nThe soft spot the stress-test note identifies is real and worth taking seriously. The completeness estimates used for the 2.97 ratio are conditioned on the true input parameters of injected mocks, but real clusters are binned by their measured masses, ages, and distances. If the measured parameters of old, distant clusters are biased – e.g., masses underestimated because only a few bright giants are recovered – those clusters fall into lower-mass bins with smaller detection probabilities, and the inverse correction inflates the estimated old-cluster count. The paper never re-fits recovered mock clusters to test how measured parameters scatter around true ones. I don't think this kills the result, because the age and distance trends are strong and the hot-model test shows the direction is hard to reverse, but it does mean the quoted 2.97 ± 0.11 is formally only a Poisson error. The systematic error from parameter-recovery bias could be comparable to or larger than that.\n\nA second, minormiss: the selection function products aren't released in this paper, only promised. That limits immediate use by others.\n\nWho this is for: anyone doing Gaia-era cluster census work, or using open clusters as tracers of Galactic structure. It deserves a serious referee. I would send it to review, but the authors should be asked to quantify the parameter-recovery systematics, even if only with a simplified mock-refit experiment, and to release the selection function maps and trained models.","headline":"Impressive injection-recovery selection function; the old-cluster excess is likely real, but the headline 2.97 ratio carries unquantified systematics from using true parameters for completeness and measured ones for real clusters.","tokens_in":29195,"tokens_out":1573,"would_cite":true,"duration_ms":19041,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After correcting for selection effects, old open clusters are $2.97\\pm0.11$ times more common at a Galactocentric radius of 13 kpc than in the solar neighbourhood, and this excess is physical rather than a bias of the census.","keywords":["open clusters","Galactic anticentre","selection function","completeness","Gaia DR3","HDBSCAN","Galactic structure","cluster formation"],"falsifier":"Run a deeper, independent blind cluster search in the anticentre and count young ($\\log t<8.5$) clusters of mass 250 to 2000 $M_\\odot$ at $R_{\\rm GC}\\approx13$ kpc: if the completeness-corrected young fraction comes out well above $27.5\\%\\pm5.7\\%$, the claimed $2.97\\pm0.11$ excess of old clusters is wrong.","tokens_in":28086,"feed_emoji":"🌌","tokens_out":12298,"duration_ms":106255,"temperature":0.7,"pith_summary":"This paper builds an empirical selection function for the HR24 open-cluster catalogue and uses it to ask whether the outer Milky Way really contains a different mix of old and young star clusters than the solar neighbourhood. The authors generate 192,318 realistic mock clusters, inject them into Gaia DR3 data, and try to recover them with the same blind HDBSCAN-based search that produced the catalogue. They find that cluster mass, distance, extinction, and age control detectability, with old clusters much harder to recover because they contain few bright stars. After correcting for these biases, they conclude that clusters older than $\\log t = 8.5$ are $2.97 \\pm 0.11$ times more common at a Galactocentric radius of 13 kpc than near the Sun, and that this excess cannot be explained by observational bias. If true, the result means the outer disc is not forming young clusters massive enough to be seen in Gaia, while a population of old clusters remains hidden there.","feed_headline":"Old star clusters outnumber young ones in the outer Milky Way","feed_subtitle":"A bias-corrected Gaia census finds old clusters 2.97 times more common at 13 kpc than near the Sun.","key_machinery":"The load-bearing object is an empirical selection function built by injecting realistic mock open clusters into Gaia DR3 and attempting to recover them with the same HDBSCAN blind search used to construct the catalogue. A gradient-boosted classifier converts the injection-recovery outcomes into a detection probability $f_{\\rm detected}$ as a function of longitude, latitude, distance, mass, age, and proper motion; cluster core and tidal radii are dropped after showing negligible effect. This function is then used to correct the observed cluster counts in mass, distance, and age bins, and to forward-model what a kinematically 'warm' versus 'hot' outer-disc cluster population would look like.","core_discovery":"The central claim is that the observed excess of old open clusters in the Galactic anticentre is a real property of the Milky Way, not a selection artefact. Using 147,639 recovered mock clusters out of 192,318 injected, the paper builds a completeness model and finds that the fraction of clusters younger than $\\log t = 8.5$ drops from $81.7\\%\\pm9.2\\%$ in the solar neighbourhood to $27.5\\%\\pm5.7\\%$ at $R_{\\rm GC}=13$ kpc, equivalent to old clusters being $2.97\\pm0.11$ times more common there. The paper further argues that this deficit of young outer-disc clusters is best interpreted as a limit on the mass of clusters that can form in the outer Galaxy, that many low-mass old clusters remain undiscovered, and that the two most distant known clusters, Berkeley 29 and Saurer 1, are likely the high-latitude tip of that hidden population.","pith_inferences":["Inference: applying the same injection-recovery machinery to the whole sky would separate algorithm-specific detection biases from physical gradients, allowing a direct test of whether the $2.97$ ratio persists beyond the anticentre.","Inference: the migration explanation carries a testable chemical signature: old clusters in the outer disc should show inner-disc metallicities and abundance patterns if they migrated outward, whereas a lower destruction rate predicts a shallower age gradient in the anticentre.","Inference: the selection function predicts specific dust-obscured low-latitude zones between $R_{\\rm GC}=14$ and $19$ kpc where dozens of hidden low-mass old clusters should reside, a prediction that deeper astrometric or co-added ground-based surveys could check directly."],"forward_implications":["Completeness-corrected counts place the young-cluster fraction at $27.5\\%\\pm5.7\\%$ at $R_{\\rm GC}=13$ kpc, versus $81.7\\%\\pm9.2\\%$ near the Sun, so old clusters are $2.97\\pm0.11$ times more common in the outer disc.","The outer Galaxy is not forming young clusters massive enough to be identified in Gaia DR3, so $R_{\\rm GC}\\sim13$ kpc is likely a limit for massive cluster formation rather than for star formation itself.","The old-cluster census in the anticentre is probably very incomplete, with many low-mass or low-latitude old clusters still undiscovered; Berkeley 29 and Saurer 1 are likely the visible high-latitude part of that population.","The observed asymmetry of the cluster warp, with more old clusters below the disc than above it, survives selection correction at the $3.7\\sigma$ level and would appear roughly twice as strong in a complete census."],"supporting_citations":[{"why":"The HR24 cluster catalogue whose selection function this paper constructs, supplying cluster parameters, masses, and radii used in the correction.","marker":"Hunt & Reffert (2024)"},{"why":"Supplies the HDBSCAN + cluster significance test blind-search pipeline that is replicated exactly for injection and recovery.","marker":"Hunt & Reffert (2023)"},{"why":"Provides the ages and extinctions used for real clusters, which define the old/young split that drives the result.","marker":"Cavallo et al. (2024)"},{"why":"Supplies the 3D dust map that sets extinction for mock clusters, with extinction one of the dominant detectability parameters.","marker":"Green et al. (2019)"},{"why":"The source of the position- and magnitude-dependent selection function used to emulate the Gaia subsample quality cuts in HR23.","marker":"Castro-Ginard et al. (2023)"},{"why":"Supplies the flared vertical density profile with scale height growing as $t^{1.3}$, used in the 'warm' forward model.","marker":"Cantat-Gaudin et al. (2020)"},{"why":"Supplies the age-velocity relation used in the 'warm' kinematic model for corrected cluster counts.","marker":"Tarricq et al. (2021)"},{"why":"Supplies the MWPotential2014 potential used to compute tidal radii, orbits, and rotation in the simulations and forward model.","marker":"Bovy (2015)"},{"why":"The theory of a galactocentric-dependent maximum cluster mass that the paper invokes to interpret the deficit of young massive outer-disc clusters.","marker":"Pflamm-Altenburg & Kroupa (2008)"}],"fun_headline_variants":["Outer Milky Way old clusters outnumber young 3:1 after bias fix","Galactic anticentre's old cluster excess is real, not selection bias","Old clusters 3x more common in Milky Way's outer disc - no bias","Outer Galaxy lacks young clusters - ancient ones dominate","Milky Way's anticentre hides more old clusters - young missing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that the simulated cluster population and the catalogue's mass and age estimates reproduce the true outer-disc population, so that the correction bins are not systematically mis-scaled.","fun_headline_variants_meta":{"raw":{"variants":["Outer Milky Way old clusters outnumber young 3:1 after bias fix","Galactic anticentre's old cluster excess is real, not selection bias","Old clusters 3x more common in Milky Way's outer disc - no bias","Outer Galaxy lacks young clusters - ancient ones dominate","Milky Way's anticentre hides more old clusters - young missing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000675,"raw_usage":{"total_tokens":3137,"prompt_tokens":1075,"completion_tokens":2062,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":691,"completion_tokens_details":{"reasoning_tokens":1967}},"tokens_in":691,"tokens_out":2062,"duration_ms":16655,"temperature":1.0,"reasoning_tokens":1967,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:43:58.175399+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a deeper, independent blind cluster search in the anticentre and count young ($\\log t<8.5$) clusters of mass 250 to 2000 $M_\\odot$ at $R_{\\rm GC}\\approx13$ kpc: if the completeness-corrected young fraction comes out well above $27.5\\%\\pm5.7\\%$, the claimed $2.97\\pm0.11$ excess of old clusters is wrong.","supporting_citations":[{"cited_title":"& Kroupa, P","cited_arxiv_id":null,"evidence_quote":"The theory of a galactocentric-dependent maximum cluster mass that the paper invokes to interpret the deficit of young massive outer-disc clusters."}],"review_version":2}