{"id":"95979b20-9fb3-4c4a-911e-b67ae45ff74b","arxiv_id":"2506.00240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Adding Gaia-like biases to STARFORGE simulated groups under-recovers masses by up to 80%, overestimates sizes by up to 10x, inflates virial parameters by up to 10x, and leaves traceback ages reliable only for truly unbound systems.","lead":"This paper adds realistic Gaia-like selection effects to simulated stellar groups from the STARFORGE suite, finding that mass, size, and velocity dispersion measurements can be wrong by up to 100%, with low-mass groups hit hardest. The same biases can inflate the virial parameter tenfold, which undermines claims that kinematic traceback ages are reliable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Synthetic observed groups are not re-identified after biases: HDBSCAN is run only on true simulation data, so the reported virial-inflation factor and traceback threshold may not apply to Gaia-selected associations.","rationale":"The paper's central message is a caution about interpreting observed virial parameters and traceback ages, and that message is supported by the simulation-based synthetic observations. The most load-bearing step in constructing those synthetic observations is the transfer of biases to the group properties; the weakest part of that step is that group membership is fixed from the unbiased simulation before the biases are applied (§2.4, §2.7). Because the entire measured property set—mass, half-mass radius, velocity dispersion, virial parameter, and traceback age—depends on membership, not re-running the group finder on the biased catalog leaves the synthetic observed population inconsistent with the actual Gaia detection process. This is an internal methodological gap, not simply an external question about whether STARFORGE represents typical Milky Way clouds, and it is directly testable. The reader identified the bias model as a weak assumption but focused on specific omissions (reddening, radial-velocity incompleteness); the group-reidentification issue is a more structural omission that could change the quantitative claims. I recommend keeping the reader's CONDITIONAL verdict: the concern is real and warrants an additional check, but it does not by itself overturn the qualitative cautionary conclusion about the unreliability of observed virial parameters and traceback ages. The Bx100 exclusion from the traceback analysis (§3.3) is a secondary model-selection issue that I do not treat as the primary concern.","tokens_in":24794,"tokens_out":7773,"duration_ms":77963,"concrete_test":"Re-run the analysis of §2.7 on the full synthetic observed catalogs: for each simulation snapshot at 20, 25, and 30 Myr, apply the magnitude limits, binary-removal rules, and 15% source loss to the entire stellar population, then apply the same HDBSCAN group-finding procedure used in §2.4 (with the same N_min and k) to the biased catalog. Match recovered groups to true groups by membership overlap, and recompute α_vir,obs and the traceback fractional errors. If the recovered-groups-only α_vir,obs values and the fractional-traceback-error versus α_vir relation are statistically consistent with Figure 4 and Figure 8, the concern is resolved; if the inflation factor or the α_vir>2 threshold shifts materially, the quantitative conclusions in §3.1 and §3.3 require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.4 explicitly identifies groups by running HDBSCAN on the unbiased simulation data, and only then applies the observational biases of §2.7. The synthetic observed groups therefore inherit the true membership of the simulation groups, whereas real Gaia associations are themselves products of clustering algorithms applied to the biased catalog (e.g., HDBSCAN in SPYGLASS). This difference matters for the quantitative claims, because the measured mass, half-mass radius, velocity dispersion, and hence the virial parameter are all defined relative to the recovered membership. If group finding were re-run after the biases, some low-mass or heavily incomplete groups would not be detected, other groups might merge or fragment, and interlopers could be included; the population of detected groups would not be identical to the population of true groups. The factor-of-ten virial inflation in Figure 4 and the 20% traceback accuracy threshold in Figure 8 are computed without this detection step, so they describe how the properties of a known group are biased, not necessarily how the properties of groups that appear in a Gaia-based catalog are biased. A selection effect could make the observed sample preferentially complete, reducing the virial inflation for detected groups, or could act in the opposite direction. Either way, the transfer of these numbers to real associations is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper uses the STARFORGE MHD simulation suite, post-processed with 200 Myr N-body integrations, to build synthetic observations of young stellar groups. After identifying groups with HDBSCAN on the unbiased simulation data, the authors apply observational biases: Gaia magnitude limits and saturation, unresolved and wide binary cuts, a 15% source-loss rate, and distance-dependent resampling. They measure how inferred total mass, half-mass radius, and velocity dispersion deviate from true values, finding up to 100% errors in mass and radius for groups below ~100 Msun, and a systematic upward bias in the virial parameter that can reach a factor of ten. They compare the resulting synthetic properties to the Cepheus Far North (CFN) association and match four of its seven subgroups. Finally, they test linear kinematic traceback ages against the true dynamical age from the simulations, reporting that traceback is accurate to within 20% only when the true virial parameter exceeds about 2, and that observational biases can erase the relation between traceback error and virial parameter. They also find no correlation between the dynamical-age minus stellar-age difference and the embedded-phase duration.","tokens_in":25112,"tokens_out":6213,"duration_ms":60578,"significance":"If the results hold, the paper has a direct and important implication for Gaia-era studies of young associations: observed virial parameters and traceback ages can be substantially biased, so associations that appear unbound may not be suitable for traceback dating. The Monte Carlo synthetic observation procedure is clearly specified, and the central error curves are emergent outputs of the simulations rather than parameters tuned to match observed associations. The comparison with CFN is concrete and falsifiable, and the finding that massive stars are preferentially missing from the observed group mass is a useful caution for mass-based dynamical state estimates. The main limitations are that the bias model is incomplete, the group-finding step is not part of the synthetic observation loop, and a few load-bearing analysis choices (matching threshold, exclusion of Bx100) are not tested for sensitivity. These issues are addressable in a revision.","major_comments":[{"comment":"The group-finding step is not part of the synthetic observation loop. HDBSCAN is run on the unbiased simulation data, and only afterwards are the biases of §2.7 applied; the manuscript states this explicitly at the end of §2.4. Real Gaia associations are the output of clustering algorithms applied to the biased catalog, as the paper notes for SPYGLASS in §2.7.1. Because mass, half-mass radius, velocity dispersion, virial parameter, and traceback ages are all membership-conditional quantities, the factor-of-ten virial inflation in Figure 4 and the 20% traceback accuracy threshold in Figure 8 describe the bias of a known group, not the bias of groups that would be detected in a Gaia-like survey. Selection effects could make detected groups preferentially complete or preferentially incomplete, and interloper contamination is not included. The authors should rerun HDBSCAN on the synthetic observed samples and repeat the measurements, or at minimum quantify detection completeness and membership contamination as a function of true group properties.","section":"§2.4, §2.7, Figs. 4 and 8"},{"comment":"The Bx100 model is excluded from the traceback-virial analysis because late-forming stars contaminate the traced trajectories. This exclusion matters for the central traceback claim: Bx100 is the model with the longest star formation duration, and the statement that traceback errors fall below 20% for alpha_vir > 2 is obtained only after removing this model. If similar contamination occurs in real associations, the threshold will not transfer. Please provide an observable criterion for identifying this contamination (for example, a proxy based on age spread or the fraction of stars formed after expansion begins), or show how the Figure 8 relation changes when Bx100 is included.","section":"§3.3, Fig. 8"},{"comment":"The matching criterion D <= 0.6 is introduced without justification or sensitivity analysis. The statement that four of seven CFN groups have good matches, and the conclusion that no single simulation reproduces most CFN groups, depend on this threshold. Changing the threshold to 0.4 or 0.8 would likely change which CFN groups are matched. Please justify the threshold, for example by comparing the D distribution of true analogue pairs with that of random pairs, or report the number of matches as a function of D.","section":"§2.7.2, Eq. (3), §3.2"},{"comment":"The bias model omits reddening and radial-velocity incompleteness. The Discussion acknowledges these omissions, but the abstract and conclusions quote quantitative error magnitudes (up to 100% in mass, factor-of-ten virial inflation, 20% traceback threshold) without stating that these are lower limits or conditional on the included bias set. Since only a small fraction of Gaia sources have radial velocities with km/s-level uncertainties, the velocity dispersion and traceback measurements in real samples are made on a much smaller and kinematically selected subset, which can plausibly change the traceback error distribution in Figure 8. Please add a simple bracketing test, for example a reddening screen and an RV-only subsample, so that the reader can see how much the quoted error bars and thresholds could grow.","section":"§2.7, §4"}],"minor_comments":[{"comment":"The column labels t_dyn and t*_dyn appear reversed relative to the text: §2.6 defines t*_dyn as the best-case estimate using all members and t_dyn as the estimate after observational biases, but Table 1 reports error bars on t*_dyn and no error bars on t_dyn, suggesting the columns are swapped.","section":"Table 1 and §2.6"},{"comment":"There are several typos: 'the the' in the Introduction, 'realty' in §2.3, 'viral' instead of 'virial' in the Discussion, 'Biasses' in the Figure 1 caption, and 'afactoroften' in the Figure 3 caption.","section":"§1, §2.3, §4, Fig. 1, Fig. 3"},{"comment":"The definition of t_dyn,true as the time when the median mutual distance is minimized with at least 50% of members present is reasonable but should be tested for sensitivity to the 50% threshold; the traceback accuracy claim depends on this specific definition.","section":"§2.6"},{"comment":"The caption says 'Side panels shows the distribution of groups from different models ... (R3 and alpha1)', but all six models appear to be shown; please clarify the wording.","section":"Fig. 6 caption"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and useful paper for MNRAS, and the main conclusions are defensible in principle. The requested revision is substantive but feasible: rerunning the group finder after applying biases, sensitivity tests for the matching threshold and the Bx100 exclusion, and bracketing the omitted biases. The paper's heavy reliance on the STARFORGE series is normal for a companion paper, and I see no novelty or attribution concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. The paper takes STARFORGE-bound/unbound groups, runs them through a Monte Carlo Gaia-style completeness recipe, and shows that inferred masses are recovered at 20-90%, radii can be overestimated by up to 10x for low-mass groups, and virial parameters can be inflated by up to a factor of ten. It also argues that kinematic traceback ages are only good to ~20% for genuinely unbound systems (true alpha_vir > 2), and since observed virial parameters can be biased upward, you often cannot tell if a given association's traceback age is trustworthy. That is a substantive, within-subfield result, and the caution it urges is sensible.\n\nWhat it does well: the sampling procedure is clearly specified, the error curves follow from it, and the authors are upfront about the representativeness assumption and about omitted biases like reddening and radial-velocity incompleteness. The CFN comparison is honest — they match 4 of 7 groups, and they explicitly say no single simulation matches most CFN groups. The central claim, that completeness biases inflate virial parameters and undermine traceback reliability, holds up.\n\nSoft spots, in order of importance. First, the stress-test concern is correct: HDBSCAN group identification is run on the unbiased simulation data (Section 2.4), and only then are biases applied. Real Gaia associations are themselves products of clustering on biased catalogs, so the reported factors describe how a known group's properties shift, not how the population of groups recovered from a Gaia catalog is biased. A selection effect could dilute, or in principle amplify, the virial inflation for detected groups. This does not break the central argument, but it means the factor-ten number should be presented as an upper bound for a known group, not as a blanket statement about observed associations. Second, the matching criterion D <= 0.6 is arbitrary and not calibrated; it drives the CFN analogue claims. Third, Bx100 is excluded from the traceback analysis with a stated reason, but that exclusion reduces coverage of the model grid. Fourth, data and code are not released, so the numbers are not independently checkable; a robustness test re-running HDBSCAN on the biased samples would also be the natural fix.\n\nFor whom: anyone working with Gaia-based young association catalogs, virial parameters, or traceback ages. It deserves a serious referee; I would send it out. The identification issue should be addressed, ideally with a re-clustering experiment, before the quantitative bias claims are used to reinterpret observations.","headline":"A useful, mostly sound quantification of Gaia biases on young association properties, but the group-identification step is done before biases are applied, so the quoted bias factors describe known groups rather than groups recovered from a real Gaia catalog.","tokens_in":25622,"tokens_out":1433,"would_cite":true,"duration_ms":17589,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Observational biases can inflate the virial parameter of stellar associations by a factor of ten.","keywords":["STARFORGE simulations","stellar associations","virial parameter","Gaia observational biases","kinematic traceback age","Cepheus Far North","star formation","synthetic observations"],"falsifier":"Measure the actual census of missing luminous members in a Gaia-selected association such as CFN using a deep, spatially complete survey that recovers saturated and unresolved massive stars. If the true missing mass is small, the claimed factor-of-ten virial inflation would not apply to real groups; alternatively, compute traceback ages for a sample of associations with independently confirmed virial parameters above 2 and compare with photometric ages — if the errors exceed 20%, the accuracy threshold fails.","tokens_in":24603,"feed_emoji":"🌌","tokens_out":5408,"duration_ms":49128,"temperature":0.7,"pith_summary":"The paper argues that Gaia-era measurements of young stellar associations carry severe hidden biases: missing, unresolved, or saturated stars distort measured mass, size, and velocity dispersion so much that inferred virial parameters can overshoot reality by a factor of ten. Using the STARFORGE magnetohydrodynamical simulations as a stand-in for typical Milky Way clouds, the authors build synthetic observations and show that kinematic traceback ages are reliable to within 20% only when the association is genuinely unbound, with a true virial parameter above 2. Because observational biases can make a bound-looking group appear unbound, a high observed virial ratio does not guarantee that a traceback age can be trusted. The paper also finds that the Cepheus Far North association likely formed in a low-density environment similar to the simulations, and that the difference between dynamical and stellar ages does not measure the duration of the embedded phase.","feed_headline":"Gaia biases can inflate stellar-association virial ratios tenfold","feed_subtitle":"Kinematic traceback ages need truly unbound groups; observed virial parameters may be fake.","key_machinery":"The central machinery is a synthetic-observation pipeline applied to STARFORGE groups: each simulated group is placed at a chosen distance, filtered through PARSEC isochrones with Gaia magnitude limits (roughly 3 to 20.7), stripped of unresolved binaries closer than 1 arcsecond and wide binaries within 10,000 AU, hit with a 15% random source-loss rate, and resampled 100 times to build percentile error bars. The virial parameter, defined as $\\alpha_{\\rm vir} = 5 R_h \\sigma_{1d}^2/(G M_{\\rm tot})$, is the quantity that carries the argument; the pipeline shows that the correlated biases in $M_{\\rm tot}$ and $R_h$ inflate $\\alpha_{\\rm vir}$ by up to a factor of ten, while the traceback age is computed by linearly extrapolating the median mutual distance between stars back to its minimum.","core_discovery":"In the paper's own terms, the discovery is that the standard quantities used to characterize observed stellar associations — total mass, half-mass radius, and velocity dispersion — are differentially corrupted by observational incompleteness, and these corruption patterns combine to systematically inflate the virial parameter by up to an order of magnitude. Masses are always underestimated because bright massive stars saturate in Gaia and faint members fall below its detection limit; radii of small groups are overestimated because poor sampling spreads the measured distribution; velocity dispersion is the least affected. Since the virial parameter enters as $\\alpha_{\\rm vir} = 5 R_h \\sigma_{1d}^2/(G M_{\\rm tot})$, the net effect is a large upward bias. Therefore an observed association that looks unbound may actually contain loosely bound stars whose slow expansion skews traceback ages. The authors establish that for groups with true virial parameters above 2, linear traceback of stellar motions recovers the expansion age within 20%, but the observational inflation of the virial parameter makes it hard to know whether a given group actually meets that condition.","pith_inferences":["If the tenfold virial inflation holds for real Gaia-discovered associations, then many systems currently classified as unbound expanding associations may in fact retain a significant fraction of loosely bound stars, and their expansion ages derived from traceback would be systematically too young.","A natural test is to re-measure group masses with infrared or deep optical surveys that recover saturated massive stars; if recovered masses rise substantially, the virial-inflation bias is present in real data, and traceback studies should require an independent boundness check.","The omission of reddening and radial-velocity incompleteness from the bias recipe suggests the real error budget for field associations could be even larger; including them would likely strengthen the paper's qualitative conclusions rather than reverse them.","The CFN result that massive stars are not preferentially in groups may reflect genuine low-density formation rather than bias, but a deeper census of CFN's high-mass end would settle whether the contrast with simulations is real."],"forward_implications":["For groups with true mass below about 100 $M_\\odot$, recovered masses fall between 20% and 90% of the true value, and half-mass radii can be overestimated by up to a factor of ten.","Velocity dispersion is the most trustworthy observable, with relative errors below about 20% for all groups and below 5% for the most massive ones.","Kinematic traceback ages are accurate to within 20% only for associations whose true virial parameter exceeds 2; because observed virial parameters can be inflated tenfold, a high observed value is not enough to justify trusting a traceback age.","Four of the seven Cepheus Far North subgroups have analogues in the STARFORGE simulations, indicating CFN formed in a low-density, typical Galactic-cloud environment without preferential massive-star placement in groups.","There is no correlation between the stellar-dynamical age difference and the duration of the embedded phase, because star formation continues while the cloud disperses."],"supporting_citations":[{"why":"Supplies the STARFORGE simulation code and setup that generates the stellar systems analyzed here.","marker":"Grudić et al. (2021)"},{"why":"Defines the group-identification, N-body post-processing, and group catalogues that this paper extends.","marker":"Farias et al. (2023)"},{"why":"Provides the SPYGLASS young-star survey, the 15% source-loss rate, and the CFN membership basis.","marker":"Kerr et al. (2021)"},{"why":"Provides the CFN subgroup properties, the traceback-age definition, and the mass-correction recipe.","marker":"Kerr et al. (2022)"},{"why":"Is the observational traceback study whose photometric-traceback age discrepancy the paper reinterprets.","marker":"Miret-Roig et al. (2024)"},{"why":"Supplies the PARSEC isochrones used to map observable luminosity ranges at a given distance.","marker":"Bressan et al. (2012)"},{"why":"Defines Gaia magnitude limits and astrometric accuracy used in the synthetic observations.","marker":"Brown et al. (2018)"},{"why":"Defines the virial parameter whose observational inflation is the core result.","marker":"Bertoldi & McKee (1992)"}],"fun_headline_variants":["Gaia biases inflate stellar virial ratios tenfold","Observational biases can fake unbound stellar associations","Inflated virial parameters from Gaia undermine traceback ages","Stellar group dynamics skewed by Gaia's missing massive stars","Virial parameter inflation: a Gaia bias trap for stellar ages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that STARFORGE simulations represent typical Milky Way giant molecular clouds and that the synthetic-observation recipe captures the dominant Gaia biases; if either fails, the error magnitudes, the tenfold virial inflation, and the traceback accuracy threshold do not transfer to real associations.","fun_headline_variants_meta":{"raw":{"variants":["Gaia biases inflate stellar virial ratios tenfold","Observational biases can fake unbound stellar associations","Inflated virial parameters from Gaia undermine traceback ages","Stellar group dynamics skewed by Gaia's missing massive stars","Virial parameter inflation: a Gaia bias trap for stellar ages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1593,"prompt_tokens":1040,"completion_tokens":553,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":656,"tokens_out":553,"duration_ms":6223,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:09:50.123267+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the actual census of missing luminous members in a Gaia-selected association such as CFN using a deep, spatially complete survey that recovers saturated and unresolved massive stars. If the true missing mass is small, the claimed factor-of-ten virial inflation would not apply to real groups; alternatively, compute traceback ages for a sample of associations with independently confirmed virial parameters above 2 and compare with photometric ages — if the errors exceed 20%, the accuracy threshold fails.","supporting_citations":[],"review_version":1}