{"id":"8600b4b6-6f91-483b-aba5-eec12dc1851d","arxiv_id":"2505.17775","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Infall time estimation from projected radius and line-of-sight velocity is capped at roughly 2.5 Gyr by orbital overlap, making a simple linear partition as accurate as any method.","lead":"This paper tests five methods for estimating when galaxies fell into galaxy clusters, using the TNG300-1 cosmological simulation. It finds all methods are limited to about 2.5 billion years of accuracy because galaxies on different orbits look the same in projected measurements, and a simple linear partition works as well as the more complex recipes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed 'fundamental limit' is not tested against extra observables, and Sec. 5.2's two-estimate RMSE advantage relies on an oracle classification not demonstrated to be achievable.","rationale":"The reader's weakest_assumption was the untested population split in Sec. 5.2, and that is exactly the load-bearing weakness I identified. The Table 1 RMSEtwo calculation is internally self-consistent (the GMM fits are shown in Fig. 10), but it computes the minimum of two errors for each galaxy, which requires knowing the true infall time. No observable is used to make the earlier/later assignment, despite Sec. 5.2 asserting that physical properties might help. This is not an internal inconsistency in the simulation analysis, but it is a gap between the displayed RMSEtwo and any statement about what observers can achieve. In addition, the abstract's claim of a 'fundamental' limit would only be fundamental if no additional observable could reduce the dispersion. The paper does not attempt such an impossibility proof; it tests only R and V and mentions extra properties as a possible mitigation (Sec. 5.2, Sec. 5.3). Therefore the honest verdict is CONDITIONAL: accept the comparison of the five methods and the 2.5-2.7 Gyr dispersion as measured, but require the population-split validation (and, ideally, a test of additional observables) before the ~1.5 Gyr two-estimate claim can be presented as a practical result. I agree with the reader's assessment, and my proposed test is a concrete, feasible check using the same TNG300-1 data.","tokens_in":17451,"tokens_out":1604,"duration_ms":14959,"concrete_test":"Train a classifier on observable galaxy properties (e.g., stellar mass, SFR, color, local density) to predict which GMM component (earlier vs later infall) an observed galaxy belongs to in the linear-partition zones, using the TNG300-1 sample with true tfall withheld for validation. If the resulting RMSEtwo with predicted assignments is not below ~2.0 Gyr (or no better than the single median RMSEmedian), then the claimed 'two estimates' improvement is an oracle artifact and the abstract's 'fundamental' language should be softened to 'with only R and V'.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that orbital overlap fundamentally caps accuracy at ~2.53 Gyr is established only for estimators built from R2D and Vlos. An information-theoretic lower bound would require considering the full observable set; the paper itself gestures at 'various physical properties' (Sec. 5.2) but never tests whether such properties actually break the overlap. The key numerical proposal (two estimates, RMSEtwo ~1.0-1.5 Gyr in Table 1) is computed by assigning each galaxy to whichever of two GMM peaks is closer, without using any observable to make that choice. As the reader notes, this is an oracle bound: an observer does not know the true infall time, so the per-galaxy assignment is unattainable. Since the paper's headline improvement to ~1.5 Gyr depends entirely on this untested and possibly unattainable assignment, the 'fundamental' wording is stronger than the evidence. Some support exists: Table 1 reports interloper fractions, and Fig. 12 shows zones 7-8 are >50% interlopers, so any practical application must address contamination. The methods comparison itself is robust and the 2.5-2.7 Gyr figures are credible; the overreach is in calling the limit fundamental conditional only on (R,V).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses the TNG300-1 simulation to construct projected phase-space (R–V) diagrams for 136 massive clusters (408 line-of-sight projections) and systematically compares five methods for estimating galaxy infall times: projected radius, caustic distance, a newly proposed linear partition, the R17 five-zone scheme, and the P19 eight-zone scheme. All methods are evaluated on the same simulated galaxy sample with a common infall-time definition (first crossing of 3 R200). The authors report that all methods achieve similar root-mean-square errors (RMSE ≈ 2.5–2.7 Gyr), close to the intrinsic pixel-level dispersion of 2.53 Gyr, and that the simple linear partition performs marginally best. They attribute this irreducible-looking scatter to orbital overlap: galaxies on different orbital phases with different infall times occupy the same regions of the R–V diagram. Two possible improvements are explored: restricting analysis to dynamically relaxed clusters (gain of only ~0.1 Gyr) and replacing a single median infall time per zone with two Gaussian-mixture estimates per zone, which lowers the per-zone RMSE to about 1.0–1.5 Gyr under the assumption that galaxies can be reliably classified into earlier- and later-infall populations. The paper concludes that orbital overlap fundamentally limits infall-time accuracy and recommends linear partitioning as a simple and robust empirical estimator.","tokens_in":17882,"tokens_out":4277,"duration_ms":36575,"significance":"If the central dispersion analysis is correct, the paper provides a useful quantitative benchmark for the many observational studies that adopt R–V diagram infall-time estimators: it shows that differences between existing methods are negligible compared with the intrinsic scatter and that a simple linear partition captures the trend as well as more complex recipes. The comparison is performed on a well-known public simulation with a clear and consistent normalization, and the authors are appropriately cautious about the small differences between methods (≈0.1 Gyr). The two-estimate idea is a constructive suggestion, but the headline improvement to ≲1.5 Gyr currently rests on an untested oracle classification, and the word 'fundamentally' overstates the scope of the result, which is conditional on using only projected radius and line-of-sight velocity. With revision, the paper can serve as a valuable reference for the infall-time estimation literature.","major_comments":[{"comment":"The RMSEtwo values in Table 1 are computed by comparing each galaxy's true infall time with both GMM peak estimates and retaining the smaller error. No observable property is used to choose between t1 and t2; the calculation assumes the authors can 'reliably divide galaxies into earlier-infall and later-infall populations.' This is an oracle bound, not a demonstrated observational deliverable. The improvement from RMSEmedian ≈ 2.2–2.7 Gyr to RMSEtwo ≈ 0.8–1.5 Gyr therefore depends entirely on an untested classification step. To make the claim credible, the authors should either (a) demonstrate a practical classification using the galaxy properties they mention (color, star formation rate, gas fraction) and recompute the realized RMSE, or (b) explicitly and consistently label the two-estimate numbers as an idealized upper bound on achievable accuracy.","section":"Sec. 5.2, Table 1"},{"comment":"The statement that 'orbital overlap fundamentally limits the accuracy of infall time estimation' is established only for estimators built from R2D and Vlos. The paper itself suggests that additional physical properties of galaxies may help partially mitigate the degeneracy (Sec. 5.2; last paragraph of Sec. 5.3), yet those properties are never tested. The R–V diagram is a projection of a higher-dimensional space; a lower bound derived from it does not imply a fundamental limit on the full observable set. The 'fundamental' language in the abstract and conclusions should be softened to something like 'currently limits the accuracy of R–V-based estimators,' or the authors should attempt a concrete test of whether adding, e.g., color or star formation rate reduces the dispersion below 2.53 Gyr.","section":"Abstract and Sec. 5.3, 6"},{"comment":"The linear partition slope k is optimized on the full TNG sample by maximizing the Spearman correlation between tfall and dlinear on that same sample, and the RMSE for the linear partition is then evaluated on the same data. This in-sample fit gives the linear partition a potential advantage over methods whose boundaries were fixed externally (R17, P19) and over the projected-radius and caustic methods that do not use fitted parameters. Although the reported accuracy differences are small (~0.1 Gyr), the 'slightly outperforms' conclusion should be verified out-of-sample, for example by dividing the cluster sample into training and test sets or by using a nested cross-validation to select the slope.","section":"Sec. 4.1, Figs. 4–5"}],"minor_comments":[{"comment":"The paragraph beginning 'Additionally, the fractions of galaxies covered by each method...' is repeated almost verbatim in the text after Fig. 6; one of the two copies should be removed.","section":"Sec. 4.2, Fig. 6"},{"comment":"The definition Vlos = |vi + H0 × ri| uses an absolute value, so the resulting R–V diagram is folded. The text later says 'we use only the absolute value |V|', but it would be clearer to state the folding convention directly at Eq. (2), since the absolute value is not the standard definition of a line-of-sight velocity.","section":"Sec. 2.4, Eq. (2)"},{"comment":"The virial mass Mvir = 3R200 σ_los^2 / G assumes a particular form of the virial theorem and a spherical mass distribution; a brief justification of why Mvir/M200 is a suitable dynamical-state indicator would be helpful, especially since the paper finds only a weak dependence on it.","section":"Sec. 5.1, Eq. (7)"},{"comment":"The text notes that a two-component GMM does not capture all multimodality (e.g., 'at least three peaks are visible in Zone 5'). A short statement about sensitivity to the number of components and to the GMM initialization would make the two-estimate table more reproducible.","section":"Sec. 5.2"},{"comment":"The interloper analysis is well placed and appropriately warns about zones 7–8, but the paper does not quantify how interloper contamination affects the RMSE of the other methods, which use different zone boundaries. A sentence acknowledging that the interloper impact is likely zone-dependent would be useful.","section":"Sec. 5.5, Fig. 12"}],"recommendation":"major_revision","confidential_remarks":"The systematic method comparison and the intrinsic-dispersion measurement are solid and should be of genuine interest to the A&A readership. The main concerns are the oracle nature of the two-estimate result and the overstatement implied by 'fundamentally limits'. If the authors can either validate a practical population classification or clearly reframe the two-estimate RMSE as an idealized bound, and qualify the 'fundamental' claim to the R–V observables actually tested, the paper could be suitable for publication. The in-sample slope optimization is a smaller issue but should be addressed with an out-of-sample check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful methods-comparison paper. It puts four published infall-time estimators plus a simple linear partition on the exact same TNG300-1 R-V diagrams and shows they all land within ~0.1 Gyr of each other, with the linear partition at least as good as the fancier recipes. The ~2.5 Gyr floor from orbital overlap is real and well illustrated. The main overreach is the two-estimate route: the <1.5 Gyr numbers assume you already know which orbital population a galaxy belongs to, which is exactly the information an R-V diagram does not give you.\n\nWhat's new: the same-dataset comparison is a first, and the orbital-overlap decomposition into six orbital phases, each with per-population scatter ~1-1.5 Gyr, makes a clean case that the scatter is mostly mixing of phases, not method failure. The linear partition is a nice addition—simple, robust, and competitive. Numbers in Table 1 are a practical resource for observers.\n\nWhere it's soft: (1) 'Fundamentally limits' overreaches; the paper only tests R2D and Vlos. The authors themselves suggest color/SFR/gas fraction in Sec. 5.2 but never test them. So the floor is a floor for (R,V) estimators, not a fundamental limit. (2) The RMSEtwo <1.5 Gyr result is an oracle bound, not an achievable method, and the text says 'assuming that we can reliably divide galaxies into earlier-infall and later-infall populations.' That assumption is untested; until someone shows a physical property that breaks the degeneracy, quote it as an upper bound on what population-split methods might achieve. (3) The linear partition slope k is fit in-sample. The difference is small and the slope is stable across mass subsamples, but the RMSE is still mildly optimistic. (4) Interlopers: zones 7-8 are >50% interlopers; the authors note this, but it limits the practical reach of the tabulated values.\n\nVerdict: the central comparison and the 2.53 Gyr dispersion measurement hold up. The paper deserves a serious referee; the main revision should be softening the 'fundamental' language and clearly flagging the two-estimate numbers as an oracle test. I'd use the linear partition values in future work and cite this as the reference benchmark for the accuracy floor.","headline":"A clean methods comparison shows all R-V infall-time estimators are capped near 2.5-2.6 Gyr by orbital overlap, but the paper's <1.5 Gyr two-estimate improvement and 'fundamental' wording rest on an untested oracle split.","tokens_in":18216,"tokens_out":3763,"would_cite":true,"duration_ms":27334,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Orbital overlap caps galaxy infall-time accuracy at about 2.6 Gyr.","keywords":["galaxies: clusters: general","galaxies: evolution","galaxy infall time","R-V diagram","projected phase space","orbital overlap","caustic profiles","cluster dynamics"],"falsifier":"Run a simulation with known infall times and test whether adding a measured galaxy property such as colour, star formation rate, or gas fraction to the R-V position lowers the estimation RMSE below 2.5 Gyr; if any property does, the claimed orbital-overlap ceiling is not fundamental.","tokens_in":17258,"feed_emoji":"🔭","tokens_out":10756,"duration_ms":77868,"temperature":0.7,"pith_summary":"Astronomers estimate when a galaxy first fell into a cluster from its position on the R-V diagram, the plot of projected cluster-centric radius against line-of-sight velocity. This paper compares five recipes for doing that on the same simulated galaxy sample and asks how accurate any of them can be. It concludes that all methods are essentially equivalent, with errors set by the intrinsic 2.53 Gyr dispersion of infall times in the diagram, so estimating a single infall time from R and V alone is capped near 2.6 Gyr. The paper attributes this ceiling to orbital overlap: galaxies in different orbital phases occupy the same projected positions and have very different infall histories. It then shows that two estimates per zone instead of one median lowers the scatter to about 1.5 Gyr, but only under the assumption that observers can tell the two populations apart.","feed_headline":"Galaxy infall times can't be pinned below 2.6 Gyr","feed_subtitle":"Five methods tie near 2.6 Gyr; two estimates may reach 1.5 Gyr if early and late infallers can be told apart.","key_machinery":"The R-V diagram, with $R=R_{2D}/R_{200}$ and $V=V_{\\rm los}/\\sigma_{\\rm los}$, is the observational projection of the infall problem. The evaluation is carried by the RMSE of true infall time around the median in each diagram pixel or zone, which quantifies the fundamental scatter. The paper introduces a linear distance $d_{\\rm linear}=(|V|-kR)/\\sqrt{1+k^2}$ from the origin to an oblique partition line, calibrates the optimal slope $k=-3.7$ in 2D via the Spearman rank correlation, and uses it as the reference estimator. The orbital-overlap argument rests on a two-component Gaussian mixture fit per zone for the two-estimate improvement, and on a six-component fit of the global infall-time distribution compared with six orbital-phase populations defined by pericentre and apocentre counts.","core_discovery":"The paper's central claim is that the infall time of a galaxy cannot be recovered from the R-V diagram more accurately than the diagram's intrinsic dispersion, measured here as 2.53 Gyr, and that no partition of this diagram decisively beats any other: a simple linear cut has RMSE 2.56 Gyr, while caustic, radius, and the two published zone recipes all fall within roughly 0.1 Gyr of it. The reason is orbital overlap, demonstrated by decomposing the full infall-time distribution into six Gaussian components that align with six orbital-phase populations; each population has its own scatter below about 1.5 Gyr, but the populations occupy the same projected positions, so a galaxy's location in the R-V diagram is a mixture of very different infall histories. Two improvement routes are established: selecting dynamically relaxed clusters lowers the dispersion by about 0.1 Gyr, and switching from one median estimate to two estimates per zone reduces the scatter to about 1.5 Gyr or less, assuming the two populations can be identified. The paper also notes that outermost zones are dominated by interlopers and should be interpreted cautiously.","pith_inferences":["A testable next step is to check whether any single galaxy property, such as colour, star formation rate, or gas fraction, separates the two Gaussian populations per zone; if it does, the two-estimate accuracies become achievable with existing cluster surveys.","The paper's use of 'fundamental' applies to estimators built only from R and V; non-kinematic observables are an open route the paper itself suggests but does not test.","A cleaner causal test of the orbital-overlap explanation would define orbital populations purely from dynamical histories in the simulation, then measure how separable those populations are in observable space, rather than choosing population boundaries to match a Gaussian mixture.","If the orbital-overlap ceiling holds, infall-time analyses in galaxy-evolution studies should shift from point estimates to explicitly reporting the bimodal distribution of possible infall times."],"forward_implications":["Any recipe restricted to projected radius and line-of-sight velocity inherits the same roughly 2.5 Gyr intrinsic scatter, so the choice among partitions is secondary.","A simple linear cut through the R-V diagram is the cheapest reliable estimator, matching or slightly beating the more elaborate zone recipes.","Restricting samples to dynamically relaxed clusters reduces the typical error by only about 0.1 Gyr, so cluster selection is not a strong lever.","Reporting two plausible infall times per galaxy rather than one can reduce per-zone scatter to about 1.5 Gyr or less, if the earlier-infall and later-infall populations can be separated.","Outer zones of the R-V diagram are more than half interlopers, so infall-time estimates in the outermost parts of the diagram should be treated as unreliable."],"supporting_citations":[{"why":"Introduced orbital-library infall-time probability densities in the R-V diagram and reported a 2.58 Gyr accuracy figure that this study re-evaluates against a common dataset.","marker":"Oman et al. (2013)"},{"why":"Supplied the five-zone R17 discretization compared here, built to maximize population purity per zone.","marker":"Rhee et al. (2017)"},{"why":"Supplied the eight-zone quadratic P19 discretization compared here, the current standard for zone-based infall-time estimation.","marker":"Pasquali et al. (2019)"},{"why":"Popularized caustic profiles for separating infall stages in the R-V diagram; the caustic tracer evaluated here follows this line.","marker":"Noble et al. (2013)"},{"why":"Established the caustic technique in cluster phase space that underlies the caustic method.","marker":"Diaferio & Geller (1997)"},{"why":"Derived the near-hyperbolic caustic shape that motivates the caustic tracer's form.","marker":"Regos & Geller (1989)"},{"why":"Produced the simulation suite from which the 136-cluster and 10,087-galaxy samples are drawn.","marker":"Springel et al. (2018)"},{"why":"Built the SubLink merger trees used to trace galaxy histories and define first-infall times.","marker":"Rodriguez-Gomez et al. (2015)"}],"fun_headline_variants":["Infall time accuracy floored at 2.6 Gyr by orbital overlap","Two infall estimates shrink error to 1.5 Gyr: study","Simple linear cut rivals complex recipes for galaxy infall times","Galaxy infall times: all methods within 0.1 Gyr, no winner","Orbital overlap sets 2.6 Gyr limit on infall time estimates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The only route the paper offers below 2.6 Gyr assumes observers can reliably classify galaxies as early-infall or late-infall from their physical properties, a classification the paper assumes but never tests.","fun_headline_variants_meta":{"raw":{"variants":["Infall time accuracy floored at 2.6 Gyr by orbital overlap","Two infall estimates shrink error to 1.5 Gyr: study","Simple linear cut rivals complex recipes for galaxy infall times","Galaxy infall times: all methods within 0.1 Gyr, no winner","Orbital overlap sets 2.6 Gyr limit on infall time estimates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1693,"prompt_tokens":1104,"completion_tokens":589,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":486}},"tokens_in":720,"tokens_out":589,"duration_ms":5683,"temperature":1.0,"reasoning_tokens":486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:40:17.618802+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a simulation with known infall times and test whether adding a measured galaxy property such as colour, star formation rate, or gas fraction to the R-V position lowers the estimation RMSE below 2.5 Gyr; if any property does, the claimed orbital-overlap ceiling is not fundamental.","supporting_citations":[{"cited_title":"A., Hudson , M","cited_arxiv_id":null,"evidence_quote":"Introduced orbital-library infall-time probability densities in the R-V diagram and reported a 2.58 Gyr accuracy figure that this study re-evaluates against a common dataset."},{"cited_title":"2019, Monthly Notices of the Royal Astronomical Society, 484, 1702, aDS Bibcode: 2019MNRAS.484.1702P","cited_arxiv_id":null,"evidence_quote":"Supplied the eight-zone quadratic P19 discretization compared here, the current standard for zone-based infall-time estimation."},{"cited_title":"G., Webb , T","cited_arxiv_id":null,"evidence_quote":"Popularized caustic profiles for separating infall stages in the R-V diagram; the caustic tracer evaluated here follows this line."},{"cited_title":"& Geller , M","cited_arxiv_id":null,"evidence_quote":"Established the caustic technique in cluster phase space that underlies the caustic method."},{"cited_title":"& Geller , M","cited_arxiv_id":null,"evidence_quote":"Derived the near-hyperbolic caustic shape that motivates the caustic tracer's form."},{"cited_title":"2015, , 449, 49","cited_arxiv_id":null,"evidence_quote":"Built the SubLink merger trees used to trace galaxy histories and define first-infall times."}],"review_version":1}