{"id":"86758483-d2af-49f9-834e-dad5ac2767a9","arxiv_id":"2505.10487","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A machine-learning ranking lets planetary ephemerides drop 147 of 343 Main Belt asteroids while keeping postfit Mars residuals within about 20 cm of the full 343-asteroid model and improving mass uncertainties by about 15 percent.","lead":"The authors use a boosted decision tree to rank 343 Main Belt asteroids by how strongly their masses affect Mars tracking residuals, then show that 147 can be dropped from the planetary ephemeris model without degrading the fit. The result is a faster, lower-dimensional model with about 15% tighter uncertainties on the remaining asteroid masses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No control selection is tested: the gradual degradation in Fig. 3 could reflect only the number of asteroids removed, not the BDT ordering, so the ranking claim is not yet supported.","rationale":"The reader identifies the ±10% linearized training target as the weakest assumption. I agree that extrapolating from ±10% mass changes to complete removal is a real risk, but the full-refit validation already tests the actual nonlinear removal for the one BDT sequence, so that risk is partially mitigated. The more decisive and, in my view, more directly load-bearing gap is the absence of any control ordering: the paper never compares the BDT sequence with random subsets, a mass-based ordering, or a reverse ranking. The full-refit results are genuine evidence that the specific 196-asteroid model works, and the extrapolation check is a positive feature, but neither demonstrates that the BDT ranking, rather than the number of removed objects, is responsible for the result. This concern can be settled by a concrete computational experiment with the same fitting pipeline, and the verdict should remain CONDITIONAL pending that control. If the control shows no advantage for the BDT ordering, the central claim should be downgraded to a reduced-model demonstration; if the control confirms the ordering matters, the paper's methodology claim would be substantially strengthened.","tokens_in":19625,"tokens_out":6683,"duration_ms":72222,"concrete_test":"Run the identical cumulative-removal full-fit procedure for (i) at least 10 random permutations of the 343-asteroid list and (ii) a ranking by smallest fitted INPOP25b mass, computing sigma_MEX and the residual indicator RI at each removal count up to 200. Compare the BDT RI curve to the spread of the control curves, especially at N_removed = 147. If the BDT curve lies inside the control band, or if the RI minimum at 147 is not significantly below the control minima, then the central ranking claim fails and the paper would need to be reframed as a demonstration that some reduced model works, not that the BDT selects it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest evidence for the central claim is the cumulative-removal refit curve in Sec. 2.4. This evidence, however, does not separate the effect of the BDT ordering from the effect of removing a large number of asteroids. Any monotone removal order will show residuals that increase with the number of removals, and a plateau below 20 cm for the first 200 removals shows that the removed set as a whole is collectively unimportant; it does not show that this particular 196-asteroid subset is the one selected by the ranking, nor that another ordering (random, smallest-mass-first, or a simple size cut) would not give the same or better residual indicator. The paper's statement that 'The ranking is validated by the increase of the differences to the reference solution with the increase of the ranking' (Sec. 2.4) is exactly the property that any cumulative removal order possesses. Because the training target is built with ±10% linearized MEX residuals while the actual operation is 100% mass removal, the full-refit loop is the only place where the true nonlinear response is measured, and it is run on a single ranking trajectory. Without a control or reverse-ranking experiment, the causal role of the BDT in producing INPOP25c is not established. The reduced model may be valid and useful, but the headline claim that the BDT identifies which asteroids can be omitted is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a boosted decision tree (BDT) ranking of the 343 Main Belt asteroids used as point masses in the INPOP planetary ephemerides. The BDT is trained on a synthetic dataset of ±10% random mass perturbations, with the target being the linearized change in Mars Express (MEX) residual chi-square. Starting from the INPOP25b ephemeris, the authors remove asteroids one by one in increasing BDT-importance order, performing a full planetary fit after each removal. They find that removing up to 200 asteroids keeps the postfit MEX sigma within 20 cm of the reference, and they select INPOP25c, a solution with 147 asteroids removed (196 remaining), as the minimum of a residual indicator that combines fit and extrapolation dispersions. They report that INPOP25c improves the conditioning of the fit, reduces average mass uncertainties by about 15%, and cuts integration time by about half. The fitted masses are then compared with independent mass estimates from albedo-based neural networks, the literature, and Wasserstein barycenters.","tokens_in":19777,"tokens_out":5722,"duration_ms":59524,"significance":"If the central claims hold, the work would provide a practical and inexpensive way to reduce the dimension of the main-belt model in planetary ephemerides, with concrete benefits for conditioning, mass-parameter uncertainties, and computational cost. The paper's strongest evidence is the repeated full ephemeris refit after each cumulative removal, which is a much more demanding validation than a simple proxy test. The comparison of the INPOP25c masses with independent albedo-based and literature estimates is also a valuable sanity check. However, the central methodological claim that the BDT ranking identifies which asteroids can be omitted is not yet supported, because no alternative removal order is tested; and the training/application mismatch between ±10% linearized perturbations and 100% removal leaves open the possibility that the reduced model works for reasons unrelated to the specific ranking. The absence of public code or data also limits reproducibility of the machine-learning component. The contribution is potentially significant for the ephemerides community, but it needs stronger validation before the ranking claim can be accepted.","major_comments":[{"comment":"The validation of the ranking is not controlled. The statement that 'the ranking is validated by the increase of the differences to the reference solution with the increase of the ranking' describes a property that any cumulative removal order necessarily possesses: residuals grow as more asteroids are removed. The plateau below 20 cm for up to 200 removals shows that the removed set as a whole is collectively unimportant, but it does not show that this particular 196-asteroid subset is the one identified by the BDT, nor that a random ordering, a reverse ranking, or a simple size/mass-based cut would not perform equally well. Without such control experiments, the causal role of the BDT in producing INPOP25c is not established, and the paper's central claim about the ranking is unsupported.","section":"Sec. 2.4, Fig. 3"},{"comment":"The training set is built from random mass variations of only ±10% around the INPOP21a postfit values, with a target defined by linearized residual changes, while the actual operation is complete removal of the asteroid, i.e., a 100% change. The paper states that the 'validity of such assumptions is going to be proved by the results presented in Sect. 2.4', but those results use the same MEX-based objective and the same full global chi-square for validation; they do not independently probe nonlinearities or interactions among the many simultaneously removed asteroids. This is a load-bearing gap because the reduced model's success might be insensitive to the ordering even if the BDT ranking is unreliable outside the training range. A focused nonlinear test, such as direct integration of a few removal scenarios or an ordering based on full-removal chi-square changes for a subset, would materially strengthen the claim.","section":"Sec. 2.1.3, Eqs. (3)-(5)"},{"comment":"The selection of INPOP25c as the model with 147 asteroids removed relies on the minimum of the residual indicator RI, which is computed on the same fitted quantities and the same extrapolation interval used to define 'best'. No uncertainty on the RI values, no error bars, and no stability analysis of the minimum across the training-set sizes and hyperparameters are reported. Because the same index is used both to evaluate candidates and to select the final solution, the exact number of removals may be noise-driven. The authors should report the RI values at the minimum, the sensitivity of the minimum to the chosen training set, and ideally a statement of whether solutions between, say, 100 and 200 removals are statistically indistinguishable.","section":"Sec. 2.4.1, Fig. 4"}],"minor_comments":[{"comment":"The captions of Fig. 5 and Fig. 6 refer to 'INPOP26c' while the text uses 'INPOP25c'; also Sec. 3.1 states that 196 masses are fitted by both INPOP25b and INPOP25c, whereas the Fig. 5 caption says 193 masses. Please reconcile these numbers.","section":"Fig. 5 and Fig. 6 captions; Sec. 3.1"},{"comment":"There are several typos: 'preform' should be 'perform', 'wether' should be 'whether', and the Table 1 header 'noise-to-signal ration' should be 'ratio'.","section":"Sec. 3.2.1 and Sec. 3"},{"comment":"The notation ∆^(O−C)|O is not defined; please explain the hat operator and the restriction to O (MEX) explicitly, since Eqs. (4)-(5) are central to the training-set construction.","section":"Sec. 2.1.3, Eqs. (4)-(5)"},{"comment":"The sentence 'For this study, we used the INPOP25a datasets' is ambiguous after the paper states that the ranking is implemented starting from INPOP25b; clarify whether 'datasets' refers to the observation data set or to the ephemeris version used as the starting model.","section":"Sec. 2.4"},{"comment":"The text mentions p-values for the Kolmogorov-Smirnov tests, but Fig. 7 does not show them; please report the p-values in the text or in the figure.","section":"Sec. 3.2.1"},{"comment":"The residual indicator RI combines σ_MEX,fit and σ_MEX,ext in quadrature without an explicit justification for this particular combination; please state the units and whether the two dispersions are weighted equally by construction or by choice.","section":"Sec. 2.4.1, RI definition"},{"comment":"Minor inconsistencies in formatting include 'several trainingsets' (missing space) and inconsistent capitalization of 'INPOP25C' versus 'INPOP25c' in Sec. 3.2.2.","section":"Sec. 2.3 and Sec. 3.2.2"}],"recommendation":"major_revision","confidential_remarks":"The main weakness is the absence of any control ordering in the cumulative-removal validation; this is fixable within the scope of the paper by adding random, reverse-BDT, and mass-based removal curves. I would also encourage the editor to weigh whether the methodological novelty is sufficient, since the comparison with existing selection strategies (e.g., Kuchynka & Folkner 2013) is only qualitative and no code or data are released. The manuscript would be much stronger with a short reproducibility statement or a link to the trained ranking."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a useful engineering result, not a deep scientific breakthrough. The authors train XGBoost to rank the 343 INPOP main-belt asteroids by importance derived from linearized MEX residuals, then sequentially remove asteroids and do full ephemeris refits. The resulting 196-asteroid model INPOP25c holds postfit MEX residuals within about 20 cm, cuts integration time roughly in half, and tightens mass uncertainties by about 15% because the Jacobian is better conditioned. Those numbers are credible because they come from full refits, not just from the surrogate model.\n\nWhat I like: the validation loop is honest. Each removal step is a full fit with the complete observation set, not just the ranking metric. The physical plausibility checks against albedo-based posteriors and the Kretlow/Carry mass catalogs are reasonable and show no red flags. The paper also tells you where the assumptions are weak, which I appreciate.\n\nThe soft spots, in order of weight. First, there is no control ordering. The paper interprets the slow rise of sigma_MEX over the first 200 removals as validation of the BDT ranking. But any ordering that removes unimportant asteroids first will show a plateau; you need a random or reverse-ranking trajectory to show the BDT actually identified the right set. Without that, the claim that the BDT identifies which asteroids can be omitted is not established, even though the reduced model itself clearly works. Second, the ranking is trained on plus-or-minus 10% linearized mass perturbations and applied to 100% removals. The authors say the results prove the assumption, but the refit metric overlaps with the training objective, so that proof is partly circular. Third, the choice of 147 removals is the minimum of a residual indicator with no error bars; the differences between adjacent points look small. Minor: no code, data, or hyperparameters are released, so independent reproduction is impossible.\n\nBottom line: the reduced ephemeris is a real product and worth having. The paper deserves a serious referee and probably acceptance after a control experiment. I would ask for a random-removal and maybe a mass-cut comparison, plus artifact release. For someone working on planetary ephemerides, this is a useful reference; I would cite the INPOP25c result, with a caveat about the ranking claim.","headline":"A credible reduced-asteroid ephemeris built with a BDT ranking, but the ranking's causal role is under-tested because there is no control ordering.","tokens_in":20440,"tokens_out":2265,"would_cite":true,"duration_ms":22449,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Boosted decision trees rank the 343 main-belt asteroids by effect on Mars residuals, and removing the 147 lowest-ranked ones yields a 196-asteroid model with no significant degradation of the fit.","keywords":["boosted decision trees","asteroid ranking","Main Belt asteroids","planetary ephemerides","Mars Express residuals","point-mass model","mass uncertainties","INPOP"],"falsifier":"Recompute the BDT ranking with training perturbations spanning the full removal range, for example setting individual masses to zero or using ±100% mass changes, and refit the ephemerides; if the 147-removal set no longer keeps σMEX oscillations below 20 cm, the linearized ±10% training assumption is the breaking point. A second test would be to refit with a non-MEX observable, such as InSight or Mars orbiter ranging, and check whether residuals remain within 20 cm.","tokens_in":19325,"feed_emoji":"☄️","tokens_out":6875,"duration_ms":65332,"temperature":0.7,"pith_summary":"The paper claims that a boosted decision tree can rank the 343 Main Belt asteroids currently treated as point masses by how much each one changes Mars Express ranging residuals, and that removing the 147 lowest-ranked asteroids leaves a 196-asteroid model, called INPOP25c, whose postfit residual quality is essentially unchanged. This matters because unknown asteroid masses are a main limitation on Mars orbit prediction, and cutting the number of fitted masses by more than 140 makes the least-squares fit better conditioned, reduces average mass uncertainties by about 15%, and roughly halves integration time. The ranking is trained on random mass variations of ±10% around previous postfit values using a linearized approximation of the residual change, then validated by full planetary refits in which asteroids are removed cumulatively from least to most important. The central practical claim is that a much smaller point-mass model can replace the 343-asteroid model without degrading the ephemeris.","feed_headline":"Decision-tree ranking cuts modeled asteroids from 343 to 196","feed_subtitle":"Removing 147 main-belt asteroids keeps Mars ranging residual drift under 20 cm and sharpens mass fits.","key_machinery":"The load-bearing object is a supervised ranking of asteroid masses by their marginal effect on Mars ranging residuals, learned by boosted regression trees. A gradient-boosted decision tree is trained on a dataset pairing random asteroid mass perturbations, uniform within ±10% of INPOP21a postfit masses, with the implied change in the Mars Express residual χ2, computed through the linearized residual formula $\\Delta \\widehat{(O-C)}|_O \\equiv \\frac{\\partial (O-C)|_O}{\\partial m}\\Delta m$. The tree ensemble assigns a relative importance to each of the 343 masses, and that importance order is used as a removal sequence. The validation machinery is the residual indicator $RI = \\sqrt{\\sigma^2_{\\rm MEX,fit} + \\sigma^2_{\\rm MEX,ext}}$, combining in-fit and extrapolated Mars Express residuals, whose minimum selects the 147-removal solution.","core_discovery":"The central discovery is that the relative importance of asteroid masses for the Mars Express residual fit, learned by gradient-boosted regression trees, is a reliable guide for deleting asteroids from the dynamical model altogether. Removing objects in increasing order of importance keeps the oscillation in the postfit σMEX below 20 cm up to about 200 removals, and the chosen solution INPOP25c removes 147 asteroids and keeps 196. The reduced model has a conditioning number about 70% lower, average mass uncertainties about 15% smaller, and integration time about 52% lower in user time, while the fitted masses remain consistent with an independent albedo-based posterior and with literature mass estimates. The paper presents this as evidence that the BDT ranking identifies which asteroids can be omitted from the point-mass model without significant degradation of the planetary ephemeris.","pith_inferences":["Extending beyond the tested case, the same ranking procedure could likely be applied to trans-Neptunian objects or other perturbing populations; the authors mention this as future work but do not test it, and the computational cost of building a comparable training set would need to be addressed first.","Because the tree-based importance reflects correlations among asteroids, the method could be extended to identify groups of asteroids that can be replaced by a single effective mass, potentially reducing the model further than the 196-asteroid solution.","If the ranking proves stable under larger training perturbations, the reduced model could guide which asteroid masses most need independent determination, for example by spacecraft flybys or occultations, to keep Mars ephemerides accurate."],"forward_implications":["A 196-asteroid point-mass model can replace the 343-asteroid model without significant degradation of postfit Mars Express residuals.","The conditioning number of the least-squares fit drops by about 70%, and average fitted mass uncertainties improve by about 15%.","Integration of the planetary ephemeris takes about 52% less user time and 30% less real time.","The mass estimates from the reduced model are statistically consistent with an independent posterior built from neural-network albedo predictions.","The close match between σMEX and global χ2 trends indicates that Mars Express residuals are a reliable proxy for global ephemeris quality."],"supporting_citations":[{"why":"Supplies the original list of 343 individual Main Belt asteroids whose point-mass modeling is the target of reduction.","marker":"Williams (1984)"},{"why":"Provides the INPOP21a postfit masses used as the center of the training perturbations and the baseline ephemeris construction.","marker":"Fienga et al. (2021)"},{"why":"The gradient-boosting implementation used to train the regression and obtain the variable importance ranking.","marker":"Chen & Guestrin (2016)"},{"why":"Supplies the statistical learning background for boosted decision trees and their interpretability.","marker":"Hastie et al. (2009)"},{"why":"Documents the extrapolation limitation of planetary ephemerides and the density-boundary Monte Carlo procedure used in building INPOP versions.","marker":"Fienga et al. (2019)"},{"why":"Establishes the widespread correlation among asteroid perturbations in planetary orbitography fits, motivating the selection problem.","marker":"Kuchynka & Folkner (2013)"},{"why":"Provides the neural-network albedo catalog used to build an independent asteroid mass posterior for validating INPOP25c masses.","marker":"Murray (2023)"},{"why":"Supplies an updated literature compilation of asteroid mass determinations compared against INPOP25c.","marker":"Kretlow (2020)"},{"why":"Supplies a widely used literature mass database compared against INPOP25c.","marker":"Carry (2012)"}],"fun_headline_variants":["Boosted trees rank asteroids, drop 147 without hurting Mars fit","Decision tree pruning cuts asteroid count from 343 to 196","Machine learning identifies 147 removable asteroids for ephemeris","BDT ranking trims asteroid model while keeping residuals under 20 cm","Data-driven asteroid selection shrinks model, stabilizes Mars orbit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking is learned from small, ±10% mass perturbations and a linearized (straight-line) approximation of how the residuals respond, but it is then used to justify deleting asteroids entirely, a 100% change, and the validation refits use the same Mars-Express-based objective, so the order and size of the optimal reduced set could change if the linear ordering is not valid outside its training range.","fun_headline_variants_meta":{"raw":{"variants":["Boosted trees rank asteroids, drop 147 without hurting Mars fit","Decision tree pruning cuts asteroid count from 343 to 196","Machine learning identifies 147 removable asteroids for ephemeris","BDT ranking trims asteroid model while keeping residuals under 20 cm","Data-driven asteroid selection shrinks model, stabilizes Mars orbit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00038,"raw_usage":{"total_tokens":1932,"prompt_tokens":773,"completion_tokens":1159,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":389,"completion_tokens_details":{"reasoning_tokens":1072}},"tokens_in":389,"tokens_out":1159,"duration_ms":8963,"temperature":1.0,"reasoning_tokens":1072,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:09:16.220702+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the BDT ranking with training perturbations spanning the full removal range, for example setting individual masses to zero or using ±100% mass changes, and refit the ephemerides; if the 147-removal set no longer keeps σMEX oscillations below 20 cm, the linearized ±10% training assumption is the breaking point. A second test would be to refit with a non-MEX observable, such as InSight or Mars orbiter ranging, and check whether residuals remain within 20 cm.","supporting_citations":[{"cited_title":"1984, Icarus, 57, 1, https://doi.org/10.1016/0019-1035(84)90002-2","cited_arxiv_id":null,"evidence_quote":"Supplies the original list of 343 individual Main Belt asteroids whose point-mass modeling is the target of reduction."},{"cited_title":"2021, Notes Scientifiques et Techniques de l'Institut de Mecanique Celeste, 110","cited_arxiv_id":null,"evidence_quote":"Provides the INPOP21a postfit masses used as the center of the training perturbations and the baseline ephemeris construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the widespread correlation among asteroid perturbations in planetary orbitography fits, motivating the selection problem."},{"cited_title":"2023, , 4, 90, 10.3847/PSJ/acd381","cited_arxiv_id":null,"evidence_quote":"Provides the neural-network albedo catalog used to build an independent asteroid mass posterior for validating INPOP25c masses."},{"cited_title":"2020, in European Planetary Science Congress, EPSC2020--690, 10.5194/epsc2020-690","cited_arxiv_id":null,"evidence_quote":"Supplies an updated literature compilation of asteroid mass determinations compared against INPOP25c."}],"review_version":1}