{"id":"38aa493c-2695-464c-8ed8-61a7bfab4c14","arxiv_id":"2412.02785","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Both the equal-mass-embryo scenario and the high-mass-ratio impact scenario form Uranus and Neptune analogues with comparable low probability (~0.1-1%), so neither is dynamically preferred.","lead":"This paper runs 14,000 computer simulations of the final stages of Uranus and Neptune formation, comparing two giant impact scenarios: collisions between roughly equal-mass embryos and collisions where a massive proto-planet is struck by a small embryo. It finds both scenarios have similarly low success rates, around 0.1 to 1 percent, when matching the observed masses, rotation periods, and orbital architecture.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3's success rates are derived from 1–15 events per 1000 simulations, so the central claim that HMR and I15 scenarios are 'within a factor of ~2' is not statistically supported.","rationale":"The central claim is a quantitative comparison of success probabilities. The most load-bearing condition for that claim is that the reported percentages accurately reflect the underlying probabilities to within a factor of ~2. With 1000 simulations per set and success rates of 0.02–1.5%, the expected event counts are 0.2–15, so the sampling uncertainty is large (Poisson coefficient of variation of 25–200% for the lower counts). The reader's weakest_assumption concerned the perfect-merging rotation model; the authors actually test that by re-scaling grazing-impact rotation periods (Section 5) and report that conclusions are qualitatively unchanged. That reduces the centrality of the rotation-model concern. The small-number statistics, however, are not addressed anywhere and directly undermine the factor-of-2 phrasing. This is an internal precision issue rather than a disagreement with external consensus. If the confidence intervals overlap for all rows, the paper's main qualitative takeaway (both scenarios are similarly unlikely to succeed, at the ~0.1–1% level) survives, but the precise 'within a factor of ~2' should be softened. The reader already recommended adding error bars, so the CONDITIONAL verdict stands; my concern reinforces that condition.","tokens_in":18029,"tokens_out":8301,"duration_ms":79283,"concrete_test":"Compute exact binomial (Clopper–Pearson) 95% confidence intervals for each percentage in Table 3 using N=1000. For each tolerance column, compare the intervals for the HMR rows (Msmall=0.5, 1, 2, 3) and the I15 rows (6, 4–8). If the intervals overlap for every column and the ratio of the largest plausible ratio (upper bound of the higher row divided by lower bound of the lower row) exceeds 2, then the 'within a factor of ~2' claim is not supported and should be reworded to 'the same order of magnitude'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 3 (Section 4) reports success rates of 0.02%–1.54% per set, each based on 1000 simulations. These percentages correspond to only 0.2–15.4 successful simulations per set. For example, the ±15% column gives 0.42% (≈4 events) for the Msmall=0.5 M⊕ HMR set and 0.16% (≈1.6 events) for the 6 M⊕ I15 set. The point-estimate ratio is 2.6, but the 95% Poisson confidence intervals are approximately 0.11%–1.07% and 0.02%–0.58% respectively. These intervals overlap substantially and permit true ratios ranging from well below 1 to over 10. The same issue affects every row: with counts this small, the data cannot support the 'within a factor of ~2' wording in the Abstract and Section 4. The paper provides no confidence intervals, jackknife, or bootstrap, and does not discuss this small-number limitation. The qualitative conclusion that both scenarios are rare (order 0.1–1%) is likely robust, but the quantitative factor-of-2 comparison is not.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether the formation of Uranus and Neptune through giant impacts with high mass-ratio impactors (small embryos of 0.5–3 M⊕ hitting ~13–17 M⊕ protoplanets) is dynamically as likely as the previously studied equal-mass-ratio embryo scenario (I15). The authors run 12 sets of 1000 N-body simulations including gas disk migration and tidal damping, classify outcomes, compute rotation periods of collision products via a two-body angular momentum approximation, and then compare the fraction of simulations that produce both planets with plausible masses and rotation periods. They conclude that the two scenarios have broadly similar success rates, within a factor of ~2, with overall probabilities of order 0.1–1%. The paper also tests a simplified correction for grazing collisions and a scenario with more than two initial protoplanets.","tokens_in":18281,"tokens_out":5361,"duration_ms":54143,"significance":"If the central quantitative claim were fully supported, this paper would be a valuable contribution to the ice-giant formation literature: it is the first to test the high-mass-ratio impact scenario, motivated by SPH simulations, within a planet-formation N-body framework rather than assuming the impact happens. The authors provide a systematic parameter study with 1000 simulations per set, transparently describe their initial conditions and acknowledge their ad-hoc nature, and test the robustness of their rotation-period results to a simplified grazing-impact correction. The qualitative conclusion that both scenarios are rare (order 0.1–1%) appears robust and is itself a useful result. However, the more specific quantitative comparison between scenarios is not yet supported by the statistics reported, and the mass-matching component of the success criterion is partly imposed by the initial conditions rather than predicted.","major_comments":[{"comment":"The central quantitative claim in the Abstract and Section 4 that the two scenarios have success rates \"within a factor of ~2\" is not supported by the reported statistics, because each percentage in Table 3 corresponds to only 0.2–15.4 successful simulations out of 1000. For example, the ±15% column gives 0.42% (≈4 events) for the Msmall = 0.5 M⊕ HMR set and 0.16% (≈1.6 events) for the I15 6 M⊕ set; the 95% Poisson confidence intervals for these counts are approximately 0.11%–1.07% and 0.02%–0.58%, respectively, which overlap substantially and permit true ratios ranging from well below 1 to over 10. The paper should report confidence intervals for every success rate and either soften the factor-of-2 wording to a qualitative statement or demonstrate that the ratio is robust to counting uncertainty.","section":"Section 4, Table 3"},{"comment":"The statement in the Abstract that the simulations broadly match the masses, mass ratio, and rotation periods is partly by construction for the HMR scenario. The protoplanet initial masses are chosen (Table 1) so that a single collision with an embryo of mass Msmall yields exactly the present-day masses of Uranus and Neptune, and the success criterion in Section 4 for HMR runs only requires that both protoplanets experience at least one collision with a small embryo; final masses are not checked. For the I15 runs, a mass threshold larger than 12 M⊕ is imposed, so the two scenarios are not evaluated with the same mass-selection criterion. The authors should either verify final masses in the reported statistics or explicitly state that mass matching is imposed rather than predicted.","section":"Section 2 and Section 4"},{"comment":"The sensitivity of the central comparison to the perfect-merging assumption is underreported. The authors state that when grazing collisions (impact parameter greater than 0.8, about 30% of collisions) are assigned rotation periods a factor of 2 longer, the probabilities of Table 3 change \"by up to a factor of a few,\" which is larger than the claimed factor-of-2 difference between scenarios. Because this uncertainty is comparable to or larger than the effect being measured, the recalculated probabilities should be presented explicitly, for example as an additional table or column, so that the reader can judge whether the factor-of-2 claim survives. The current qualitative statement that the conclusions remain broadly unchanged is not sufficient for the quantitative claim in the Abstract.","section":"Section 5"}],"minor_comments":[{"comment":"There are several typographical errors: \"1 Merath\" in the Abstract should be \"1 M⊕\", and \"Planetary dinamics\" in the keywords should be \"Planetary dynamics\".","section":"Abstract and Keywords"},{"comment":"When comparing to the original I15 scenario, the paper should state more prominently that the I15-like runs in this work place Jupiter at 5.2 au, whereas most original I15 simulations placed Jupiter at 3.5 au; this difference may affect the collision statistics and thus the direct comparison to the earlier work.","section":"Section 2"},{"comment":"The x-axis label says \"see 1,\" which should read \"see Table 1.\"","section":"Figure 2"},{"comment":"The five-giant-planet scenario is described only qualitatively, stating that success fractions are \"very similar\" to those in Table 3, but no numbers are provided; a table or a quantitative sentence would allow the reader to assess this claim.","section":"Section 4.1"},{"comment":"In the sentence \"this scenario is refereed to as the high-mass-ratio scenario,\" \"refereed\" should be \"referred.\"","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper's main quantitative conclusion is not yet supported by the reported counting statistics, and the mass-matching component of the success criterion is partly an artifact of the initial conditions. Both issues are fixable with additional analysis, so I recommend major revision rather than rejection. The qualitative conclusion that both scenarios are rare is likely robust."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe new thing here is the first N-body formation test of the high-mass-ratio giant impact scenario, which previously lived only in SPH papers. The authors run 1000 simulations per set across 12 sets, with type-I migration and gas damping, and compare against the equal-mass I15 scenario. That is useful and long overdue, and the paper is refreshingly honest about its own limitations.\n\nThe main qualitative result—both scenarios produce Uranus/Neptune analogues with the right masses, mass ratio, and rotation periods only at the 0.1–1% level—is probably robust. The rotation period distributions behave as expected: low-mass impactors give slower spins, equal-mass impacts give too-fast spins. The authors also test sensitivity to grazing collisions and find their qualitative conclusions survive.\n\nWhere I part company is the \"within a factor of ~2\" claim. Table 3 success rates come from a handful of events per 1000 simulations. The ±15% column gives 0.42% (≈4 events) for the best HMR set and 0.16% (≈1.6 events) for the I15 6 M⊕ set. Poisson uncertainties on those numbers are huge and overlapping; the true ratio could be below 1 or above 10. Every row has the same problem. So the data cannot support the factor-of-2 language in the Abstract and Section 4. The qualitative statement that both scenarios are comparably rare is fine; the quantitative comparison is not.\n\nOther soft spots are the ad hoc initial conditions (acknowledged) and the perfect-merging collision treatment. The latter is partially mitigated by their Section 5 recalculation with grazing-impact corrections, but those corrections change the probabilities by \"up to a factor of a few\"—which reinforces how fragile the factor-of-2 comparison is. The paper also ships no code or data, only \"on reasonable request,\" which makes independent verification harder.\n\nThe citation pattern is fine; they build on I15 and the SPH literature appropriately. No missing references jumped out.\n\nWho is this for? Ice giant formation modellers and people interpreting SPH impact results in a dynamical context. It deserves a serious referee, but with a clear instruction: add confidence intervals or bootstraps on Table 3 and soften the factor-of-2 wording. The qualitative result is worth publishing; the statistical overreach is not.\n\nRecommendation: send to peer review with required revisions. The work is honest, useful, and the central qualitative claim is solid.","headline":"Useful first N-body test of the high-mass-ratio impact scenario, but the 'within a factor of ~2' comparison is not supported by the small-number statistics.","tokens_in":18803,"tokens_out":3107,"would_cite":true,"duration_ms":30712,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the two main giant-impact formation channels for Uranus and Neptune—collisions among roughly equal-mass embryos versus collisions with much smaller impactors—match the observed masses, obliquities, and rotation…","keywords":["Planetary formation","Giant impacts","Uranus","Neptune","Obliquity","Rotation period","N-body simulations","Type-I migration"],"falsifier":"Run smoothed-particle hydrodynamics impact simulations of the actual collision geometries recorded in the both-collide outcomes from each scenario, including fragmentation and hit-and-run physics, and compare the resulting rotation periods with the Maclaurin-spheroid estimates used here; if the two scenarios diverge by more than the assumed factor-of-2 grazing correction, the paper's equal-success-rate conclusion fails.","tokens_in":17789,"feed_emoji":"🪐","tokens_out":14513,"duration_ms":158287,"temperature":0.7,"pith_summary":"This paper tries to decide whether Uranus and Neptune were more likely finished by impacts between bodies of roughly equal mass or by a large proto-planet absorbing a much smaller embryo. The authors run thousands of N-body simulations of migrating protoplanets in a gas disk and compare how often each scenario produces planets matching Uranus and Neptune's masses, mass ratio, obliquities, and rotation periods. They find that the two scenarios succeed at nearly the same low rate, with overall success probabilities of order 0.1–1% and within a factor of about 2 of each other. The reason for the tie is complementary: equal-mass collisions happen often but leave planets spinning too fast, while high-mass-ratio collisions give the right spins but are rare because small embryos are scattered rather than accreted. If correct, this means the formation of the ice giants was a lucky outcome of chaotic accretion, and formation models alone cannot tell the two impact channels apart.","feed_headline":"Rival impact scenarios tie for explaining Uranus and Neptune","feed_subtitle":"Equal-mass and high-mass-ratio collisions both match the ice giants' masses, tilts, and spins at roughly the same low odds.","key_machinery":"Two pieces of machinery carry the argument. First, the N-body model adds type-I migration, eccentricity damping, and inclination damping from a one-dimensional gas disk, and this determines whether small embryos are scattered away or actually collide with the proto-Uranus and proto-Neptune; this is why the two scenarios have different collision frequencies. Second, the rotation period of every post-impact planet is computed in post-processing from angular momentum conservation using a Maclaurin spheroid moment of inertia, a simplified treatment that the paper validates against smooth-particle hydrodynamics simulations. The paper's central product is the joint statistic: the fraction of simulations in which both proto-planets survive, each experience at least one giant impact, the outer Solar System architecture is preserved, and the final rotation periods fall within a chosen fraction of the observed averaged value. That joint fraction is the quantity that comes out nearly equal, within a factor of about 2, for the two competing scenarios.","core_discovery":"On the paper's own terms: in N-body simulations of the final accretion of Uranus and Neptune, the high-mass-ratio scenario—a proto-Uranus and proto-Neptune of roughly 13 and 16 Earth masses, each absorbing a single small embryo of 0.5–3 Earth masses—produces rotation periods that peak near the observed roughly 16.5-hour values, but the required collisions occur in only about 1.16% of simulations. The equal-mass-ratio scenario produces collisions about ten times more often, about 10.25%, but the resulting planets generally spin too fast. Combining collision frequency with rotation-period agreement, the paper reports overall success rates of roughly 0.1–1% for each scenario, differing by no more than a factor of about 2 across rotation-period tolerance intervals of ±15% to ±100%. The paper concludes that, from a planet-formation perspective, there is no clear statistical preference for either giant-impact channel, and that reproducing Uranus and Neptune is a fortuitous outcome of the chaotic accretion phase.","pith_inferences":[],"forward_implications":["If the claim is right, the high-mass-ratio scenario's better spin match is offset by its rarer collisions, so neither the equal-mass nor the high-mass-ratio scenario can be rejected on dynamical grounds alone.","The low absolute success rates, of order 0.1–1%, imply that forming Uranus and Neptune is a rare, chance outcome of the giant-impact phase rather than a generic pathway of gas-disk accretion.","If some unknown mechanism removes angular momentum after accretion, such as interaction with the surrounding gas disk, the equal-mass scenario would become strongly preferred, with its success rate rising by up to an order of magnitude.","Starting with three or four protoplanets instead of two does not improve the success rate, so a five-giant-planet initial configuration offers no clear advantage for matching Uranus and Neptune.","Because roughly 30% of collisions have impact parameters above 0.8, fragmentation at those grazing impacts could lengthen rotation periods by up to a factor of 2 and shift the quantitative success rates by a factor of a few, though the authors argue the qualitative tie remains.","One concrete test: take the actual impact geometries recorded in the both-collide simulations from each scenario and run them in a smoothed-particle hydrodynamics impact code that allows fragmentation and hit-and-run. If the resulting spin periods differ systematically from the Maclaurin-spheroid estimate by more than the assumed factor for grazing impacts, the paper's equal-success-rate conclusio","The paper leaves untested whether varying the gas disk's lifetime or surface density would break the factor-of-2 tie and also whether the leftover embryos and co-orbital objects seen in many HMR simulations would survive the later Solar System instability, potentially adding an independent observational constraint on the scenario.","The simplified perfect-merging assumption may not bias the comparison symmetrically: fragmentation at high impact parameters tends to remove angular momentum from near-equal-mass collisions more effectively, so a fuller hydrodynamical treatment could widen, rather than narrow, the gap between the two scenarios."],"supporting_citations":[],"fun_headline_variants":["Ice giant spins: rival impact scenarios tie in odds","Uranus and Neptune: two impact models, similar low probability","Giant impact simulations fail to favor either ice giant origin","Equal-mass or lopsided impacts: same chance for Uranus and Neptune","Ice giant formation: both giant impact paths equally improbable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Final rotation periods are computed assuming that collisions merge perfectly, with only a simplified correction for grazing impacts; if a substantial fraction of collisions instead fragment or involve hit-and-run, as suggested by the authors' own estimate of roughly 30% high-impact-parameter events, the resulting spins could shift by up to a factor of about 2 and change the success rates in Table 3.","fun_headline_variants_meta":{"raw":{"variants":["Ice giant spins: rival impact scenarios tie in odds","Uranus and Neptune: two impact models, similar low probability","Giant impact simulations fail to favor either ice giant origin","Equal-mass or lopsided impacts: same chance for Uranus and Neptune","Ice giant formation: both giant impact paths equally improbable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000246,"raw_usage":{"total_tokens":1606,"prompt_tokens":1082,"completion_tokens":524,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":454}},"tokens_in":698,"tokens_out":524,"duration_ms":5382,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:06:33.718037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run smoothed-particle hydrodynamics impact simulations of the actual collision geometries recorded in the both-collide outcomes from each scenario, including fragmentation and hit-and-run physics, and compare the resulting rotation periods with the Maclaurin-spheroid estimates used here; if the two scenarios diverge by more than the assumed factor-of-2 grazing correction, the paper's equal-success-rate conclusion fails.","supporting_citations":[],"review_version":1}