{"id":"c70e623f-ce58-442f-8ffe-ca4f4f49e8ad","arxiv_id":"2607.03635","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Comparing seven impact-atmosphere models via Monte Carlo bombardment shows 2–3 order-of-magnitude spreads and net atmospheric growth for Venus, Earth, and Mars.","lead":"Monte Carlo runs of seven impact-atmosphere models (plus composites) on Venus, Earth, and Mars show final pressures that can differ by two to three orders of magnitude, yet most produce net growth of +0.01 to +100 bar. The spread implies single-model forecasts are unreliable and that early bombardment may have been a major volatile source.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"The net-growth claim rests on Appendix-B patches that systematically suppress non-physical losses; without them several models become unusable or reverse sign, so the qualitative conclusion is not independent of the interventions.","rationale":"The reader’s weakest-assumption diagnosis is exactly the load-bearing concern: the comparability and physical meaning of the patched models. The paper is transparent about the interventions and supplies open code, so the claim remains inspectable; the appropriate verdict is therefore still CONDITIONAL, with the same high confidence once the caveats are kept front-and-center. No stronger objection (internal inconsistency, missing data, etc.) is required; the patches themselves are the soft spot that must be stress-tested by the concrete re-run above.","tokens_in":31562,"tokens_out":520,"duration_ms":5085,"concrete_test":"Re-run the full 30-member Monte-Carlo suite for present-day Earth and Mars with every Appendix-B zeroing/capping rule disabled (or replaced by a hard NaN abort). If more than half of the individual-model trajectories reverse from net growth to net loss, or if the composite can no longer be evaluated for >10 % of impactors, the qualitative claim that “most models result in net growth” does not survive without the patches.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper’s central qualitative claim—that most models produce net growth of +0.01 to +100 bar and that impacts were therefore a significant early volatile source—depends on the extensive algorithmic interventions catalogued in Appendix B. Those interventions zero non-physical losses (Svetsov 2000 when the exponential term diverges for small r_imp / dense atmospheres; Svetsov 2007 when gains become negative or complex; Shuvalov when ξ < 0 or χ_a → ∞), force ζ ≥ 0, cap gains, and apply the Svetsov obliquity factor outside its original domain. Because the same Monte-Carlo impactor sequence is used for every model, any systematic bias introduced by these zeroings or caps is shared across the comparison. The composite further discards any model that still requires a patch inside its preferred size window, so the reported growth is conditioned on the very regions the authors themselves flag as non-physical. The reader correctly flags the patches as the weakest assumption; the load-bearing risk is that the net-growth result is an artifact of those patches rather than a robust property of the underlying physics.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript uses Monte Carlo sampling of 5e6 impactors (masses ~10^20 kg) to evolve initial atmospheres (0.006–92.5 bar) on Venus, Earth and Mars under seven published impact-alteration models (Sector/Vickery–Melosh, Pham, Svetsov 2000/2007, Genda & Abe, Shuvalov, Kegerreis) plus two literature composites and one new piecewise composite that averages models only inside their preferred size windows. After documenting extensive algorithmic patches (Appendix B) needed to keep the formulas from diverging or becoming negative/complex, the authors report that single-model runs starting from present-day pressures produce final pressures that differ by 2–3 orders of magnitude, that most models and starting conditions yield net growth of +0.01 to +100 bar, and that both an early-Mars (1 bar) and early-Earth (0.25 bar) atmosphere grow under bombardment. They conclude that using any single model is risky and that impacts were likely a significant early volatile source.","tokens_in":31960,"tokens_out":1272,"duration_ms":19221,"significance":"A systematic, apples-to-apples Monte-Carlo comparison of the existing impact-erosion/gain scalings is valuable; the community has long applied these formulas outside their original domains without quantifying the resulting scatter. Public release of the code (Huffman & Johnston 2025) and the transparent interquartile envelopes strengthen reproducibility. If the net-growth result survives scrutiny of the Appendix-B interventions, the work would tighten the “cosmic shoreline” argument and supply a useful prior for early-atmosphere volatile budgets on the terrestrial planets and rocky exoplanets.","major_comments":[{"comment":"Appendix B (and the “after” panels of Figs. 4 & 6) catalogues numerous non-physical fixes—zeroing infinite Svetsov-2000 losses when the exponential term diverges for small r_imp/dense atmospheres, forcing Svetsov-2007 gains that become negative or complex to zero, capping gains at 10^30 kg, setting ξ≤0 or χ_a\to∞ in Shuvalov, applying the Svetsov obliquity factor outside its derivation domain, etc. Because the same impactor sequence is used for every model, these interventions systematically suppress large losses. The central qualitative claim (most models produce net growth of +0.01–+100 bar; impacts were a significant early volatile source) therefore rests on the patches. A load-bearing sensitivity test is required: re-run the Monte-Carlo suite with the patches disabled (or replaced by hard domain cuts that simply discard the offending impactors) and report how the median ΔP and the gro","section":null},{"comment":"§4.2 and Fig. 6f: the authors’ composite averages gain and loss only inside preferred size windows and discards any model that still requires an Appendix-B patch. This procedure guarantees that the reported composite growth is conditioned on the very regions the authors themselves flag as non-physical. The discontinuities that remain are acknowledged as non-physical, yet the composite is still used to claim that “the atmospheres of Venus, Earth and Mars tend to grow.” Either justify why arithmetic averaging of patched models is physically preferable to the existing de Niem or Schlichting composites, or replace the average with a transparent envelope that shows the full range of the unpatched component models.","section":null},{"comment":"§5 and Table 3: the paleopressure case studies (Mars P0=1 bar, Earth P0=0.25 bar) are presented as thought experiments, yet the text still ranks models by how closely P_initial+ΔP_bombardment matches modern pressure after literature loss estimates. Because those loss estimates themselves carry large systematic uncertainties, and because the bombardment ΔP already incorporates the Appendix-B patches, the ranking of Schlichting and Kegerreis as “most reliable” is not robust. Either remove the ranking or quantify how the ranking shifts when the patches are varied.","section":null}],"minor_comments":[{"comment":"Abstract and §2.2: the impactor count is written “5x10^6” and “5,000,000”; standardize scientific notation and state the corresponding total mass range once in the abstract.","section":null},{"comment":"Fig. 3 caption and §2.1: the preferred size ranges are shown graphically but never tabulated with explicit numerical bounds; a short table would make the composite construction reproducible without inspecting the figure.","section":null},{"comment":"Table 1 and Appendix A: several symbols (e.g., ζ, χ_imp, f_M) are defined only in the appendix; a brief “notation” paragraph or inline definitions would help readers who consult only the main text.","section":null},{"comment":"§2.1.2: the choice n=250 for the Pham efficiency factor is described as a logarithmic average of the extremes; a one-sentence sensitivity check (n=10 vs n=2400) would quantify how free this parameter remains.","section":null},{"comment":"Throughout: occasional typographic slips (“Svestov”, “matmgain”, missing spaces around operators) should be cleaned in copy-editing.","section":null}],"recommendation":"major_revision","confidential_remarks":"The comparison itself is a useful community service and the code release is commendable. The load-bearing issue is whether the net-growth conclusion survives the removal of the ad-hoc patches; if the authors can show that it does (or can clearly bound the effect of the patches), the paper becomes a solid contribution. If the growth largely disappears once the patches are disabled, the manuscript would need substantial reframing. Scope is appropriate for JGR: Planets."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful new result is the systematic Monte-Carlo comparison itself: same 5e6-impactor sequences run through Pham, both Svetsovs, Genda & Abe, Shuvalov, Kegerreis, sector, plus de Niem and Schlichting composites, plus their own piecewise average. Code is public, IQRs are clean, and the 2–3-order spread in final pressure when every model is forced outside its preferred size range is real and well-documented. That alone makes the “any single model is risky” claim solid.\n\nThey also do the honest thing of cataloguing every algorithmic patch in Appendix B (zeroing divergent Svetsov losses, capping gains, forcing ζ ≥ 0, applying the Svetsov obliquity factor outside its derivation domain, discarding any model that still needs a fix inside its preferred window). The paleopressure case studies are correctly labelled thought experiments, not fits. Citation pattern is appropriate; free parameters are few and declared.\n\nThe soft spot is exactly the one the stress-test flags, and it is load-bearing for the stronger claim. Net growth of +0.01 to +100 bar and the “impacts were a significant early volatile source” language rest on those patches. Without them several models become unusable or reverse sign for dense atmospheres or small impactors. The composite further conditions on the regions that do not require patches, so the reported growth is not independent of the interventions. The discontinuities in the piecewise composite are non-physical by the authors’ own admission. Quantitative bar-level forecasts are therefore not trustworthy; the qualitative warning about model disagreement is.\n\nThis is for people who already work on atmospheric evolution or late-veneer budgets and need a clear map of how badly the existing scalings disagree. It is not a new physical model. I would send it to peer review; the comparison is careful enough and the caveats are already written down. Engage if you need the scatter quantified; do not treat the net-growth numbers as robust without re-running the unpatched versions.","headline":"Clean Monte-Carlo head-to-head of seven impact-atmosphere models shows 2–3 order final-pressure scatter and frequent net growth, but the growth result is conditioned on extensive Appendix-B patches that zero non-physical losses.","tokens_in":32504,"tokens_out":517,"would_cite":true,"duration_ms":5517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Using any single model of how impacts reshape planetary atmospheres is risky because the models disagree by orders of magnitude, yet most still predict net atmospheric growth of 0.01 to 100 bar.","keywords":["impact bombardment","planetary atmospheres","Monte Carlo modeling","atmospheric erosion","volatile delivery","Venus","Earth","Mars"],"falsifier":"A new suite of three-dimensional hydrocode runs spanning the full range of impactor radii (0.3–5000 km) and initial surface pressures (0.006–92.5 bar) that either collapses the order-of-magnitude spread among existing models or confirms that most regimes still produce net atmospheric growth of the same magnitude.","tokens_in":32485,"feed_emoji":"☄️","tokens_out":963,"duration_ms":17113,"temperature":0.7,"pith_summary":"The paper asks what happens to the atmospheres of Venus, Earth, and Mars when they are hit by millions of asteroids and comets. It runs the same Monte Carlo set of five million impactors through seven published models of impact-driven gain and loss, then through a composite that uses each model only in its preferred size range. When every model is forced onto every impactor, the final surface pressures can differ by two or three orders of magnitude, so relying on any one model is unreliable. Most models and most starting pressures still produce net growth rather than erosion. The composite likewise grows all three atmospheres, with Earth growing fastest; early Mars at 1 bar and early Earth at 0.25 bar also grow. The authors conclude that impact bombardment was likely a major source of the volatiles that built secondary atmospheres in the early Solar System.","feed_headline":"Impact models disagree by orders of magnitude yet mostly grow atmospheres","feed_subtitle":"Monte Carlo runs of seven models show net growth of 0.01–100 bar, implying early impacts delivered volatiles","key_machinery":"Sequential Monte Carlo evolution of atmospheric pressure under 5 million impactors whose sizes, velocities, and volatile contents are drawn from observed distributions; each impactor updates the atmosphere under either an individual literature model or a size-restricted composite average of those models.","core_discovery":"When seven individual impact-atmosphere models are applied to the same large Monte Carlo impactor population starting from present-day pressures, final atmospheric pressures for Venus, Earth, and Mars spread over roughly two to three orders of magnitude. Most models and most initial pressures (0.006–92.5 bar) nevertheless yield net growth between +0.01 and +100 bar. A composite that restricts each component to its preferred size regime also produces net growth for all three planets, with Earth’s atmosphere growing most quickly; early Mars (1 bar) and early Earth (0.25 bar) atmospheres likewise grow. Single-model use is therefore risky, and impacts were likely a significant early volatile sou","pith_inferences":["The discontinuities in the composite imply that a single modern hydrocode campaign covering the full size and pressure range could replace the patchwork of older analytic and two-dimensional fits.","A systematic net-gain trend would shift the cosmic shoreline toward more planets retaining atmospheres after late accretion, raising the expected number of potentially habitable worlds.","Because loss algorithms are more planet-dependent than gain algorithms, comparative Venus–Earth studies may be more diagnostic of model correctness than absolute pressure evolution on one body alone."],"forward_implications":["Heavy bombardment more often thickens than thins secondary atmospheres, raising the chance of surface habitability after impact eras.","Early volatile inventories of the terrestrial planets may have been substantially supplied by impact delivery rather than solely by outgassing.","Model choice alone can change predicted final pressure by factors of 10–100, so multi-model ensembles are required for reliable evolutionary histories.","Earth’s atmosphere grows faster than Venus’s or Mars’s under identical bombardment, so planetary parameters matter as much as impactor flux.","Paleopressure reconstructions that ignore impact delivery will systematically under-estimate the later atmospheric loss needed to reach modern values."],"fun_headline_variants":["Seven impact models span 2-3 orders yet mostly grow atmospheres","Single impact model use is risky amid huge pressure spreads","Most models yield net growth of 0.01-100 bar under bombardment","Composite shows atmospheres grow; Earth's grows most quickly","Early Mars and Earth atmospheres tend to grow via impacts"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the many algorithmic patches required to stop older models from producing non-physical gains or losses still leave those models comparable, and that averaging them only inside their preferred size windows produces a meaningful composite rather than an artifact of the patches and discontinuities.","fun_headline_variants_meta":{"raw":{"variants":["Seven impact models span 2-3 orders yet mostly grow atmospheres","Single impact model use is risky amid huge pressure spreads","Most models yield net growth of 0.01-100 bar under bombardment","Composite shows atmospheres grow; Earth's grows most quickly","Early Mars and Earth atmospheres tend to grow via impacts"]},"model":"grok-4.5","effort":"low","cost_usd":0.006616,"raw_usage":{"total_tokens":1768,"prompt_tokens":902,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":66160000,"prompt_tokens_details":{"text_tokens":902,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":781,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":902,"tokens_out":85,"duration_ms":6530,"temperature":1.0,"reasoning_tokens":781,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T01:02:03.194693+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A new suite of three-dimensional hydrocode runs spanning the full range of impactor radii (0.3–5000 km) and initial surface pressures (0.006–92.5 bar) that either collapses the order-of-magnitude spread among existing models or confirms that most regimes still produce net atmospheric growth of the same magnitude.","supporting_citations":[],"review_version":1}