{"id":"8823c817-2b26-4512-9ebf-d4356a0c4dd1","arxiv_id":"2511.09716","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In a simulated operational exercise, the SOFIE physics-based SEP model delivered a 4-day forecast in ~5 hours on 1,000 cores, reproducing observed proton fluxes within a factor of 2–6.","lead":"This paper reports a practice run at NOAA's Space Weather Prediction Center in May 2025, where the physics-based SOFIE model was tested as if it were a live forecasting tool for solar energetic particle storms. It produced four-day forecasts in about five hours on 1,000 computer cores — faster than real time — a key step toward physics-based radiation warnings for astronaut missions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Setup-2 flux skill for 4 Nov 2001 rests on a post-hoc seed-injection scale (10x default) chosen to match GOES; the paper does not show this is knowable pre-event. A fixed-scale rerun would test whether the advertised factor-of-2-3 agreement survives.","rationale":"The paper's core feasibility claim — a physics-based SEP model can complete a 4-day forecast in about 5 hours on 1,000 cores — is supported by Table 2 and Fig. 9, independent of the flux-scaling issue. However, the stronger claim that Setup 2 'reproduces key features' within a factor of 2–3 is not independent of the answer: the M-FLAMPA seed scaling factor was set to 10 for this event explicitly to match GOES, so the apparent skill is partly constructed. This is exactly the reader's weakest assumption, and it justifies a CONDITIONAL rather than ACCEPT verdict. The proposed test is feasible because the model, input data, and archived outputs are public, and it would settle whether the fitted factor is load-bearing or merely a minor normalization. I keep the verdict UNCHANGED because the reader's CONDITIONAL already accounts for this concern; my read does not move it to ACCEPT or REJECT. The abstract's 'without compromising accuracy' wording is also contradicted by Table 2 (Setup 2 degrades Spearman correlation and ESP peak intensity relative to Setup 1), but that is an overstatement in the frame, not the main load-bearing issue. The paper's own Section 6.2 disclosure of the scaling factor is a credit to the authors; the concern is about the strength of the prediction claim, not about transparency.","tokens_in":24504,"tokens_out":4603,"duration_ms":51575,"concrete_test":"Rerun the 4 Nov 2001 Setup-2 simulation with the M-FLAMPA injection scaling factor set to its default 1.0, holding all other inputs fixed (grid, AWSoM-R background, EEGGL CME parameters, mean free path 0.1 au), and recompute the Table 2 metrics for >10 MeV and >100 MeV proton fluxes. If the 4-day profiles and onset/ESP peaks remain within factor of 2–3 of GOES and roughly 92% of points still fall within an order of magnitude, the calibration concern is mitigated. If the fluxes drop by ~10x and the skill metrics collapse, the advertised flux skill is an artifact of the fitted scaling factor rather than a prediction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: speed and skill. Speed is well supported: Table 2 and Fig. 9 show Setup 2 finishing a 4-day simulation in 4.86 real hours (catch-up at 1.19 hr). The skill part is the weak link. In M-FLAMPA, the seed-particle injection scaling factor is a free multiplier on the injected proton population (Section 3, Section 6.2; Table 1), with a default of 1.0. For 4 Nov 2001 the authors set it to 10.0 'to better reproduce the GOES measurements' (Section 6.2). Because the SEP intensity is approximately proportional to this factor, the quoted Setup-2 agreement — >10 MeV onset peak 2,126 pfu vs 2,804 pfu observed, 92% within an order of magnitude — is largely an artifact of post-hoc calibration. With factor=1.0, the same run would yield roughly 213 pfu for the >10 MeV onset peak (~13x below GOES) and about 11 pfu for the >100 MeV onset peak (~4x below GOES), erasing the factor-of-2-3 claim. The authors acknowledge the factor 'may need to be fine-tuned' in practice, so the paper does not demonstrate pre-event forecast skill. This is not a fatal flaw, but it is load-bearing: the operational claim of reproducing key features within factor of 2–3 is conditional on a parameter fitted to the event. The grid-resolution comparison (Table 2) and speed metrics remain valid regardless.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the first simulated operational deployment of the SOFIE physics-based SEP prediction suite during the May 2025 SWPT exercise at NOAA/SWPC. The team simulated two historical SEP events (10 September 2017 and 4 November 2001) using the AWSoM-R, EEGGL, and M-FLAMPA components and compared the resulting >10 MeV and >100 MeV proton time-intensity profiles with GOES observations. The main quantitative claim is that with a coarsened solar-corona grid (Setup 2), a 4-day simulation of the 2001 event completed in 4.86 hr on 1,000 CPU cores, catching up with real time at 1.19 hr, demonstrating faster-than-real-time forecasting. Post-exercise runs (Setups 1 and 3) compare three grid configurations; Setup 1 gives the best accuracy but takes 21.12 hr. The paper also presents synthetic white-light CME images, solar-wind plasma comparisons, and forecast metrics (Spearman correlation, factor-of-2 and order-of-magnitude hit rates). The authors explicitly disclose that the background solar wind was prepared in advance and that the M-FLAMPA injection scaling factor was tuned for the 2001 event.","tokens_in":24749,"tokens_out":7201,"duration_ms":77555,"significance":"The runtime result is a genuinely useful operational benchmark: the dated, internally consistent numbers in Table 2 and Fig. 9 (4.86/21.12/18.57 hr for Setups 2/1/3; catch-up at 1.19/10.87/4.11 hr) are concrete and reproducible, with all plotted data archived on Zenodo and the SWMF/SOFIE code publicly available. If the speed result generalizes, it directly addresses a frequently cited bottleneck for physics-based SEP forecasting. The paper also gives a valuable three-way grid-resolution comparison and is candid about parameter tuning and the exclusion of background-preparation time. However, the operational-accuracy claim is only partially supported: the 2001 Setup-2 flux agreement relies on a seed-injection scale factor set to 10.0 to match GOES, and the abstract's 'without compromising accuracy' is contradicted by the paper's own Table 2 metrics. The 2017 event, run with the default scale factor of 1.0, is the only event demonstrating genuine forecast-mode flux skill.","major_comments":[{"comment":"The Setup-2 flux agreement for the 4 November 2001 event is partly by construction. The M-FLAMPA injection scaling factor (Table 1) was set to 10.0 for this event 'to better reproduce the GOES measurements' (§6.2), whereas the default is 1.0. Since the predicted SEP intensity is approximately proportional to this factor, the claimed factor-of-2–3 onset agreement (§5.3) reflects post-hoc calibration, not a pre-event forecast. A factor-1.0 run would reduce the quoted >10 MeV onset peak from ~2,126 pfu to ~213 pfu, about 13× below the observed 2,804 pfu. Please quantify this sensitivity and either adopt a pre-specified/default scale or clearly reframe the 2001 flux skill as a calibrated demonstration. This does not affect the runtime conclusion, but it is load-bearing for the operational-accuracy claim.","section":"§5.3, §6.2, Table 1"},{"comment":"The abstract states that the coarser background grid with higher-resolution regions reduces computational cost 'without compromising accuracy.' Table 2 does not support this for Setup 2 relative to Setup 1: for >10 MeV the Spearman correlation drops from 0.929 to 0.841 and the within-factor-of-2 rate from 41.7% to 37.7%; for >100 MeV the Spearman correlation drops from 0.918 to 0.563 and the ESP peak falls from 132 to 28 pfu against 162 pfu observed. Setup 2 preserves the onset and decay phases reasonably but clearly compromises ESP-phase fidelity. Please revise the claim to specify which accuracy aspects are retained, or provide a statistical equivalence test.","section":"§6.1, Table 2, Abstract"},{"comment":"The headline '5 hours' is not an end-to-end forecast time. As the paper explicitly states, 'for both events, we prepared the background solar wind in advance, and the corresponding timing is not included' (§7). The AWSoM-R steady-state background is a necessary component of SOFIE, and its computational cost is not part of the 4.86-hr total. The abstract's 'completed ... within 5 hours' should be qualified as the eruption-triggered run assuming a precomputed background. Please report the end-to-end operational timeline (magnetogram ingestion, background restart, CME-parameter availability, M-FLAMPA run) or state in the abstract that the quoted time excludes the background.","section":"§7, §3, Abstract"}],"minor_comments":[{"comment":"In the text near Figure 3(e), 'observed by ACR' appears to be a typo for 'ACE'.","section":"§4.2"},{"comment":"The five statistical metrics (Spearman correlation, percentage within an order of magnitude, percentage within a factor of 2, median logarithmic error, median absolute logarithmic error) are presented without definitions. Since they carry the quantitative comparison, please define them in the text or cite a published standard with explicit formulas.","section":"Table 2"},{"comment":"The claim that the 2017 event's onset and peak fluxes agree 'within a factor of 2' is not supported by a table or listed numerical values. Please provide the same quantitative summary used for the 2001 event (e.g., peak values and leading times) for the 2017 event.","section":"§4.3"},{"comment":"The three setups are described in the text and in Figs. 2/5, but a compact table summarizing the exact AMR criteria (e.g., angular resolution, cone half-widths, refinement levels) for Setups 1–3 would improve reproducibility.","section":"§6.1, Fig. 8"}],"recommendation":"major_revision","confidential_remarks":"I would not reject on the basis of the fitted injection parameter; the speed result is solid and the authors are transparent about tuning. However, the paper needs either a factor=1 rerun or a substantially revised accuracy framing before the operational-skill claim can be accepted as stated. The abstract's 'without compromising accuracy' must be reconciled with Table 2, and the background-preparation caveat should be moved into the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new result here is the operational timing: SOFIE finishing a 4-day SEP simulation in about 5 hours of wall-clock on 1,000 cores (Setup 2), with catch-up to real time at 1.19 hours. That claim holds. The numbers are concrete, internally consistent across Table 2 and Figure 9, and the underlying data are on Zenodo. The paper also gives a useful three-way grid comparison (Setups 1–3) that quantifies the accuracy/runtime tradeoff, and it was run under simulated operational conditions with real forecaster input. That is a legitimate pilot milestone for physics-based SEP forecasting.\n\nThe soft spot is the flux-skill claim for the 4 Nov 2001 event. The M-FLAMPA injection scaling factor was set to 10.0 explicitly \"to better reproduce the GOES measurements\" (Section 6.2, Table 1). Since the predicted flux scales almost linearly with that factor, the factor-of-2–3 onset agreement for that event is partly assured by construction. A fixed-scale rerun would likely put the >10 MeV peak roughly an order of magnitude low. The authors state the factor \"may need to be fine-tuned\" in practice, which is honest but does not demonstrate pre-event forecast skill. The other event, 10 Sept 2017, used the default scaling of 1.0 and is therefore a more genuine prediction—though that one has its own caveats (no direct CME impact, omitted preceding CMEs for the 2001 event).\n\nThe abstract overreaches when it says the grid optimization cut cost \"without compromising accuracy.\" Table 2 shows Setup 2 degrades every metric relative to Setup 1: Spearman 0.84 vs 0.93, ESP peak 4,535 vs 27,149 pfu, and similar drops in the >100 MeV channel. The paper's own Section 6.1 acknowledges the coarser setup gives an earlier and weaker ESP phase. That internal inconsistency between abstract and table should be fixed in revision.\n\nOther concerns are minor: two-event sample, no uncertainty quantification, and the 5-hour wall-clock excludes the precomputed AWSoM-R background (they state this clearly, and daily background runs are part of the operational pipeline, so it is not a fatal omission). The citation pattern is fine; self-citations point to the model validation papers, which is appropriate.\n\nWho benefits: space weather operations teams, SEP model developers, and anyone judging whether physics-based SEP models can meet latency requirements. It deserves a serious referee and likely publication after revision—align the abstract with Table 2, show a scaling=1.0 rerun for 2001, and soften the \"robustness\" conclusion to match the two-event sample.","headline":"Useful speed demonstration and grid-comparison, but the 2001 flux skill is partly fitted via a post-hoc injection scaling factor; the abstract overstates the accuracy-vs-speed tradeoff.","tokens_in":25638,"tokens_out":2160,"would_cite":true,"duration_ms":24175,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A physics-based solar-particle model produced a 4-day radiation forecast in 5 hours during a 2025 operational exercise.","keywords":["solar energetic particles","space weather prediction","SEP forecasting","coronal mass ejection","diffusive shock acceleration","MHD solar wind model","operational testbed","radiation risk to astronauts"],"falsifier":"Take the same fast grid setup and M-FLAMPA parameters used here, fix the injection scaling factor at its default value of 1.0 before running, and apply the pipeline to a set of historical SEP events; if the predicted >10 MeV onset peaks and integral fluxes stop agreeing with spacecraft measurements to within the factor of 2–3 shown for 2001, then the accuracy demonstrated here was fitted rather than forecast.","tokens_in":24179,"feed_emoji":"☀️","tokens_out":7800,"duration_ms":81542,"temperature":0.7,"pith_summary":"This paper reports the first operational-style test of SOFIE, a physics-based model of solar energetic particles (SEPs) that couples an MHD model of the ambient solar wind, a flux-rope model of the coronal mass ejection (CME), and a field-line model of particle acceleration and transport. The authors' goal was to answer a practical question: can such a model, normally thought of as too expensive for real time, deliver SEP forecasts within the latency of a space-weather operation? They ran SOFIE on 1,000 CPU cores during a May 2025 exercise using two historical events. For the 4 November 2001 event, a deliberately coarsened solar-corona grid let the 4-day simulation finish in about 5 hours of wall-clock time — about 91 hours ahead of the event timeline — while reproducing the observed >10 MeV and >100 MeV proton profiles at onset to within a factor of 2–3. If that performance carries over to genuine events, a single pipeline could give astronauts and mission operators multi-day, all-sky radiation-flux maps within hours of a CME detection.","feed_headline":"Physics-based solar storm forecast finishes in 5 hours","feed_subtitle":"Radiation-flux maps can be ready hours after a CME is spotted, not days.","key_machinery":"The speed comes from the way SOFIE is assembled and from the adaptive grid, not from any single new physical term. AWSoM-R supplies a precomputed, stream-aligned magnetohydrodynamic (MHD) solar wind; EEGGL inserts a Gibson–Low magnetic flux rope matched to the observed active-region location and CME speed; M-FLAMPA (the Multiple Field-Line-Advection Model for Particle Acceleration) simulates diffusive shock acceleration and transport of protons along magnetic field lines advected by that wind. The operational innovation tested here is a two-level grid strategy: the global solar-corona background is coarsened by a factor of two while adaptive mesh refinement stays on the heliospheric current","core_discovery":"On its own terms, the paper establishes that the complete SOFIE chain can clear the operational bar: the full 4-day simulation of an SEP event finished in 4.86 hours of real time using the fast grid setup, and the earliest useful 10-hour forecast was available within 2.35 hours of eruption (leading time 7.65 hours). Accuracy at the operational channels: for >10 MeV, about 92% of predicted points fell within an order of magnitude of spacecraft measurements and the onset peak within a factor of 2–3; for >100 MeV, 92% within an order of magnitude and about 49% within a factor of 2. The paper also compares three grid designs and reports that the default high-resolution grid is the most accurate","pith_inferences":["The 5-hour figure excludes the precomputed ambient solar wind; a full event-to-answer latency would include the daily background-solar-wind preparation, though that pipeline runs continuously and can be restarted from saved states.","Because the injection scaling factor was tuned to 10 for the 2001 event after the fact, the cleanest prospective test of forecast skill is to fix that factor at its default before an event and measure the resulting bias across many events.","The same fast pipeline could be run as an ensemble (varying CME speed, active-region location, and mean free path) to produce probabilistic radiation-dose maps rather than single deterministic profiles — an extension the paper does not develop.","The Setup-2 ESP-phase timing error (onset about 7 hours early, peak reduced severalfold) suggests numerical diffusion from the coarser grid; a grid-refinement policy between the fast and intermediate setups might recover ESP accuracy without the full cost of the high-resolution setup."],"forward_implications":["For the 4 November 2001 event, the fast grid setup delivered the full 96-hour profile in 4.86 hours of real time, yielding a leading time of about 91 hours over the event timeline.","About 92% of the fast setup's >10 MeV predicted points fell within an order of magnitude of spacecraft measurements, with onset peak within a factor of 2–3; the default grid raised this to over 99% and Spearman correlation above 0.9, but took about 21 hours.","The coarser-corona grid removes the main computational bottleneck, which is the solar-corona domain with its small cells and high wave speeds; after the CME leaves that domain, the simulation runs faster than real time.","The paper recommends a two-phase operational workflow: run the fast setup for the first roughly 5 hours for order-of-magnitude guidance, then switch to the high-resolution setup for refined flux profiles.","The paper concludes that a physics-based SEP model can meet the latency and robustness requirements of operational space-weather prediction for crewed exploration missions."],"fun_headline_variants":["SOFIE solar storm model forecasts 4 days in 5 hours","Physics-based SEP forecast ready hours after CME, not days","Operational test: fast grid cuts SEP model run to 5 hours","Solar particle radiation maps in hours, not days, during test","SEP model clears operational bar: 5-hour run for 4-day forecast"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the particle-injection scaling factor — a free multiplier on how many protons are seeded at the shock — can be set before an event and left fixed; in this exercise it was raised from the default of 1.0 to 10.0 for the 4 November 2001 event specifically to make the modeled flux match the spacecraft measurements, and the 5-hour runtime also excludes the precomputed background solar wind.","fun_headline_variants_meta":{"raw":{"variants":["SOFIE solar storm model forecasts 4 days in 5 hours","Physics-based SEP forecast ready hours after CME, not days","Operational test: fast grid cuts SEP model run to 5 hours","Solar particle radiation maps in hours, not days, during test","SEP model clears operational bar: 5-hour run for 4-day forecast"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001167,"raw_usage":{"total_tokens":4732,"prompt_tokens":877,"completion_tokens":3855,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":3761}},"tokens_in":621,"tokens_out":3855,"duration_ms":28292,"temperature":1.0,"reasoning_tokens":3761,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:34:12.638925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same fast grid setup and M-FLAMPA parameters used here, fix the injection scaling factor at its default value of 1.0 before running, and apply the pipeline to a set of historical SEP events; if the predicted >10 MeV onset peaks and integral fluxes stop agreeing with spacecraft measurements to within the factor of 2–3 shown for 2001, then the accuracy demonstrated here was fitted rather than forecast.","supporting_citations":[],"review_version":1}