{"id":"1a1a08fc-f120-4687-acb4-bdf637b9d92d","arxiv_id":"2607.07951","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"Fully trained BiLSTM outperforms zero-shot and LoRA-adapted TSFMs on California wildfire PM2.5 under leave-one-incident-out evaluation, especially at hazardous AQI thresholds.","lead":"A compact BiLSTM beats six time-series foundation models, zero-shot and LoRA-tuned, at forecasting extreme wildfire PM2.5 under leave-one-incident-out evaluation. The result challenges the idea that large pretrained models automatically win on rare environmental extremes and gives clear deployment rules for air-quality warning systems.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified that would invert the BiLSTM-over-TSFM hierarchy under the stated LOIO protocol.","rationale":"The strongest claim is tightly scoped to the LOIO univariate benchmark and is supported by the reported tables and figures. The reader's weakest assumption correctly flags the univariate design as the main external-validity caveat, yet that design is deliberate and applied uniformly, so it does not invalidate the internal ranking. Residual parameterization for the trained models (Section 4.1) is the only plausible hidden advantage; the proposed concrete test isolates it without requiring new data. Code/data non-release keeps the verdict CONDITIONAL rather than ACCEPT, matching the reader. No stronger load-bearing flaw (leakage, metric gaming, or selective reporting) is evident in the manuscript.","tokens_in":20917,"tokens_out":473,"duration_ms":5008,"concrete_test":"Re-run the LOIO protocol with the same 48 h univariate windows but replace residual targets (Eq. 6) for the trained baselines with direct native-scale targets matching the TSFM interface; if BiLSTM's MAE rises above Chronos-2 (LoRA-FT) or its Hazardous F1 falls below 0.54, the residual parameterization (not architecture) would be driving the ranking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is an empirical ranking under a carefully controlled univariate LOIO design (Section 3.3, Table 4, Figures 4–6). The ranking is internally consistent: BiLSTM leads on MAE, RMSE, R², and exceedance F1 at every AQI threshold; zero-shot Chronos-2's RMSE/R² instability is reported transparently; LoRA narrows but does not close the gap. The reader's weakest assumption (univariate 48 h PM2.5 without covariates) is a genuine scope limitation, already acknowledged in the Discussion, but it does not undermine the claim as stated—namely that, under identical univariate inputs, no TSFM configuration surpasses the trained recurrent baselines. Expanding to multivariate inputs would test a different, broader claim about operational readiness, not the fairness of the present comparison. No internal inconsistency, leakage, or metric artifact appears load-bearing enough to reverse the hierarchy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper benchmarks time series foundation models (zero-shot TimesFM, Chronos-2, Moirai-2, Time-MoE, plus LoRA-adapted Chronos-2 and Time-MoE) against fully trained LSTM, BiLSTM, and Transformer baselines and naïve persistence for wildfire-driven PM2.5 forecasting. Using a 12-year California panel (79 sites, 1,375 incidents) and a severity-stratified leave-one-incident-out protocol, it evaluates MAE, RMSE, R², and exceedance F1 at EPA AQI thresholds over 6/12/24-hour horizons from a shared 48-hour univariate context. The central empirical claim is that BiLSTM attains the lowest MAE (5.16 µg/m³) and the highest exceedance F1 at every threshold, including Hazardous (0.63), while no zero-shot or LoRA-adapted foundation model surpasses the trained recurrent baselines on any metric; zero-shot Chronos-2 shows severe RMSE/R² tail instability that LoRA largely repairs without closing the gap.","tokens_in":21227,"tokens_out":1235,"duration_ms":11228,"significance":"If the ranking holds, the work is a useful corrective to the assumption that large pretrained TSFMs automatically dominate extreme environmental forecasting. Strengths include a carefully designed LOIO protocol that avoids within-incident leakage, severity-stratified folds, shared windows across horizons, residual parameterization for trained models, transparent reporting of Chronos-2 tail instability, and operationally relevant exceedance F1 at multiple AQI thresholds. The released multi-incident California benchmark and the deployment guidance (prefer compact trained recurrent models when multi-year fire history exists; use LoRA TSFMs only as a cold-start path) are concrete contributions for air-quality and environmental ML practice.","major_comments":[{"comment":"Section 4.1 and Table 3: the trained baselines predict residuals from the last observed standardized value (Eqs. 6–7), while zero-shot TSFMs (except Time-MoE) receive native-scale series and emit direct median quantiles. Residual parameterization is a strong inductive bias that anchors forecasts to persistence and stabilizes heavy-tailed training; the paper does not ablate a non-residual trained baseline or residual-style adaptation for TSFMs. Without that control, part of the BiLSTM advantage may be attributable to target parameterization rather than architecture or pretraining alone. A short ablation (or explicit residual-head TSFM variant) would make the hierarchy more conclusive.","section":null},{"comment":"Section 4.6 and Table 4: LoRA is applied only to Chronos-2 and Time-MoE, with a fixed budget (500 steps, r=8 for Time-MoE, lr=1e-5). The claim that \"no foundation model, zero-shot or fine-tuned, surpasses the trained recurrent baselines\" therefore rests on two adapted families only. Given that TimesFM and Moirai-2 are competitive zero-shot, either adapting them under the same LOIO protocol or justifying their exclusion is needed before the adaptation conclusion is fully general.","section":null},{"comment":"Section 5 and Discussion: all models are strictly univariate (48 h PM2.5 only). The authors correctly note this as a limitation, but the abstract and contribution list frame the result as challenging the universal dominance of larger pretrained models in environmental forecasting. That framing is stronger than the controlled univariate comparison supports. Soften the claim to the stated setting (identical univariate inputs, LOIO wildfire PM2.5) and treat multivariate/covariate-capable TSFMs as future work rather than as already refuted.","section":null}],"minor_comments":[{"comment":"Abstract vs. Table 4: Chronos-2 zero-shot RMSE is given as 23.4 µg/m³ in the abstract and 23.45 in the table; keep one consistent rounding convention.","section":null},{"comment":"Figure 6 includes a Moderate (9.1) threshold that is not listed among the primary AQI breakpoints in the abstract or Table 4; either add it to the main table or drop it from the figure for consistency.","section":null},{"comment":"Section 3.2: \"evaluation was conducted on held-out continuous weeks containing elevated-concentration episodes\" is slightly ambiguous relative to the LOIO protocol of Section 3.3; clarify that LOIO is the sole evaluation design.","section":null},{"comment":"Table 3 lists Moirai-2 as \"patch; dec.\" while Section 4.5 describes it as a patch-based encoder; align the architecture description.","section":null},{"comment":"Figure 7 caption and body: incident ID is truncated differently across places; use a single full identifier for reproducibility.","section":null},{"comment":"A few typographic issues: \"naïve\" vs \"naïve\" consistency, and the arXiv date line \"July 10, 2026\" looks like a placeholder.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The manuscript is a solid empirical systems paper rather than a methods breakthrough. Fit is good for an applied ML / environmental data science venue; for a top general ML conference the univariate scope and limited LoRA coverage would likely draw stronger pushback. No integrity or novelty-disclosure concerns. The LOIO design and transparent Chronos-2 failure mode are the main reasons I favor minor rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: under a severity-stratified leave-one-incident-out protocol on 12 years of California wildfire PM2.5, a compact BiLSTM beats every foundation-model configuration they tried—zero-shot TimesFM, Chronos-2, Moirai-2, Time-MoE, and LoRA-adapted Chronos-2 and Time-MoE—on MAE, RMSE, R2, and exceedance F1 at every EPA AQI threshold, including Hazardous. That ranking is the result worth knowing.\n\nWhat is actually new is the evaluation design more than any architecture. Prior PM2.5 work almost always uses chronological splits that can leak within a long-lived fire; this paper groups by incident, stratifies folds by peak severity, shares the same 48 h windows across 6/12/24 h horizons, and reports both aggregate error and binary exceedance F1. The Chronos-2 RMSE/R2 tail instability is reported honestly rather than buried. Residual targets for the trained models and native-scale inputs for the foundation models are handled cleanly, and LoRA is applied per fold so the protocol is preserved. The hierarchy is stable across folds and horizons.\n\nSoft spots are real but proportionate. Everything is univariate PM2.5; no wind, humidity, burned area, or fire radiative power. That is a genuine scope limit for operational claims, and the authors say so. Only two of the four foundation families get LoRA, the adaptation budget is fixed, and code/data artifacts are not shipped—only public CARB and CAL FIRE sources. None of that inverts the claim as stated: under identical univariate inputs and true incident holdout, the trained recurrent baselines win. Expanding to covariates would test a broader claim, not the fairness of this comparison.\n\nMath and metrics look solid; the citation pattern covers the relevant TSFM and wildfire-PM2.5 literature without obvious gaps. This is for people who care about extreme-event forecasting, air-quality systems, or whether foundation-model hype transfers to heavy-tailed environmental series. It deserves a serious referee. I would engage with it and expect it to survive peer review with ordinary revision.","headline":"Careful LOIO benchmark shows a trained BiLSTM beats zero-shot and LoRA TSFMs on extreme wildfire PM2.5; the ranking holds under the stated univariate design.","tokens_in":21790,"tokens_out":540,"would_cite":true,"duration_ms":5670,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A small trained BiLSTM beats zero-shot and LoRA-tuned foundation models at forecasting extreme wildfire PM2.5 under leave-one-incident-out evaluation.","keywords":["PM2.5 forecasting","wildfire smoke","time series foundation models","leave-one-incident-out","exceedance prediction","LoRA fine-tuning","deep learning","air quality"],"falsifier":"Re-run the same LOIO benchmark after adding wind, humidity, temperature inversions, burned area, or fire radiative power as covariates; if a zero-shot or LoRA-tuned foundation model then beats BiLSTM on Hazardous exceedance F1 or MAE, the central claim is overturned.","tokens_in":21806,"feed_emoji":"🔥","tokens_out":690,"duration_ms":5894,"temperature":0.7,"pith_summary":"Wildfire smoke produces rare, hazardous PM2.5 spikes that public-health systems need to forecast ahead of time, yet most air-quality models are tested on chronological splits that can leak information from the same fire. This paper builds a 12-year California benchmark of 1,375 wildfire incidents across 79 monitors and evaluates models with a leave-one-incident-out protocol that holds out entire fires. It compares four zero-shot time-series foundation models and two LoRA-adapted versions against fully trained LSTM, BiLSTM, and Transformer baselines plus naive persistence, measuring both ordinary error and whether hazardous AQI thresholds are correctly flagged at 6-, 12-, and 24-hour horizons. The trained BiLSTM wins every metric, including the hardest Hazardous band, while foundation models only modestly beat persistence and never overtake the recurrent baselines even after fine-tuning. The result matters because it challenges the assumption that large pretrained forecasters automatically dominate extreme environmental events and supplies concrete guidance on when a compact domain-trained model remains the better operational choice.","feed_headline":"Small BiLSTM beats foundation models on wildfire PM2.5","feed_subtitle":"Leave-one-incident-out tests show pretrained TSFMs never top trained recurrent baselines, even with LoRA.","key_machinery":"Leave-one-incident-out (LOIO) cross-validation: every window from a given wildfire–site episode is held out together so that models must generalize to an entirely unseen fire, preventing the within-incident leakage that chronological splits allow.","core_discovery":"Under leave-one-incident-out evaluation on California wildfire PM2.5, a fully trained BiLSTM achieves the lowest MAE (5.16 µg/m³) and the highest exceedance F1 at every EPA AQI threshold, including Hazardous (>225.5 µg/m³) at 0.63, while no zero-shot or LoRA-adapted foundation model surpasses the trained recurrent baselines on any metric.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["BiLSTM tops all foundation models on wildfire PM2.5","Trained BiLSTM beats TSFMs under leave-one-incident tests","No TSFM surpasses BiLSTM on California wildfire PM2.5","BiLSTM leads zero-shot and LoRA models for extreme PM2.5","Fully trained BiLSTM wins LOIO wildfire air-quality forecast"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That a univariate 48-hour PM2.5 history alone, without weather covariates or fire-side drivers, is a fair and sufficient input for comparing foundation models against recurrent baselines on extreme-event generalization.","fun_headline_variants_meta":{"raw":{"variants":["BiLSTM tops all foundation models on wildfire PM2.5","Trained BiLSTM beats TSFMs under leave-one-incident tests","No TSFM surpasses BiLSTM on California wildfire PM2.5","BiLSTM leads zero-shot and LoRA models for extreme PM2.5","Fully trained BiLSTM wins LOIO wildfire air-quality forecast"]},"model":"grok-4.5","effort":"low","cost_usd":0.006268,"raw_usage":{"total_tokens":1742,"prompt_tokens":950,"num_sources_used":0,"completion_tokens":101,"cost_in_usd_ticks":62680000,"prompt_tokens_details":{"text_tokens":950,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":691,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":950,"tokens_out":101,"duration_ms":6011,"temperature":1.0,"reasoning_tokens":691,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T14:49:34.419473+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same LOIO benchmark after adding wind, humidity, temperature inversions, burned area, or fire radiative power as covariates; if a zero-shot or LoRA-tuned foundation model then beats BiLSTM on Hazardous exceedance F1 or MAE, the central claim is overturned.","supporting_citations":[],"review_version":1}