Pith. sign in

REVIEW 3 major objections 23 references

Hybrid CNN-recurrent models beat pure GRU/LSTM for short-horizon weather forecasts used in agriculture.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 13:28 UTC pith:4ORVP43F

load-bearing objection Solid single-station architecture bake-off with usable code; the hybrid gains are real but tiny and the evaluation has a clear temporal-leakage soft spot. the 3 major comments →

arxiv 2607.10208 v1 pith:4ORVP43F submitted 2026-07-11 cs.LG cs.AI

Exploratory Analysis of Deep Learning Models for Forecasting Meteorological Parameters in the Agricultural Sector

classification cs.LG cs.AI
keywords meteorological forecastingreference evapotranspirationvapour pressure deficitLSTMGRU1D-CNN hybridagricultural decision supportERA5
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper compares recurrent and hybrid deep learning models for jointly forecasting five agricultural weather variables (reference evapotranspiration, vapour pressure deficit, wind speed, and the sine and cosine of wind direction) from a long hourly record at one Greek site. It tests single-layer and deep GRU and LSTM networks against 1D-CNN-GRU and 1D-CNN-LSTM hybrids on next-day (24 h) and week-ahead (168 h) tasks, scoring them with a composite Weighted Quality Score that mixes normalized error and explained variance. The central claim is that adding a one-dimensional convolutional front end yields the highest scores and improves on the best pure recurrent baselines by roughly 1.2–1.6 percent at 24 hours and only about 0.45 percent at 168 hours, so local feature extraction helps short-term forecasts more than week-ahead ones. A sympathetic reader cares because irrigation and crop-stress decisions need reliable joint forecasts of water demand and wind, and the work shows which architectural choices actually move the needle under a shared training protocol.

Core claim

Under a common training and evaluation protocol on 134k hourly ERA5-derived observations, hybrid 1D-CNN–GRU and 1D-CNN–LSTM models achieve the highest Weighted Quality Scores (0.827535 at 24 h and 0.782863 at 168 h) and improve on the strongest pure recurrent baselines by 1.22–1.63 percent at the daily horizon and 0.44–0.45 percent at the weekly horizon, showing that convolutional feature extraction is more beneficial for short-term than for week-ahead multivariate meteorological forecasting.

What carries the argument

The Weighted Quality Score (WQS = 0.8·(1 − nRMSE) + 0.2·R²), which collapses normalized multi-step, multi-variable error and explained variance into a single ranking metric used to compare all architectures.

Load-bearing premise

The claim rests on the assumption that fitting min-max scalers on the whole series and then shuffling overlapping sliding windows into train/validation/test sets produces trustworthy out-of-sample scores; if future information leaks into the scales or the splits, the reported hybrid gains are not genuine.

What would settle it

Re-run the identical model suite with scalers fit only on the training windows and with a strictly chronological (non-shuffled, non-overlapping) train/validation/test split; if the hybrid WQS advantage over the best pure recurrent models disappears or reverses, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. The manuscript presents a controlled empirical comparison of single- and multi-layer GRU/LSTM networks against hybrid 1D-CNN–GRU and 1D-CNN–LSTM models for multivariate multi-step forecasting of ET0, VPD, wind speed, and the sine/cosine of wind direction. Using 134 376 hourly ERA5-derived observations for Ioannina (2011–2026), the authors evaluate 24 h and 168 h horizons under a common training protocol and report nRMSE, R2 and a composite Weighted Quality Score (WQS). The central claim is that hybrid models attain the highest WQS (0.827535 at 24 h; 0.782863 at 168 h) and improve on the best pure-recurrent baselines by 1.22–1.63 % and 0.44–0.45 %, respectively, with convolutional feature extraction being more helpful for the shorter horizon; deep stacked recurrent nets are shown to degrade sharply.

Significance. If the reported ranking and relative gains hold under a leakage-free protocol, the work supplies a useful, reproducible station-scale bake-off for agricultural decision-support systems. Strengths include a publicly released code repository, a transparent common training regime, explicit documentation that ten-layer stacks fail, and a clear demonstration that shallow single-layer recurrent models already capture most of the predictable structure. The absolute WQS deltas are small, so the practical significance is modest, yet the comparative design and open artefacts make the study a solid reference point for subsequent agro-meteorological forecasting work.

major comments (3)
  1. §2.2 states that independent min–max scalers were “calculated on the whole time series dataset” before any split. Global extrema therefore leak future information into every train/val/test window. Because the claimed hybrid gains are only ~0.01 WQS (24 h) and ~0.003 WQS (168 h), this leakage is load-bearing for the central claim; scalers must be refit on the training partition alone and the tables recomputed.
  2. §2.2 and Table 1 generate overlapping sliding windows (lookback + horizon ≫ stride) that are then shuffled into a 70/15/15 split. Substantial temporal overlap between train and test windows remains, so the reported test metrics are not clean out-of-sample estimates. A chronological or blocked split (or explicit non-overlapping windows) is required before the 1.22–1.63 % / 0.44–0.45 % improvements can be trusted.
  3. Tables 3, 4, 7 and 9 present absolute WQS differences of order 0.001–0.01 with no confidence intervals, bootstrap, or statistical test. Given the free choice of hybrid hidden sizes (256/1024 vs 64 for the best pure LSTMs) and the hand-tuned WQS weight α = 0.8 (Eq. 7), the ranking cannot be declared robust without uncertainty quantification or matched-capacity ablations.

Circularity Check

0 steps flagged

Empirical bake-off of RNN/CNN hybrids; no derivation that reduces by construction to its inputs.

full rationale

The paper is a controlled comparative evaluation of GRU/LSTM and 1D-CNN hybrids on ERA5-derived multivariate meteorological series. Performance numbers (nRMSE, R^{2}, WQS) are obtained by training and testing concrete architectures under a fixed protocol; they are not algebraic consequences of the metric definition or of any fitted parameter that is later re-labeled a prediction. WQS itself is an explicit linear blend (α=0.8) of the two base metrics and is introduced only as a ranking device; the ranking is therefore designer-chosen but not tautological. Self-citations ([13], [20]) supply background or the WQS formula but are not load-bearing uniqueness theorems that force the architectural conclusions. No self-definitional loop, no fitted-input-called-prediction, and no ansatz smuggled via citation appear in the derivation chain. The methodological concerns (global min-max scaling, shuffled overlapping windows) affect the trustworthiness of the reported deltas but are leakage/correctness issues, not circularity. Score 1 reflects only the minor, non-load-bearing self-citation of the composite metric.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The central ranking claim rests on standard sequence-model practice plus several protocol choices that are not forced by meteorology: whole-series scaling, shuffled overlapping windows, a hand-weighted composite score, fixed training hyperparameters, and treating a single ERA5 grid cell as the agricultural target. No new physical entities are introduced; free parameters are metric and training knobs that can reorder close model scores.

free parameters (5)
  • WQS weight α
    α=0.8 in Eq. (7) assigns 80% weight to (1−nRMSE) and 20% to R²; this hand choice defines the headline ranking metric.
  • Hybrid recurrent hidden sizes
    M5 uses 256 (24 h) / 1024 (168 h) units while M6 uses 64 at both horizons; these unequal capacities are chosen by the authors and affect the hybrid vs recurrent comparison.
  • Lookback, horizon, and stride
    72/24/12 h and 336/168/48 h windows (§2.2, Table 1) define the learning problem; different operational windows could change relative model skill.
  • Dropout, LR, weight decay, batch size, early-stopping patience
    p=0.2, Adam 1e−3, weight decay 1e−4, batch 32, patience 10, etc. are fixed across architectures without systematic search (§3).
  • CNN channel width and kernel
    Two Conv1D blocks expand to 128 channels with kernel size 3 and same padding (§4); these front-end choices are free design parameters of the hybrid models.
axioms (5)
  • domain assumption ERA5 reanalysis at the nearest grid point is an adequate surrogate for local agricultural meteorological decision support.
    Invoked throughout §2.1 and the conclusions; no station-level validation against in-situ sensors is provided.
  • ad hoc to paper Fitting min–max scalers on the full series and then shuffling overlapping windows into train/val/test yields valid generalization estimates.
    Stated in §2.2; this is a strong evaluation assumption that can leak future information into training.
  • domain assumption FAO-56 Penman–Monteith ET0 and the Open-Meteo VPD fields correctly represent the intended agro-meteorological targets.
    Eqs. (1)–(2) and API variable choices in §2.1; standard agronomy but still an external modeling assumption.
  • ad hoc to paper Clipped R²∈[0,1] and nRMSE by test-range are appropriate building blocks for architecture comparison.
    §3.1 evaluation procedure; clipping poor models to R²=0 and range-normalizing RMSE shape the WQS surface.
  • domain assumption Standard GRU/LSTM/CNN sequence-learning machinery can capture the relevant temporal structure without explicit spatial or calendar features.
    Implicit in model design §4 and limitations noting missing hour-of-day / multi-station inputs.
invented entities (1)
  • Weighted Quality/Quotient Score (WQS) as used here no independent evidence
    purpose: Single scalar to rank multivariate multi-step forecasts by blending nRMSE and R².
    Defined in Eq. (7) as inspired by [20]; not a physical quantity and naming is inconsistent between abstract and body. Independent evidence is weak because α is arbitrary and the metric is paper-specific.

pith-pipeline@v1.1.0-grok45 · 23449 in / 3667 out tokens · 37798 ms · 2026-07-14T13:28:50.592570+00:00 · methodology

0 comments
read the original abstract

Accurate meteorological forecasting is essential for agricultural planning, irrigation management, and environmental decision support. This study conducts a comparative evaluation of recurrent and hybrid deep learning architectures for multivariate forecasting of reference evapotranspiration ($ET_0$), vapour pressure deficit (VPD), wind speed, and the sine and cosine components of wind direction. The analysis utilizes 134,376 hourly observations from Ioannina, Greece, spanning January 2011 to April 2026, sourced from ERA5 via the OpenMeteo Historical Weather API. Single and multi-layer GRU and LSTM networks are compared with hybrid 1D-CNN-GRU and 1D-CNN-LSTM models for two forecasting tasks: a 24-hour next-day forecast and a 168-hour week-ahead forecast. Performance is evaluated using normalized root mean squared error, the coefficient of determination, and a composite Weighted Quotient Score (WQS). The most effective purely recurrent models are a 64-unit LSTM for the 24-hour horizon, with a WQS of 0.816755, and a 1024-unit GRU for the 168-hour horizon, with a WQS of 0.779465. The hybrid CNN-GRU models achieved the highest overall scores of 0.827535 and 0.782863 for the 24-hour and 168-hour horizons, but with additionally more number of units respectively to LSTM models, while the CNN-LSTM models yield nearly identical results with substantially fewer parameters. Compared to the corresponding recurrent baselines, the hybrid models improve WQS by 1.22--1.63\% at 24 hours and by 0.44--0.45\% at 168 hours, indicating that convolutional feature extraction is more beneficial for short-term forecasting.

Figures

Figures reproduced from arXiv: 2607.10208 by Piotr Sikora, Sotirios Kontogiannis.

Figure 1
Figure 1. Figure 1: Graphs comparing daily values of Reference [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of the six forecasting architectures. M1 and [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Training and validation loss curves for the single-layer [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Observed and predicted reference evapotranspiration [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Observed and predicted reference evapotranspiration [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 5 canonical work pages

  1. [1]

    R. G. Allen, L. S. Pereira, D. Raes, and M. Smith,Crop Evapotranspiration: Guidelines for Computing Crop Wa- ter Requirements, ser. FAO Irrigation and Drainage Paper. Rome, Italy: Food and Agriculture Organization of the United Nations, 1998, no. 56

  2. [2]

    Empirical evaluation of gated recurrent neural networks on sequence modeling,

    J. Chung, C. Gulcehre, K. Cho, and Y . Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” arXiv preprint, 2014, available at: https: //arxiv.org/abs/1412.3555 (accessed 25 Mar 2021). [Online]. Available: https://doi.org/10.48550/arXiv.1412.3555

  3. [3]

    Exploring machine learning, deep learning, and explainable ai methods for seasonal precipitation prediction in south america,

    M. C. Domingos, V . A. de Santiago J ´unior, J. A. Anochi, E. H. Shiguemori, L. M. C. dos Santos, H. C. dos Santos Pereira, and A. E. C. Oliveira, “Exploring machine learning, deep learning, and explainable ai methods for seasonal precipitation prediction in south america,” arXiv preprint, 2025, submitted for publication. [Online]. Available: https://arxi...

  4. [4]

    Forecasting vapor pressure deficit for agricultural water management using machine learning in semi-arid environments,

    A. Elbeltagi, A. Srivastava, J. Deng, Z. Li, A. Raza, L. Khadke, Z. Yu, and M. El-Rawy, “Forecasting vapor pressure deficit for agricultural water management using machine learning in semi-arid environments,”Agricultural Water Management, vol. 283, p. 108302, 2023. [Online]. Available: https://doi.org/10.1016/j.agwat.2023.108302

  5. [5]

    Comparison of machine learning algorithms,

    C. Fischer, “Comparison of machine learning algorithms,” Bachelor’s thesis, FH Campus Wien, 2025. [Online]. Available: https://pub.hcw.ac.at/obvfcwhs/content/titleinfo/ 12219798

  6. [6]

    Learning to for- get: Continual prediction with lstm,

    F. Gers, J. Schmidhuber, and F. Cummins, “Learning to for- get: Continual prediction with lstm,”Neural Computation, vol. 12, pp. 2451–2471, 10 2000

  7. [7]

    ERA5 hourly data on single levels from 1940 to present,

    H. Hersbach, B. Bell, P. Berrisford, G. Biavati, A. Hor ´anyi, J. Mu˜noz Sabater, J. Nicolas, C. Peubey, R. Radu, I. Rozum, D. Schepers, A. Simmons, C. Soci, D. Dee, and J.-N. Th´epaut, “ERA5 hourly data on single levels from 1940 to present,” ECMWF Copernicus Climate Data Store, 2023, available at: https://cds.climate.copernicus.eu/doi/10.24381/ cds.adbb...

  8. [8]

    The era5 global reanalysis,

    H. Hersbach, B. Bell, P. Berrisford, S. Hirahara, A. Hor ´anyi, J. Mu ˜noz-Sabater, J. Nicolas, C. Peubey, R. Radu, D. Schepers, A. Simmons, C. Soci, S. Abdalla, X. Abellan, G. Balsamo, P. Bechtold, G. Biavati, J. Bidlot, M. Bonavita, G. De Chiara, P. Dahlgren, D. Dee, M. Diamantakis, R. Dragani, J. Flemming, R. Forbes, M. Fuentes, A. Geer, L. Haimberger,...

  9. [9]

    Long short-term mem- ory,

    S. Hochreiter and J. Schmidhuber, “Long short-term mem- ory,”Neural Computation, vol. 9, pp. 1735–1780, 11 1997

  10. [10]

    State-of-the-art in 1d convolu- tional neural networks: A survey,

    A. O. Ige and M. Sibiya, “State-of-the-art in 1d convolu- tional neural networks: A survey,”IEEE Access, vol. 12, pp. 144 397–144 419, 2024

  11. [11]

    Batch normalization: Accelerating deep network training by reducing internal covariate shift,

    S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” inProceedings of the 32nd International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 37. PMLR, 2015, pp. 448–456. [Online]. Available: https://proceedings.mlr.press/v37/ioffe15.html

  12. [12]

    Adam: A method for stochastic optimization,

    D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint, 2017, available at: https: //arxiv.org/abs/1412.6980 (accessed 15 Apr 2020). [Online]. Available: https://doi.org/10.48550/arXiv.1412.6980

  13. [13]

    Distributed systems and algorithms for measurement collection, decision making, and visualiza- tion of georeferenced information with applications in viticulture,

    S. Kontogiannis, “Distributed systems and algorithms for measurement collection, decision making, and visualiza- tion of georeferenced information with applications in viticulture,” Doctoral dissertation, Aristotle University of Thessaloniki, 2026, available at: https://ikee.lib.auth.gr/ record/368402/?ln=el (accessed 11 Mar 2026). [Online]. Available: ht...

  14. [14]

    Graphcast: Learning skillful medium-range global weather forecasting,

    R. Lam, A. Sanchez-Gonzalez, M. Willson, P. Wirnsberger, M. Fortunato, F. Alet, S. Ravuri, T. Ewalds, Z. Eaton-Rosen, W. Hu, A. Merose, S. Hoyer, G. Holland, O. Vinyals, J. Stott, A. Pritzel, S. Mohamed, and P. Battaglia, “Graphcast: Learning skillful medium-range global weather forecasting,”

  15. [15]

    Available: https://arxiv.org/abs/2212.12794

    [Online]. Available: https://arxiv.org/abs/2212.12794

  16. [16]

    PyTorch: An imperative style, high- performance deep learning library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. K ¨opf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An imperative style, high- performance deep learning library,” arXiv preprint, 2019, available at: https:/...

  17. [17]

    Scikit-learn: Ma- chine learning in Python,

    F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V . Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Ma- chine learning in Python,”Journal of Machine Learning Re- search, vol. 12, pp. 2825–2830, 2011

  18. [18]

    Daily prediction and multi-step forward forecasting of reference evapotranspiration using LSTM and Bi-LSTM models,

    D. K. Roy, T. K. Sarkar, S. S. A. Kamar, T. Goswami, M. A. Muktadir, H. M. Al-Ghobari, A. Alataway, A. Z. Dewidar, A. A. El-Shafei, and M. A. Mattar, “Daily prediction and multi-step forward forecasting of reference evapotranspiration using LSTM and Bi-LSTM models,” Agronomy, vol. 12, no. 3, p. 594, 2022. [Online]. Available: https://doi.org/10.3390/agron...

  19. [19]

    A deep learning based framework for enhanced reference evapotranspiration estimation: Evaluating accuracy and forecasting strategies,

    S. S. Sarkar, J. Bedi, and S. Jain, “A deep learning based framework for enhanced reference evapotranspiration estimation: Evaluating accuracy and forecasting strategies,” Scientific Reports, vol. 15, p. 15136, 2025. [Online]. Available: https://doi.org/10.1038/s41598-025-99713-2

  20. [20]

    Dropout: A simple way to prevent neural networks from overfitting,

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,”Journal of Machine Learning Research, vol. 15, no. 56, pp. 1929–1958, 2014. [Online]. Available: https://jmlr.org/papers/v15/srivastava14a.html 12

  21. [21]

    Water quality identification: Integrating iot sensors and deep learning for near-real-time water quality assessment,

    C. Tsolaki, G. Kokkonis, S. Valsamidis, and S. Kontogiannis, “Water quality identification: Integrating iot sensors and deep learning for near-real-time water quality assessment,” Applied Sciences, vol. 16, no. 10, 2026. [Online]. Available: https://www.mdpi.com/2076-3417/16/10/4868

  22. [22]

    Development of an in vivo sensor to monitor the effects of vapour pressure deficit (VPD) changes to improve water productivity in agriculture,

    F. Vurro, M. Janni, N. Copped `e, F. Gentile, R. Manfredi, M. Bettelli, and A. Zappettini, “Development of an in vivo sensor to monitor the effects of vapour pressure deficit (VPD) changes to improve water productivity in agriculture,”Sensors, vol. 19, no. 21, p. 4667, 2019. [Online]. Available: https://doi.org/10.3390/s19214667

  23. [23]

    Open-Meteo.com Weather API,

    P. Zippenfenig, “Open-Meteo.com Weather API,” Zenodo, 2023, available at: https://open-meteo.com/ (accessed 11 Apr 2024). [Online]. Available: https://doi.org/10.5281/ zenodo.7970649 13