Pith. sign in

REVIEW 3 major objections 3 minor 88 references

Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum

T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read In two-dimensional horizontal convection, combining a density extremum with a nonlinear equation of state reorganizes the flow into a single-roll circulation and, when mixing plumes reach full depth, enhances heat transport from the…

desk verdict The attached full text is a different paper (RadarQA), so this verdict rests entirely on the abstract: the density-extremum horizontal convection result is plausible and likely novel, but the central z-hat scaling argument is too under-supported to accept as established. read the letter →

arxiv 2508.12289 v4 pith:7LOBAJL3 submitted 2025-08-17 physics.flu-dyn

classification physics.flu-dyn
keywords horizontalconvectiondensityextremumnonlinearequationofstatemixingplumesheattransportscalingenergybudgetdirectnumericalsimulationRossby
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether the standard Oberbeck-Boussinesq picture of horizontal convection survives when the fluid's density is not a linear function of temperature, as in water near $4^\circ\mathrm{C}$. Using two-dimensional direct numerical simulations from $Ra=10^6$ to $5\times 10^{10}$, it compares linear and nonlinear equations of state with monotonic and density-extremum boundary conditions, and finds that only the extremum-plus-nonlinear case reorganizes the flow: a bicellular pattern gives way to a single roll driven by central mixing plumes. When those plumes penetrate the full cavity depth, heat transport follows a steeper scaling, $Nu \sim Ra^{1/4}$ up to $Nu \sim Ra^{1/3}$, instead of the usual Rossby $Nu \sim Ra^{1/5}$. The paper traces this to an additional potential-energy transfer term in the total energy budget that the nonlinear equation of state introduces, whose size is set by the plume height. If correct, standard energy closures under-predict heat transport in density-extremum horizontal convection once plumes reach the full depth.

What carries the argument

The load-bearing object is $\Phi_{i2}$, an additional potential-energy transfer term appearing in the total energy budget when the equation of state is nonlinear, alongside the usual Oberbeck-Boussinesq terms. Its magnitude is argued to scale with the characteristic mixing-plume height $\hat{z}$; full-depth plumes ($\hat{z} \sim H$) make the standard OB horizontal-convection kinetic-energy dissipation closure incomplete, which is how the paper moves the heat-transport exponent from $1/5$ to $1/4$–$1/3$.

What would settle it

Run the EXT-NELT configuration in a taller cavity, or at a higher Rayleigh number, where $\hat{z}$ no longer reaches $H$; if $Nu$ still follows a $Ra^{1/4}$–$Ra^{1/3}$ scaling, the plume-depth link fails. Conversely, a laboratory experiment with water near $4^\circ\mathrm{C}$ that directly measures plume penetration depth and Nusselt number would test whether $\hat{z} \sim H$ is necessary for the enhanced scaling.

Watch

Extended reading notes

Core claim

The central claim is that the nonlinearity of the equation of state, not the density-extremum boundary condition alone, is what changes horizontal convection. In the EXT-NELT configuration, the large-scale circulation shifts from two cells to a single roll, with central mixing plumes carrying the transport; when these plumes span the whole height $H$ ($\hat{z} \sim H$), the Nusselt number grows as $Ra^{1/4}$ to $Ra^{1/3}$, distinctly steeper than the Rossby $Ra^{1/5}$ seen in the other three configurations. The paper's mechanism is a new term $\Phi_{i2}$ in the global kinetic-energy balance, a potential-energy transfer produced by the nonlinear equation of state, whose magnitude is controlled by $\hat{z}$. With $\hat{z} \sim H$, this term changes the dissipation closure, and a scaling model built from it reproduces the main trends of the DNS data.

Load-bearing premise

The explanation assumes that the extra energy term $\Phi_{i2}$ is controlled by the plume height $\hat{z}$ measured from the same simulations whose heat-transport scaling it is then used to reproduce, so the mechanism is calibrated to the data rather than independently predicting them.

Editorial extensions

If this is right

  • For density-extremum horizontal convection with full-depth plumes, the standard Oberbeck-Boussinesq closure under-predicts heat transport; global energy budgets for such flows should include $\Phi_{i2}$.
  • The flow structure itself changes: the bicellular circulation becomes a single-roll, central-plume regime, with transitional anomalies in the Reynolds-number scaling.
  • The heat-transport exponent is not fixed: it can lie between $1/4$ and $1/3$ depending on plume depth, so a single power law may not describe the full parameter range.
  • Models of lakes, ice-ocean settings, or industrial systems where water near $4^\circ\mathrm{C}$ is the working fluid should not assume the Rossby $1/5$ scaling once plumes become depth-penetrating.
  • The energy-budget model offers a diagnostic route: computing $\Phi_{i2}$ and $\hat{z}$ from DNS or experiments could predict when enhanced transport begins.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the plume-height control holds, the transition from $Ra^{1/5}$ to $Ra^{1/3}$ should be continuous in $\hat{z}/H$; a testable prediction is that the effective Nusselt exponent is a monotone function of plume penetration depth.
  • The single-roll reorganization suggests that three-dimensional simulations may show a qualitatively different large-scale flow, since two-dimensional central plumes are often sensitive to confinement; the scaling claim should be checked in three dimensions.
  • One can use the $\Phi_{i2}$ budget term to design controlled experiments: changing the temperature of the cold boundary relative to $4^\circ\mathrm{C}$ should tune $\hat{z}$ and produce a predictable shift in the Nusselt scaling.
  • Analogous potential-energy transfer terms may appear for other non-Oberbeck-Boussinesq nonlinearities, such as salinity or compressibility effects, so the same energy-budget analysis could identify enhanced transport in those settings.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript as identified by its title and abstract reports a two-dimensional direct numerical simulation study of horizontal convection with a nonlinear equation of state exhibiting a density extremum near 4°C. The abstract claims that in the EXT-NELT configuration the flow reorganizes from a bicellular structure into a single-roll circulation driven by full-depth mixing plumes, that heat transport is enhanced from the Rossby scaling Nu ~ Ra^{1/5} to Nu ~ Ra^{1/4}–Ra^{1/3}, and that this enhancement arises from an additional potential-energy transfer term Phi_{i2} introduced by the nonlinear equation of state, with the term's magnitude controlled by the characteristic plume height z-hat. The submitted full text, however, is not this paper: it is an unrelated manuscript on multi-modal quality analysis of weather radar forecasts (RadarQA, arXiv:2508.12291), containing no governing equations, simulation setup, energy-budget analysis, or numerical results for horizontal convection. The abstract-level claims are therefore the only reviewable content of the physics submission.

Significance. If the abstract's claims were fully substantiated, the result would be of genuine interest to the horizontal-convection and geophysical fluid dynamics communities: it would indicate that the standard Oberbeck-Boussinesq energy closure underpredicts heat transport in density-extremum configurations once plumes penetrate the full cavity depth, and it would identify a specific energy-budget term, Phi_{i2}, responsible for the change. The reported reorganization from a bicellular to a single-roll circulation is a plausible and potentially impactful observation for flows near the density maximum of water. However, the paper as submitted offers no machine-checked proofs, reproducible code, derivations, figures, or data tables for any of these claims; the actual manuscript body is a different work. The only quantitative statement available, that the model 'captures the main trends of the numerical data,' is too weak to constitute validation, and the claimed exponent range 1/4–1/3 is presented without fitting statistics or a stated Ra window.

major comments (3)
  1. [Full text] The body of the submission is a different manuscript: the entire 'Full Text' section is the RadarQA paper on weather radar forecast quality analysis, with different authors, different subject matter, and no equations or results relating to horizontal convection, the density extremum, Phi_{i2}, z-hat, or Nu-Ra scaling. None of the central claims of the title and abstract can be checked because the supporting derivation, DNS setup, boundary conditions, equation-of-state form, and data-analysis procedure are absent. This is a load-bearing defect: the manuscript is internally inconsistent and cannot be reviewed as a physics paper.
  2. [Abstract, second paragraph] The explanatory mechanism appears circular: the new potential-energy transfer term Phi_{i2} is asserted to be controlled by the characteristic plume height z-hat, and z-hat is diagnosed from the same EXT-NELT DNS runs whose heat-transport scaling the model then reproduces. The abstract states this only as a 'scaling argument' with 'suggests,' and no independent derivation of the Phi_{i2}(z-hat) relation is provided. As presented, the model is calibrated to the data it explains rather than offering an out-of-sample prediction; a test on a different Rayleigh-number range or an alternative equation-of-state parametrization would be needed to establish the claimed 1/4-to-1/3 exponent.
  3. [Abstract, first paragraph] The claim of 'enhanced heat transport scaling ranging from Nu ~ Ra^{1/4} to Nu ~ Ra^{1/3}' spans two distinct power-law exponents, and the abstract does not say whether this is a single fitted exponent that drifts with Rayleigh number, a crossover between two asymptotic regimes, or a fit with substantial uncertainty. Without the figures and fitting details that would appear in the full text, the reader cannot assess whether the data support a new dissipation closure or merely a transitional, non-asymptotic effect.
minor comments (3)
  1. [Abstract, second paragraph] The term Phi_{i2} is introduced without a definition or an equation number, so its physical content is not accessible from the abstract alone.
  2. [Abstract, first paragraph] The abbreviations EXT, MON, LENT, and NELT are used without expansion, and the nonlinear equation-of-state form is never stated.
  3. [Full text] The author list and reference list of the full text belong to the RadarQA manuscript, which is unrelated to the title and abstract of the submission; this mismatch should be resolved by the authors before any further review.

Circularity Check

1 steps flagged · score 5.0 of 10

The Phi_i2-vs-z-hat scaling argument is calibrated to the same DNS data whose Nu exponent it explains.

  1. fitted input called prediction [Abstract, second paragraph (energy-budget discussion of Phi_i2 and z-hat)]
    "The scaling argument suggests that the magnitude of this contribution is controlled by the characteristic plume height ($\hat{z}$). Specifically, when plumes penetrate the entire cavity depth ($\hat{z} \sim H$), as observed in the EXT-NELT case, the global kinetic energy dissipation is no longer described by the standard OB HC energy closure alone. The resulting model captures the main trends of the numerical data and provides a possible energy budget interpretation of the enhanced transport observed in this two-dimensional configuration."

    The explanatory variable z-hat is diagnosed from the same DNS runs whose heat-transport scaling the model then reproduces. The magnitude of the new energy term Phi_i2 is said to be controlled by z-hat, and the full-depth condition z-hat ~ H is itself 'observed in the EXT-NELT case' from the same simulations. Thus the claimed agreement between the resulting model and the numerical data is a consistency check rather than an independent prediction; the scaling argument is calibrated to the data it explains. The phrase 'captures the main trends of the numerical data' reinforces that the model is fitted to the observed Nu(Ra) behavior rather than predicting it from first principles.

full rationale

The central empirical finding—that EXT-NELT reorganizes into a single roll and exhibits enhanced Nu scaling—is a DNS observation and is not circular by itself. No self-citation chains or imported uniqueness theorems appear in the abstract. However, the proposed energy-budget mechanism is not independent: Phi_i2 is stated to be controlled by z-hat, and z-hat is taken from the same simulations whose Nu trend is then reproduced. The abstract's own wording ('scaling argument suggests', 'as observed in the EXT-NELT case', 'model captures the main trends of the numerical data') shows that the explanatory step uses diagnosed flow quantities rather than an a priori prediction. The Nu~Ra^{1/4} to Nu~Ra^{1/3} range also spans two power laws, consistent with a non-asymptotic transition rather than a new dissipation closure, though that concern is about correctness rather than circularity. Because the load-bearing control variable is internal to the data being explained, the mechanistic account is partially circular, warranting a score of 5. The supplied full text is an unrelated RadarQA manuscript, so no additional physics equations could be checked beyond the abstract.

Assumptions & free parameters 3 free parameters · 4 assumptions · 1 invented entities

The abstract-level audit shows the central claim rests on a reference scaling law (Rossby Ra^{1/5}), a modeling choice for the nonlinear equation of state, the two-dimensionality of the DNS, and a new energy term Phi_{i2} whose stated control parameter is the plume height measured in the same simulations. The heat-transport exponent and its prefactor are effectively fitted over the Ra window, and z-hat is a data-derived input to the interpretive model. No genuinely independent falsifiable handle is offered for Phi_{i2} in the abstract.

free parameters (3)
  • Heat transport scaling exponent gamma = 1/4 to 1/3 for EXT-NELT
    The central claim is a fitted power law Nu ~ Ra^gamma over the window 10^6 to 5 x 10^10; the abstract reports a range rather than a converged exponent with uncertainty.
  • Scaling prefactor C in Nu = C Ra^gamma = not stated
    Any fitted power law carries a prefactor; it is not reported in the abstract.
  • Characteristic plume height z-hat = up to ~H (full cavity depth) in EXT-NELT
    The energy budget model takes the plume penetration height from the same DNS fields that produce the Nu data; the model's output therefore tracks an input measured from the data.
assumptions (4)
  • domain assumption Standard Oberbeck-Boussinesq horizontal convection energy closure and Rossby scaling Nu ~ Ra^{1/5} are the correct reference behavior for the MON/LENT cases.
    The paper's core contrast is against this baseline, which it imports from prior literature; the abstract gives no derivation of the baseline's validity for the simulated window.
  • domain assumption Two-dimensional DNS captures the reorganization and scaling that matter for this configuration.
    All results are 2D; the abstract itself restricts the enhanced-transport claim to "this two-dimensional configuration," so any extension to 3D is an extrapolation.
  • domain assumption A specific nonlinear equation of state approximating the 4°C density extremum (functional form and parameters not given in the abstract) drives the EXT-NELT results.
    The magnitude of Phi_{i2} and the position of the extremum relative to the thermal boundary conditions depend on this choice, which is not stated in the abstract.
  • ad hoc to paper The new potential-energy transfer term Phi_{i2} is controlled by the characteristic plume height z-hat, and z-hat ~ H in the EXT-NELT case.
    This is the load-bearing link of the interpretation; the abstract introduces it as "the scaling argument suggests," and the plume height is measured from the same simulations whose scaling it explains.
invented entities (1)
  • Phi_{i2}, an additional potential-energy transfer term from the nonlinear equation of state
    purpose: Adds a channel of potential-to-kinetic energy transfer that modifies the global dissipation balance when plumes reach the full cavity depth, supporting the enhanced Nu scaling.
    The term is diagnosed and its stated control parameter, plume height, is measured in the same simulations that establish the scaling, so the abstract offers no independent falsifiable handle; it is a ledger entry within the interpretation, not a testable entity outside the simulation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum." pith.science (2026). https://pith.science/paper/7LOBAJL3

@misc{pith2026250812289,
  author       = {Pith},
  title        = {Pith review of: Flow reorganization and transport enhancement in two-dimensional horizontal convection near a density extremum},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7LOBAJL3}},
  note         = {Machine review of arXiv:2508.12289}
}
abstract

Horizontal convection serves as a canonical model for geophysical and industrial flows. While the Oberbeck-Boussinesq approximation is well established, the impact of a nonlinear equation of state, specifically the density extremum of water near $4^\circ\mathrm{C}$, remains underexplored. Here we investigate this effect using two-dimensional direct numerical simulations over the Rayleigh number range $10^6 \le Ra \le 5\times 10^{10}$. We examine four configurations, contrasting extremum (EXT) and monotonic (MON) buoyancy boundary conditions against linear (LENT) and nonlinear (NELT) equations of state. Our results reveal that the EXT-NELT case undergoes a pronounced reorganization of the large-scale flow, evolving from a bicellular structure to a single-roll circulation driven by central `mixing plumes'. This reorganization manifests as transitional anomalies in the $Re$ scaling, while the emergence of full-depth plumes alters the heat transport mechanism. Consequently, distinct from the Rossby scaling ($Nu \sim Ra^{1/5}$) observed in the reference cases, the EXT-NELT case exhibits an enhanced heat transport scaling ranging from $Nu \sim Ra^{1/4}$ to $Nu \sim Ra^{1/3}$. To interpret this behaviour, we examine the total energy budget and identify an additional potential-energy transfer term, \(\Phi_{i2}\), arising from the nonlinear equation of state. The scaling argument suggests that the magnitude of this contribution is controlled by the characteristic plume height ($\hat{z}$). Specifically, when plumes penetrate the entire cavity depth ($\hat{z} \sim H$), as observed in the EXT-NELT case, the global kinetic energy dissipation is no longer described by the standard OB HC energy closure alone. The resulting model captures the main trends of the numerical data and provides a possible energy budget interpretation of the enhanced transport observed in this two-dimensional configuration.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

88 extracted references · 25 canonical work pages

  1. [1]

    Agrawal, K

    H. Agrawal, K. Desai, Y . Wang, X. Chen, R. Jain, M. Johnson, D. Batra, D. Parikh, S. Lee, and P. Anderson. Nocaps: Novel object captioning at scale. In ICCV, 2019

  2. [2]

    Antol, A

    S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh. VQA: Visual question answering. In ICCV, 2015

  3. [3]

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025

  4. [4]

    Banerjee and A

    S. Banerjee and A. Lavie. METEOR: An automatic metric for mt evaluation with improved correlation with human judgments. In ACL Workshops, 2005

  5. [5]

    K. Bi, L. Xie, H. Zhang, X. Chen, X. Gu, and Q. Tian. Pangu-weather: A 3d high-resolution model for fast and accurate global weather forecast. arXiv preprint arXiv:2211.02556, 2022

  6. [6]

    X. Bi, D. Chen, G. Chen, S. Chen, D. Dai, C. Deng, H. Ding, K. Dong, Q. Du, Z. Fu, et al. Deepseek LLM: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024

  7. [7]

    J. Chen, P. Zhou, Y . Hua, D. Chong, M. Cao, Y . Li, Z. Yuan, B. Zhu, and J. Liang. Vision-language models meet meteorology: Developing models for extreme weather events detection with heatmaps. arXiv preprint arXiv:2406.09838, 2024

  8. [8]

    K. Chen, T. Han, J. Gong, L. Bai, F. Ling, J.-J. Luo, X. Chen, L. Ma, T. Zhang, R. Su, et al. Fengwu: Pushing the skillful global medium-range weather forecast beyond 10 days lead. arXiv preprint arXiv:2304.02948, 2023

Show all 88 references
  1. [9]

    X. Chen, H. Fang, T.-Y . Lin, R. Vedantam, S. Gupta, P. Dollár, and C. L. Zitnick. Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325, 2015

  2. [10]

    Z. Chen, W. Wang, Y . Cao, Y . Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271, 2024

  3. [11]

    Davis, B

    C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part i: Methodology and application to mesoscale rain areas. Monthly Weather Review, 134(7):1772–1784, 2006

  4. [12]

    Davis, B

    C. Davis, B. Brown, and R. Bullock. Object-based verification of precipitation forecasts. part ii: Application to convective rain systems. Monthly Weather Review, 134(7):1785–1795, 2006

  5. [13]

    Donaldson, R

    R. Donaldson, R. M. Dyer, and M. J. Kraus. An objective evaluator of techniques for predicting severe weather events. Preprints, Ninth Conf. on Severe Local Storms, Norman, OK, Amer. Meteor. Soc, 1975

  6. [14]

    J. P. Finley. Tornado predictions. American Meteorological Journal. A Monthly Review of Meteorology and Allied Branches of Study (1884-1896), 1884

  7. [15]

    Y . Gao, H. Wu, R. Shu, H. Dong, F. Xu, R. Chen, Y . Yan, Q. Wen, X. Hu, K. Wang, et al. Oneforecast: A universal framework for global and regional weather forecasting. arXiv preprint arXiv:2502.00338, 2025

  8. [16]

    Z. Gao, X. Shi, H. Wang, Y . Zhu, Y . B. Wang, M. Li, and D.-Y . Yeung. EarthFormer: Exploring space-time transformers for earth system forecasting. In NeurIPS, 2022

  9. [17]

    Z. Gao, C. Tan, L. Wu, and S. Z. Li. Simvp: Simpler yet better video prediction. In CVPR, 2022. 10

  10. [18]

    Q. Ge, W. Sun, Y . Zhang, Y . Li, Z. Ji, F. Sun, S. Jui, X. Min, and G. Zhai. LMM-VQA: Advancing video quality assessment with large multimodal models. arXiv preprint arXiv:2408.14008, 2024

  11. [19]

    T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Zhang, D. Rojas, G. Feng, H. Zhao, et al. ChatGLM: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024

  12. [20]

    F. Gofa, D. Boucouvala, P. Louka, and H. Flocas. Spatial verification approaches as a tool to evaluate the performance of high resolution precipitation forecasts. Atmospheric Research, 2018

  13. [21]

    J. Gong, L. Bai, P. Ye, W. Xu, N. Liu, J. Dai, X. Yang, and W. Ouyang. Cascast: Skillful high-resolution precipitation nowcasting via cascaded modelling. arXiv preprint arXiv:2402.04290, 2024

  14. [22]

    Goyal, T

    Y . Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh. Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR, 2017

  15. [23]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  16. [24]

    X. He, Z. Zhou, W. Zhang, X. Zhao, H. Chen, S. Chen, and L. Bai. DiffSR: Learning radar reflectivity synthesis via diffusion model from satellite observations. In ICASSP, 2025

  17. [25]

    K. A. Hilburn, I. Ebert-Uphoff, and S. D. Miller. Development and interpretation of a neural-network-based synthetic radar reflectivity estimator using goes-r satellite observations. Journal of Applied Meteorology and Climatology, 2020

  18. [26]

    E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, W. Chen, et al. Lora: Low-rank adaptation of large language models. In ICLR, 2022

  19. [27]

    Hurst, A

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. GPT-4o system card. arXiv preprint arXiv:2410.21276, 2024

  20. [28]

    Jaech, A

    A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  21. [29]

    Z. Jia, Z. Zhang, J. Qian, H. Wu, W. Sun, C. Li, X. Liu, W. Lin, G. Zhai, and X. Min. VQA � : Visual question answering for video quality assessment. arXiv preprint arXiv:2411.03795, 2024

  22. [30]

    I. T. Jolliffe and D. B. Stephenson. Forecast verification: a practitioner’s guide in atmospheric science. John Wiley & Sons, 2012

  23. [31]

    B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024

  24. [32]

    F. Li, R. Zhang, H. Zhang, Y . Zhang, B. Li, W. Li, Z. Ma, and C. Li. Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895, 2024

  25. [33]

    W. Li, X. Zhang, S. Zhao, Y . Zhang, J. Li, L. Zhang, and J. Zhang. Q-Insight: Understanding image quality via visual reinforcement learning. arXiv preprint arXiv:2503.22679, 2025

  26. [34]

    B. Lin, Y . Ye, B. Zhu, J. Cui, M. Ning, P. Jin, and L. Yuan. Video-llava: Learning united visual representation by alignment before projection. arXiv preprint arXiv:2311.10122, 2023

  27. [35]

    C.-Y . Lin. Rouge: A package for automatic evaluation of summaries. In ACL, 2004

  28. [36]

    F. Liu, Y . Wang, T. Wang, and V . Ordonez. Visual news: Benchmark and challenges in news image captioning. arXiv preprint arXiv:2010.03743, 2020

  29. [37]

    H. Liu, C. Li, Q. Wu, and Y . J. Lee. Visual instruction tuning. In NeurIPS, 2023

  30. [38]

    Y . Liu, H. Duan, Y . Zhang, B. Li, S. Zhang, W. Zhao, Y . Yuan, J. Wang, C. He, Z. Liu, et al. MMBench: Is your multi-modal model an all-around player? In ECCV, 2024

  31. [39]

    H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525, 2024

  32. [40]

    C. Ma, Z. Hua, A. Anderson-Frey, V . Iyer, X. Liu, and L. Qin. WeatherQA: Can multimodal language models reason about severe weather? arXiv preprint arXiv:2406.11217, 2024. 11

  33. [41]

    Marino, M

    K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi. OK-VQA: A visual question answering benchmark requiring external knowledge. In CVPR, 2019

  34. [42]

    A. H. Murphy. The finley affair: A signal event in the history of forecast verification. Weather and forecasting, 1996

  35. [43]

    Palmer and R

    W. Palmer and R. Allen. Note on the accuracy of forecasts concerning the rain problem. US Weather Bureau, 1949

  36. [44]

    H. A. Panofsky and G. W. Brier. Some applications of statistics to meteorology . Mineral Industries Extension Services, College of Mineral Industries, Pennsylvania State University, 1958

  37. [45]

    Papineni, S

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002

  38. [46]

    Ravuri, K

    S. Ravuri, K. Lenc, M. Willson, D. Kangin, R. Lam, P. Mirowski, M. Fitzsimons, M. Athanassiadou, S. Kashem, S. Madge, et al. Skilful precipitation nowcasting using deep generative models of radar.Nature, 2021

  39. [47]

    Rempel, F

    M. Rempel, F. Senf, and H. Deneke. Object-based metrics for forecast verification of convective development with geostationary satellite data. Monthly Weather Review, 145(8):3161–3178, 2017

  40. [48]

    Robinson, J

    M. Robinson, J. Evans, and B. Crowe. En route weather depiction benefits of the nexrad vertically integrated liquid water product utilized by the corridor integrated weather system. In Conference on aviation, range and aerospace meteorology, american meteorological society, 2002

  41. [49]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017

  42. [50]

    Schwenk, A

    D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi. A-OKVQA: A benchmark for visual question answering using world knowledge. In ECCV, 2022

  43. [51]

    H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y . Liu, and H. Li. Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. In NeurIPS, 2024

  44. [52]

    Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y . Li, Y . Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024

  45. [53]

    D. B. Stephenson, B. Casati, C. Ferro, and C. Wilson. The extreme dependency score: A non-vanishing measure for forecasts of rare events. Meteorological Applications: A journal of forecasting, practical applications, training techniques and modelling, 2008

  46. [54]

    Stock, K

    J. Stock, K. Hilburn, I. Ebert-Uphoff, and C. Anderson. Srvit: Vision transformers for estimating radar reflectivity from satellite observations at scale. arXiv preprint arXiv:2406.16955, 2024

  47. [55]

    K. Sun, J. Pan, Y . Ge, H. Li, H. Duan, X. Wu, R. Zhang, A. Zhou, Z. Qin, Y . Wang, et al. Journeydb: A benchmark for generative image understanding. In NeurIPS, 2023

  48. [56]

    Touvron, T

    H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  49. [57]

    Veillette, S

    M. Veillette, S. Samsi, and C. Mattioli. SEVIR: A storm event imagery dataset for deep learning applications in radar and satellite meteorology. In NeurIPS, 2020

  50. [58]

    F. Wang, M. Chen, X. He, Y . Zhang, F. Liu, Z. Guo, Z. Hu, J. Wang, J. Xu, Z. Li, et al. Omniearth- bench: Towards holistic evaluation of earth’s six spheres and cross-spheres interactions with multimodal observational earth data. arXiv preprint arXiv:2505.23522, 2025

  51. [59]

    W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y . Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al. Cogvlm: Visual expert for pretrained language models. In NeurIPS, 2024

  52. [60]

    Y . Wang, M. Long, J. Wang, Z. Gao, and P. S. Yu. PredRNN: Recurrent neural networks for predictive learning using spatiotemporal lstms. In NeurIPS, 2017

  53. [61]

    Y . Wang, Y . Zeng, J. Zheng, X. Xing, J. Xu, and X. Xu. VideoCoT: A video chain-of-thought dataset with active annotation tool. arXiv preprint arXiv:2407.05355, 2024. 12

  54. [62]

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE TIP, 2004

  55. [63]

    C. J. Willmott and K. Matsuura. Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Climate research, 2005

  56. [64]

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, C. Li, W. Sun, Q. Yan, G. Zhai, et al. Q-Bench: A benchmark for general-purpose foundation models on low-level vision. arXiv preprint arXiv:2309.14181, 2023

  57. [65]

    H. Wu, Z. Zhang, E. Zhang, C. Chen, L. Liao, A. Wang, K. Xu, C. Li, J. Hou, G. Zhai, et al. Q-Instruct: Improving low-level visual abilities for multi-modality foundation models. In CVPR, 2024

  58. [66]

    H. Wu, H. Zhu, Z. Zhang, E. Zhang, C. Chen, L. Liao, C. Li, A. Wang, W. Sun, Q. Yan, et al. Towards open-ended visual quality comparison. In ECCV, 2024

  59. [67]

    Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y . Ma, C. Wu, B. Wang, et al. Deepseek- vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024

  60. [68]

    J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y . Fan, K. Dang, et al. Qwen2.5-omni technical report. arXiv preprint arXiv:2503.20215, 2025

  61. [69]

    A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024

  62. [70]

    J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou. mplug-owl3: Towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840, 2024

  63. [71]

    Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang. mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration. In CVPR, 2024

  64. [72]

    Z. You, J. Gu, Z. Li, X. Cai, K. Zhu, C. Dong, and T. Xue. Descriptive image quality assessment in the wild. arXiv preprint arXiv:2405.18842, 2024

  65. [73]

    Z. You, Z. Li, J. Gu, Z. Yin, T. Xue, and C. Dong. Depicting beyond scores: Advancing image quality assessment through multi-modal language models. In ECCV, 2024

  66. [74]

    Z. You, X. Cai, J. Gu, T. Xue, and C. Dong. Teaching large language models to regress accurate image quality scores using score distribution. In CVPR, 2025

  67. [75]

    Young, A

    P. Young, A. Lai, M. Hodosh, and J. Hockenmaier. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2014

  68. [76]

    D. Yu, X. Li, Y . Ye, B. Zhang, C. Luo, K. Dai, R. Wang, and X. Chen. Diffcast: A unified framework via residual diffusion for precipitation nowcasting. In CVPR, 2024

  69. [77]

    Q. Yu, Z. Zhang, R. Zhu, Y . Yuan, X. Zuo, Y . Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025

  70. [78]

    Zhang, X

    P. Zhang, X. Dong, B. Wang, Y . Cao, C. Xu, L. Ouyang, Z. Zhao, H. Duan, S. Zhang, S. Ding, et al. Internlm- xcomposer: A vision-language large model for advanced text-image comprehension and composition. arXiv preprint arXiv:2309.15112, 2023

  71. [79]

    Zhang, V

    T. Zhang, V . Kishore, F. Wu, K. Q. Weinberger, and Y . Artzi. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675, 2019

  72. [80]

    Zhang, H

    Y . Zhang, H. Yu, M. Zhang, Y . Yang, and Z. Meng. Uncertainties and error growth in forecasting the record-breaking rainfall in zhengzhou, henan on 19–20 july 2021. Science China Earth Sciences, 2022

  73. [81]

    Zhang, M

    Y . Zhang, M. Long, K. Chen, L. Xing, R. Jin, M. I. Jordan, and J. Wang. Skilful nowcasting of extreme precipitation with nowcastnet. Nature, 2023

  74. [82]

    Zhang, Z

    Z. Zhang, Z. Jia, H. Wu, C. Li, Z. Chen, Y . Zhou, W. Sun, X. Liu, X. Min, W. Lin, et al. Q-Bench-Video: Benchmarking the video quality understanding of lmms. arXiv preprint arXiv:2409.20063, 2024

  75. [83]

    Zhang, H

    Z. Zhang, H. Wu, Y . Zhou, C. Li, W. Sun, C. Chen, X. Min, X. Liu, W. Lin, and G. Zhai. LMM-PCQA: Assisting point cloud quality assessment with LMM. In ACM MM, 2024. 13

  76. [84]

    Zhang, T

    Z. Zhang, T. Kou, S. Wang, C. Li, W. Sun, W. Wang, X. Li, Z. Wang, X. Cao, X. Min, et al. Q-Eval-100K: Evaluating visual quality and alignment level for text-to-vision content. arXiv preprint arXiv:2503.02357, 2025

  77. [85]

    X. Zhao, W. Xu, B. Liu, Y . Zhou, F. Ling, B. Fei, X. Yue, L. Bai, W. Zhang, and X.-M. Wu. Msearth: A benchmark for multimodal scientific comprehension of earth science. arXiv preprint arXiv:2505.20740, 2025

  78. [86]

    Zhong, Z

    Q. Zhong, Z. Sun, H. Chen, J. Li, and L. Shen. Multi model forecast biases of the diurnal variations of intense rainfall in the beijing-tianjin-hebei region. Science China Earth Sciences, 2022

  79. [87]

    M. Zhou, J. Wu, M. Chen, and L. Han. Comparative study on the performance of convlstm and convgru in classification problems—taking early warning of short-duration heavy rainfall as an example. Atmospheric and Oceanic Science Letters, 2024

  80. [88]

    high value retain

    Y . Zhou, Y . Wang, X. He, R. Xiao, Z. Li, Q. Feng, Z. Guo, Y . Yang, H. Wu, W. Huang, et al. Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning. arXiv preprint arXiv:2506.10521, 2025. 14 Appendix A Overview This Appendix i...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.