REVIEW 5 major objections 5 minor 47 references
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Feeding extreme-value return levels into a weighted benchmark yields forward-looking resilience rankings of solar plants, and the paper applies this to two Portuguese PV sites.
desk verdict The central scaling equation is dimensionally invalid, so the forward-looking benchmark and the B2>B1 ranking rest on an undefined quantity; the EVT-correlation pipeline is clearly described but needs a corrected probability and sensitivity analysis before it supports conclusions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are (1) the Dynamic Identification of Significant Correlation (DISC)-Thresholding algorithm, which computes $\Delta\rho_{ij} = \rho_{ij}^{\text{extreme}} - \rho_{ij}$ for every variable pair and flags changes above the 90th percentile or below the 10th percentile as High Positive or High Negative Correlation, and (2) the EVA-Driven Weighted Benchmarking score $B_i = S_j \sum_i w_i C_{\text{extreme}}(V_i)$, where $S_j = E_c \times P(X>x)$ combines the normalized historical extreme-event frequency with the return-level exceedance probability, and $w_i$ are user-supplied weights. The first component tells the analyst which associations between weather variables and production actually change during extreme-low-production days; the second converts those extreme-condition statistics, scaled by projected return values and their confidence intervals, into a single comparable score per plant.
What would settle it
Recompute $B_i$ for both plants under several defensible weight vectors (e.g., equal weights, weights proportional to each variable's absolute correlation with production, or leave-one-out weights) and check whether Joao remains the higher-scoring plant; if the ordering flips, the paper's resilience conclusion collapses. A complementary out-of-sample check is to hold out the most recent years and test whether the plant with the higher EVDBM score actually exhibits more frequent or deeper low-production extremes in the holdout period.
Extended reading notes
Core claim
The central claim is that benchmarking scores can be made forward-looking by multiplying a historical frequency factor $E_c$ by an exceedance probability $P(X>x)$ derived from the fitted extreme value distribution, and then weighting the extreme-condition statistics of related variables by pre-assigned importance weights $w_i$. In the two-plant photovoltaic study, this procedure assigns plant Joao (B2) a higher benchmarking score than plant Zarco (B1), which the paper interprets as evidence that Joao is more sensitive to adverse conditions and less resilient. The paper states this directly in Section 4.3: "Since B2 has the highest EVDBM score it is clear that B2 (joao plant) is more sensitive to adverse conditions and is less resilient to fluctuations in related circumstances." If correct, the method converts historical extremes plus projected return values into a vulnerability ranking that static historical benchmarks cannot provide.
Load-bearing premise
The final ranking rests on the manually preset variable weights in Table 8, which the paper calls "dummy weights" chosen "for testing purposes"; since the score is linear in those weights and no sensitivity analysis is given, a different reasonable weight set could reverse the conclusion that Joao is less resilient.
Editorial extensions
If this is right
- Plant operators could use EVDBM scores to rank facilities by projected vulnerability before a rare low-output event occurs, not just after reviewing historical performance.
- Because scores are computed from return values and their confidence intervals, the benchmark carries an explicit uncertainty range that widens with return period.
- The DISC algorithm flags which pairwise correlations between weather variables and production shift most during extreme-low days, giving a diagnostic list of variables to monitor.
- The same algorithmic pipeline can be reapplied to any dependent variable with extreme values and correlated covariates, making the score a generic benchmarking tool.
Reading between the lines
- A direct testable extension is a weight-sensitivity analysis: recomputing $B_i$ under equal weights, data-driven weights, or weights derived from regression coefficients would show whether the Joao-less-resilient ranking is robust or an artifact of the 'dummy' proportions in Table 8.
- The 90th/10th percentile thresholds in the DISC algorithm are preset; varying them would test whether the flagged high-correlation pairs, and therefore the variable weights used in benchmarking, are stable across threshold choices.
- The same pipeline transfers to other extreme-low events, such as financial drawdowns or hospital admissions, where a return level plays the role of low PV production; the paper lists these as intended future applications but does not run them.
- Because the paper acknowledges stationarity as only partially addressed, a natural next step is to make the exceedance probability time-dependent with climate or market covariates, turning the benchmark into an early-warning indicator rather than a static ranking.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes the Extreme Value Dynamic Benchmarking Method (EVDBM), a three-stage pipeline that (i) applies extreme value analysis (POT with a GPD fit) to a dependent variable, (ii) uses a novel 'DISC-Thresholding' algorithm to flag variables whose Pearson correlations change between extreme and normal periods, and (iii) computes a weighted benchmark score B_i for each use case, scaled by a factor S_j that combines historical extreme-event frequency with EVA return levels. The method is demonstrated on hourly production data from two Portuguese PV plants (Zarco and Joao), with the claims that the scores are 'forward-looking' and that the Joao plant (B2) is less resilient to adverse conditions. The paper also lists limitations concerning data quality, stationarity, and weight subjectivity.
Significance. If the central calculation were valid, the EVDBM idea would be a useful addition to the EVA benchmarking literature: it explicitly couples tail-risk estimation with multi-criteria performance comparison and is demonstrated on real operational data from two PV plants. The DISC-Thresholding procedure is simple and transparent, and the authors are honest about several limitations. However, as the manuscript stands, the load-bearing scaling factor in Eq. (6) is not a probability, the GPD fits are not documented, and the final ranking depends on unexamined 'dummy' weights. These issues prevent the results from supporting the paper's main claims, though they are, in principle, addressable in a major revision.
major comments (5)
- [Section 3.3, Eq. (6)] The quantity P_j(X > x) = (1/T_return) × x(T)_CI is not a probability: the right-hand side has units of production per time (e.g., kWh/year if return levels are in kWh), it can be negative or exceed 1, and the text's own example (1/T = 0.2) ignores the multiplicative return level. Because S_j = E_c × P_j in Eq. (7) is the only mechanism that makes the benchmark 'forward-looking,' the adjusted scores in Figure 8 and the B2 > B1 conclusion are undefined as written.
- [Section 3.3, Step 1 items (a)-(c) and Eq. (9)] The algorithm treats the return value and the two confidence-interval endpoints as three 'probabilities' P_j(X > x0), P(X > x_lower), and P(X > x_upper), but no transformation to a dimensionless quantity in [0,1] is specified. Consequently, B_i = S_j × Σ b(V_i) mixes a dimensionless weighted score with this ill-defined factor, and the three curves in Figure 8 have no clear probabilistic interpretation.
- [Section 4.1.2 / 4.2.2, Tables 3 and 6] The GPD fits are not documented with parameter estimates, threshold-selection details, or proper diagnostic checks; the reported R² = 0.997 and p-value = 0.000 are not defined, and p-values cannot be exactly zero. In addition, the confidence-interval columns are internally inconsistent: for Zarco's 1-year return value of 0.35, the reported 'Lower CI' is 1.66 and 'Upper CI' is 0.11, so the lower bound exceeds the upper bound. These values feed directly into Step 1(b)-(c), making the benchmark inputs unreliable.
- [Section 4.3, Table 8] The variable weights in Table 8 are explicitly described as 'dummy' and chosen 'for testing purposes,' and the final ranking B2 > B1 is linear in these weights. No sensitivity analysis is provided, and the paper's own limitation section acknowledges that weighting subjectivity can bias comparisons. Therefore the statement in Section 4.3 that the Joao plant 'is more sensitive to adverse conditions and is less resilient' is not robust to reasonable alternative weight choices.
- [Section 3.2 / Section 4.1.3] The DISC-Thresholding algorithm classifies correlation changes as 'significant' based only on 10th/90th percentiles of the empirical distribution of Δρ_ij, without any hypothesis test, confidence interval, or measure of sampling variability. The term 'significant' is therefore not statistically justified, and the identified HPC/HNC pairs in Figure 6 should be interpreted as data-driven threshold exceedances rather than significance findings.
minor comments (5)
- [Section 4 and Table 1] The time range is inconsistent: the text in Section 4 states a 3-hour window from 13:00 to 16:00, while Table 1 lists 'Time-Range 13:00-14:00'.
- [Section 3.3] The equation numbering is confused: Step 2 labels b(V_i) = w_i × C_extreme(V_i) as Eq. (7), but Step 3 references Eqs. (6), (7), and (8) before introducing Eq. (9); Eq. (8) is never defined.
- [Section 2.1] The L-kurtosis and L-skewness formulas are duplicated verbatim in consecutive paragraphs, which disrupts the presentation.
- [Section 4.1.2 and Table 3] The text says that the return values are 'mostly negative,' but the values in Table 3 and Table 6 are all positive; this discrepancy should be clarified.
- [Appendix] The abbreviation 'HNN' is used in Figure 6 and in Section 4.1.3, but the abbreviation list defines 'HNC' as 'High Negative Correlation'; the notation should be unified.
Circularity Check
The forward-looking benchmark is a rescaled fitted return value, so the resilience ranking is forced by Eqs. (6)-(9) rather than independently predicted.
-
fitted input called prediction
[Section 3.3, Eqs. (6), (7), (9); Section 4.3]
"The probability of exceeding a threshold P (X > x0) accounts for the future risk and can be estimated as follows: Pj(X > x) = 1/Treturn × x(T )CI (6) ... Sj = Ec × P (X > x) (7) ... Bi = Sj × Pn i=1 b(Vi) (9). Since B2 has the highest EVDBM score it is clear that B2 (joao plant) is more sensitive to adverse conditions and is less resilient."
Eq. (6) defines the 'probability of exceeding a threshold' as the fitted return value x(T) multiplied by 1/T. Thus P_j is not an independent probability; it is a rescaled return level. Because S_j = E_c × P and B_i = S_j × Σ w_i C_extreme(V_i), the final benchmarking score is, by construction, proportional to the GPD-fitted return value from the same historical extremes that define the extreme conditions. The conclusion that Joao is less resilient is then forced by the fitted return values: Joao's Table 6 return values (e.g., 1.05 at 1-year) exceed Zarco's Table 3 values (0.35), so the 'probability' and hence B2 are larger for every confidence-interval case.
full rationale
The derivation chain is mostly self-contained: the EVA fitting, DISC-Thresholding, and data processing are standard and not circular. The self-citations ([5], [6], [28]) are contextual and not load-bearing; no uniqueness theorem or ansatz is smuggled in via citation. The central circularity is in the benchmarking step: Eq. (6) defines P_j(X > x) as (1/T_return) × x(T)_CI, where x(T) is the return value fitted from the same historical low-production extremes used to define the extreme dataset. Consequently, the scaling factor S_j (Eq. 7) and the final benchmarking score B_i (Eq. 9) are proportional to that fitted return value. The 'forward-looking' adjustment therefore has no content beyond the fitted return level, and the claim that B2/Joao is less resilient is a direct arithmetic consequence of Joao's larger fitted return values in Table 6 compared with Zarco's in Table 3, multiplied by the manually preset 'dummy' weights in Table 8. The paper acknowledges weight subjectivity in its Limitations, but the more fundamental issue is that the alleged exceedance probability is not a probability but a rescaled return value, so the ranking is forced by construction rather than independently predicted. This is partial circularity: one of the paper's central 'predictions' reduces to a fitted parameter renamed as a probability, while the rest of the methodology remains independent and reproducible.
Assumptions & free parameters
free parameters (4)
- Percentile threshold for extreme-low production =
25%
- Time window =
13:00-16:00
- DISC threshold percentiles =
P90 = 90th, P10 = 10th
- Variable weights w_i =
[0.05, 0.1, 0.2, 0.05, 0.4, 0.2]
assumptions (4)
- standard math Fisher-Tippett-Gnedenko theorem and Pickands-Balkema-De Haan theorem justify GPD/POT fits
- domain assumption Stationarity of the extreme-value process
- domain assumption Pearson correlation adequately captures dependence between weather variables and production
- domain assumption Threshold exceedances form a Poisson process and are approximately independent
Cite this review
Pith. "Pith review of Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis." pith.science (2026). https://pith.science/paper/KYQG45I3
@misc{pith2026241113608,
author = {Pith},
title = {Pith review of: Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/KYQG45I3}},
note = {Machine review of arXiv:2411.13608}
}
read the original abstract
This paper presents an innovative approach to Extreme Value Analysis (EVA) by introducing the Extreme Value Dynamic Benchmarking Method (EVDBM). EVDBM integrates extreme value theory to detect extreme events and is coupled with the novel Dynamic Identification of Significant Correlation (DISC)-Thresholding algorithm, which enhances the analysis of key variables under extreme conditions. By integrating return values predicted through EVA into the benchmarking scores, we are able to transform these scores to reflect anticipated conditions more accurately. This provides a more precise picture of how each case is projected to unfold under extreme conditions. As a result, the adjusted scores offer a forward-looking perspective, highlighting potential vulnerabilities and resilience factors for each case in a way that static historical data alone cannot capture. By incorporating both historical and probabilistic elements, the EVDBM algorithm provides a comprehensive benchmarking framework that is adaptable to a range of scenarios and contexts. The methodology is applied to real PV data, revealing critical low - production scenarios and significant correlations between variables, which aid in risk management, infrastructure design, and long-term planning, while also allowing for the comparison of different production plants. The flexibility of EVDBM suggests its potential for broader applications in other sectors where decision-making sensitivity is crucial, offering valuable insights to improve outcomes.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
S. Y. Novak, Extreme value methods with applications to finance, Mono- graphs on Statistics and Applied Probability 122 (2011) 22
work page 2011
-
[2]
F. Longin, Extreme events in finance: A handbook of extreme value theory and its applications, John Wiley & Sons, 2016
work page 2016
- [3]
-
[4]
Y. Chiu, F. Chebana, B. Abdous, D. B´ elanger, P. Gosselin, Mortal- ity and morbidity peaks modeling: An extreme value theory approach, Statistical methods in medical research 27 (5) (2018) 1498–1512
work page 2018
-
[5]
D. P. Panagoulias, D. N. Sotiropoulos, G. A. Tsihrintzis, An extreme value analysis-based systemic approach in healthcare information sys- tems: The case of dietary intake, Electronics 12 (1) (2022) 204
work page 2022
-
[6]
D. P. Panagoulias, D. N. Sotiropoulos, G. A. Tsihrintzis, Extreme value analysis for dietary intake based on weight class, in: 2022 13th Interna- tional Conference on Information, Intelligence, Systems & Applications (IISA), IEEE, 2022, pp. 1–7
work page 2022
-
[7]
A. Ghorbel, A. Trabelsi, Energy portfolio risk management using time- varying extreme value copula methods, Economic Modelling 38 (2014) 470–485
work page 2014
- [8]
Show all 47 references
-
[9]
Doherty, M
K. Doherty, M. Folley, T. Whittaker, R. Doherty, Extreme value analysis of wave energy converters, in: ISOPE International Ocean and Polar Engineering Conference, ISOPE, 2011, pp. ISOPE–I
2011
-
[10]
Arsenault, Y
E. Arsenault, Y. Wang, M. P. Chapman, Towards scalable risk anal- ysis for stochastic systems using extreme value theory, arXiv preprint arXiv:2203.12689 (2022)
2022 arXiv
-
[11]
Szigeti, T
M. Szigeti, T. Ferenci, L. Kov´ acs, The use of block maxima method of extreme value statistics to characterise blood glucose curves, in: 2020 IEEE 15th International Conference of System of Systems Engineering (SoSE), IEEE, 2020, pp. 433–438
2020
-
[12]
H. Chen, T. Zhao, Modeling power loss during blackouts in china us- ing non-stationary generalized extreme value distribution, Energy 195 (2020) 117044
2020
-
[13]
Westerlund, W
P. Westerlund, W. Naim, Extreme value analysis of power system data, in: ITISE 2019-International Conference on Time Series and Forecast- ing, 25-27 September 2019 Granada (Spain), Vol. 1, 2019, pp. 322–327. 29
2019
-
[14]
Ahmed, J
H. Ahmed, J. H. Einmahl, C. Zhou, Extreme value statistics in semi- supervised models, Journal of the American Statistical Association (2024) 1–14
2024
-
[15]
Gilleland, M
E. Gilleland, M. Ribatet, A. G. Stephenson, A software review for ex- treme value analysis, Extremes 16 (2013) 103–119
2013
-
[16]
Makkonen, M
L. Makkonen, M. Tikanm¨ aki, An improved method of extreme value analysis, Journal of Hydrology X 2 (2019) 100012
2019
-
[17]
Benstock, F
D. Benstock, F. Cegla, Extreme value analysis (eva) of inspection data and its uncertainties, Ndt & E international 87 (2017) 68–77
2017
-
[18]
Coles, J
S. Coles, J. Bawa, L. Trenner, P. Dorazio, An introduction to statistical modeling of extreme values, Vol. 208, Springer, 2001
2001
-
[19]
Xu, Proceedings of 2013 World Agricultural Outlook Conference, Springer, 2014
S. Xu, Proceedings of 2013 World Agricultural Outlook Conference, Springer, 2014
2013
-
[20]
G. A. Tsihrintzis, C. L. Nikias, Fast estimation of the parameters of alpha-stable impulsive interference, IEEE transactions on signal pro- cessing 44 (6) (1996) 1492–1503
1996
-
[21]
Jebli, F.-Z
I. Jebli, F.-Z. Belouadha, M. I. Kabbaj, A. Tilioua, Prediction of solar energy guided by pearson correlation using machine learning, Energy 224 (2021) 120109
2021
-
[22]
Benesty, J
J. Benesty, J. Chen, Y. Huang, On the importance of the pearson cor- relation coefficient in noise reduction, IEEE Transactions on Audio, Speech, and Language Processing 16 (4) (2008) 757–765
2008
-
[23]
Sedgwick, Pearson’s correlation coefficient, Bmj 345 (2012)
P. Sedgwick, Pearson’s correlation coefficient, Bmj 345 (2012)
2012
-
[24]
Carrino, Data versus survey-based normalisation in a multidimen- sional analysis of social inclusion, Italian Economic Journal 2 (2016) 305–345
L. Carrino, Data versus survey-based normalisation in a multidimen- sional analysis of social inclusion, Italian Economic Journal 2 (2016) 305–345
2016
-
[25]
Vafaei, R
N. Vafaei, R. A. Ribeiro, L. M. Camarinha-Matos, Data normalisation techniques in decision making: case study with topsis method, Interna- tional journal of information and decision sciences 10 (1) (2018) 19–38. 30
2018
-
[26]
A. A. Hancock, E. N. Bush, D. Stanisic, J. J. Kyncl, C. T. Lin, Data normalization before statistical analysis: keeping the horse before the cart, Trends in pharmacological sciences 9 (1) (1988) 29–32
1988
-
[27]
H. Abdi, L. J. Williams, et al., Normalizing data, Encyclopedia of re- search design 1 (2010) 935–938
2010
-
[28]
D. P. Panagoulias, E. Sarmas, V. Marinakis, G. A. Tsihrintzis, Cir- cumstance evaluation using extreme value analysis on charging station data: The case of dei blue in greece, in: International Conference on In- formation, Intelligence, Systems, and Applications, Springer, 202...
2023
-
[29]
Sarmas, E
E. Sarmas, E. Spiliotis, E. Stamatopoulos, V. Marinakis, H. Doukas, Short-term photovoltaic power forecasting using meta-learning and nu- merical weather prediction independent long short-term memory mod- els, Renewable Energy 216 (2023) 118997
2023
-
[30]
Sarmas, N
E. Sarmas, N. Dimitropoulos, V. Marinakis, Z. Mylona, H. Doukas, Transfer learning strategies for solar power forecasting under data scarcity, Scientific Reports 12 (1) (2022) 1–13
2022
-
[31]
Sarmas, M
E. Sarmas, M. Kleideri, N. Matias, C. Pereira, A. R. Antunes, Photo- voltaic power production dataset (2024). doi:10.17632/dbh93b6vp8.2
2024 doi
-
[32]
Ilias, E
L. Ilias, E. Sarmas, V. Marinakis, D. Askounis, H. Doukas, Unsupervised domain adaptation methods for photovoltaic power forecasting, Applied Soft Computing 149 (2023) 110979
2023
-
[33]
G. C. Hegerl, S. Br¨ onnimann, A. Schurer, T. Cowan, The early 20th century warming: Anomalies, causes, and consequences, Wiley Interdis- ciplinary Reviews: Climate Change 9 (4) (2018) e522
2018
-
[34]
Shukla, Predictability in the midst of chaos: A scientific basis for climate forecasting, science 282 (5389) (1998) 728–731
J. Shukla, Predictability in the midst of chaos: A scientific basis for climate forecasting, science 282 (5389) (1998) 728–731
1998
-
[35]
Hilorme, O
T. Hilorme, O. Zamazii, O. Judina, R. Korolenko, Y. Melnikova, Forma- tion of risk mitigating strategies for the implementation of projects of energy saving technologies, Academy of Strategic Management Journal 18 (3) (2019) 1–6. 31
2019
-
[36]
Mishra, R
D. Mishra, R. Sharma, S. Kumar, R. Dubey, Bridging and buffering: Strategies for mitigating supply risk and improving supply chain per- formance, International Journal of Production Economics 180 (2016) 183–197
2016
-
[37]
Talluri, T
S. Talluri, T. J. Kull, H. Yildiz, J. Yoon, Assessing the efficiency of risk mitigation strategies in supply chains, Journal of Business logistics 34 (4) (2013) 253–269
2013
-
[38]
Truong, S
C. Truong, S. Tr¨ uck, It’s not now or never: Implications of investment timing and risk aversion on climate adaptation to extreme events, Eu- ropean Journal of Operational Research 253 (3) (2016) 856–868
2016
-
[39]
M. M. Elmassri, E. P. Harris, D. B. Carter, Accounting for strategic investment decision-making under extreme uncertainty, The British Ac- counting Review 48 (2) (2016) 151–168
2016
-
[40]
Harrison, Decision-making in conditions of extreme uncertainty, Jour- nal of management studies 14 (2) (1977) 169–178
F. Harrison, Decision-making in conditions of extreme uncertainty, Jour- nal of management studies 14 (2) (1977) 169–178
1977
-
[41]
Guthrie, Real options analysis of climate-change adaptation: invest- ment flexibility and extreme weather events, Climatic Change 156 (1) (2019) 231–253
G. Guthrie, Real options analysis of climate-change adaptation: invest- ment flexibility and extreme weather events, Climatic Change 156 (1) (2019) 231–253
2019
-
[42]
Pagano, I
A. Pagano, I. Pluchinotta, R. Giordano, A. B. Petrangeli, U. Fratino, M. Vurro, Dealing with uncertainty in decision-making for drinking wa- ter supply systems exposed to extreme events, Water Resources Man- agement 32 (2018) 2131–2145
2018
-
[43]
E. A. Lind, K. Van den Bos, When fairness works: Toward a general theory of uncertainty management, Research in organizational behavior 24 (2002) 181–223
2002
-
[44]
van Prooijen, J
A.-M. van Prooijen, J. Bartels, T. Meester, Communicated and at- tributed motives for sustainability initiatives in the energy industry: The role of regulatory compliance, Journal of Consumer Behaviour 20 (5) (2021) 1015–1024
2021
-
[45]
Thaler, V
P. Thaler, V. Pakalkaite, Governance through real-time compliance: the supranationalisation of european external energy policy, Journal of Eu- ropean public policy 28 (2) (2021) 208–228. 32
2021
-
[46]
S. C. Moser, J. A. Ekstrom, A framework to diagnose barriers to cli- mate change adaptation, Proceedings of the national academy of sci- ences 107 (51) (2010) 22026–22031
2010
-
[47]
G. R. Biesbroek, J. E. Klostermann, C. J. Termeer, P. Kabat, On the nature of barriers to climate change adaptation, Regional Environmental Change 13 (2013) 1119–1129. 33
2013
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.