REVIEW 2 major objections 7 minor 78 references
Beyond the Norm: A Survey of Synthetic Data Generation for Rare Events
T0 review · 2 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey argues that synthetic data generation for extreme events is a distinct field requiring its own extremeness-centered evaluation framework, not the privacy-oriented one used for general synthetic data.
desk verdict Useful map of a young field, but the AKE metric in the central evaluation framework is mis-specified and needs correction before the survey can be cited as a reliable reference. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two devices carry the argument. The first is the pairing of extreme value theory (EVT) with generative models: the Generalized Pareto Distribution from the Peaks Over Threshold method supplies tail-guided sampling, tail-aware losses, heavy-tailed latent priors, and spatial tail dependence structures that ordinary Gaussian-prior generators lack. The second is the evaluation framework itself, which separates general distributional similarity from extremeness-specific checks such as extreme coverage, extremal coefficient, extremal correlation, and tail-restricted mean squared logarithmic error. Together these devices turn 'does the synthetic data look real?' into 'does the synthetic data get the extremes right?'.
What would settle it
Compare a generic high-fidelity generator (for example, a standard GAN or diffusion model) against an EVT-enhanced generator on the same benchmark datasets using both global metrics (Wasserstein distance, FID) and extremeness-specific metrics (tail-restricted MSLE, extremal coefficient, coverage rate); if the generic model matches or beats the EVT-enhanced model on extremeness-specific metrics across multiple domains, the survey's central claim that extreme-event synthesis needs specialized methods and evaluation would be refuted.
Extended reading notes
Core claim
The paper's central claim is that synthetic extreme-event data differs fundamentally from ordinary synthetic data because its purpose is not to enable privacy-preserving sharing but to train models on rare, high-impact scenarios. It documents a body of methods that inject extreme value theory into generative models—through tail-guided sampling, tail-aware training objectives, heavy-tailed latent priors, and spatial tail dependence modeling—and organizes the field's benchmark datasets by data type, from precipitation grids and limit order books to river discharge and keystroke intervals. The distinctive contribution is the evaluation framework: eight categories (distributional similarity, dependence preservation, extreme coverage, extreme dependence, reconstruction loss, extreme magnitude accuracy, visualization diagnostics, downstream performance validation) with per-metric guidance on extremeness applicability and domain-specific adaptations. The survey also maps application domains and identifies underexplored areas such as behavioral finance, wildfires, windstorms, earthquakes, and infectious outbreaks.
Load-bearing premise
The load-bearing premise is that synthetic data generation for extreme events is a distinct subfield whose evaluation must center on extremeness rather than privacy or general fidelity; if ordinary metrics already capture tail quality, the survey's framework loses its reason for being.
Editorial extensions
If this is right
- Evaluation of rare-event generators should center on extremeness-specific metrics such as tail-restricted MSLE, extremal coefficient, extreme coverage, and domain-adapted FID rather than on global fidelity or privacy measures.
- Benchmark suites for this field should include heavy-tailed, multivariate, spatially dependent datasets with explicit extreme definitions, as the surveyed precipitation, financial, hydrological, and energy datasets do.
- EVT-enhanced generative models provide a reusable template—tail-guided sampling, tail-aware losses, heavy-tailed latent priors—that can be carried into domains the survey flags as underexplored, including earthquakes, wildfires, and infectious outbreaks.
- Downstream validation becomes a standard evaluation step: augmenting real data with synthetic extremes and measuring improvements in prediction or risk detection, as done in flood forecasting and financial market supervision.
- Domain-specific metric adaptations are necessary before generic metrics transfer, for example training an autoencoder on target data before computing FID for weather maps.
Reading between the lines
- A testable consequence of the framework is that generic high-fidelity generators will look strong on global distributional metrics yet underperform EVT-enhanced models on tail-restricted metrics and downstream extreme-event tasks; this can be checked directly on the surveyed benchmarks.
- The survey's clean split between privacy-oriented and extremeness-oriented synthesis is probably too clean: in healthcare and finance the two goals interact, since rare disease records and fraud cases are both sensitive and extreme, so future evaluation may need metrics that weigh both at once.
- The catalogued underexplored domains amount to a concrete research agenda: behavioral finance, wildfires, windstorms, earthquakes, and wide infectious outbreaks are the places where the next benchmark datasets and method comparisons will likely form.
- The eight-category evaluation framework could be turned into a practical scoring checklist or public leaderboard for rare-event generators, but the survey itself does not specify a consolidated scoring procedure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews synthetic data generation for extreme events, covering EVT-enhanced GANs and VAEs, importance sampling, diffusion models, hybrid architectures, domain-constrained generative models, and LLMs. It catalogs benchmark datasets by data type, proposes an evaluation framework organized into eight metric categories (distributional similarity, dependence preservation, extreme coverage, extremal dependence, reconstruction loss, extreme magnitude accuracy, visualization diagnostics, and downstream performance), and discusses application domains and open challenges. The paper claims to be the first comprehensive overview of synthetic data generation for extreme events, with a central contribution being the in-depth analysis of each metric's applicability to extremeness and its domain-specific adaptations.
Significance. If appropriately revised, this survey could serve as a useful structured entry point for researchers working on rare-event synthesis, consolidating a fragmented literature and providing a practical evaluation taxonomy. The dataset catalog (Tables 2 and 3) and the hands-on guidance in Section 6 are valuable, and the inclusion of a benchmark repository link supports reproducibility. The mathematical definitions of standard metrics (KS, KL, Wasserstein, FID, extremal coefficient) are mostly correct. However, the mis-specification of AKE in Eq. (11) undermines one component of the central evaluation framework and must be corrected before publication; this is an internal correctness issue, not a matter of field consensus.
major comments (2)
- [6.2.1, Eq. (11)] The definition of AKE is mis-specified. As printed, Eq. (11) computes the average absolute difference between sorted raw values of the real and synthetic datasets. That quantity is exactly the univariate 1-Wasserstein distance between the two empirical marginal distributions; it is invariant to any permutation of the samples and therefore cannot measure concordance, rank dependence, or tail co-movement. The claim that it 'corresponds to the 1-Wasserstein distance between the Kendall's dependence functions' requires the Z_i to be values of an empirical Kendall function (e.g., maxima of normalized ranks), but that transformation is absent from the definition. The same under-specification affects the cited uses in EV-GAN [2] and HTGAN [21]. Because the evaluation framework is the paper's central contribution and AKE is a listed dependence-preservation metric in Table 4, this error is load-bearing. The authors should provide a correct definition (e.g., based on the empirical Kendall distribution function) or, if the primary sources use a different metric, report that faithfully.
- [5.4, Table 2] The statement that "all tabular datasets in extreme data modeling are synthetically generated rather than derived from real-world sources" is contradicted by the paper's own dataset catalog. Table 2 lists the Market Supervision dataset [38] as a real-world dataset with over 20 features including transaction volume, market volatility, financial health, and regulatory disclosures, which appears to be tabular rather than a time series (it is placed under 'Finance Time Series' in the table). Either reclassify this dataset or qualify the claim about tabular data, since as written it is factually incorrect.
minor comments (7)
- [3.1] In the sentence after Eq. (1), the shape parameter is referred to as 'ε' but is denoted 'ξ' in the CDF; please use consistent notation.
- [6.1.2, Eq. (7)] The t-statistic formula is missing the bars on the sample means; the text defines \(\bar{x}\) and \(\bar{\tilde{x}}\), so the equation should read \((\bar{\tilde{x}} - \bar{x}) / \sqrt{s^2/n + \tilde{s}^2/m}\).
- [6.6.3, Eqs. (18) and (19)] Eq. (18) defines MSLE with \(\log(1+x_i)\) terms, while Eq. (19) uses \(\log x_i^{(j)}\) without the offset; clarify whether the "log(1+)" transformation is part of the definition or a separate tail-adapted variant.
- [1] The claim that synthetic data generation for extreme events 'differs fundamentally' from conventional approaches is asserted rather than demonstrated. Several metrics in Section 6 (KS, KL, Wasserstein, FID) are standard in general synthetic-data evaluation; the distinctiveness lies in tail-focused adaptations, not the metrics themselves. A more nuanced framing would strengthen the motivation.
- [6.2.1] Even after correcting Eq. (11), the statement that AKE is 'robust to outliers and invariant to marginal transformations' will need support: this holds for rank-based versions but not for raw-value versions.
- [5.2.3] There is a duplicated phrase in the text: 'supporting for stress testing stress testing of grid reliability'.
- [7.1] The text says 'as shown in Table 3' when referring to underexplored areas such as behavioral anomalies and DeFi disruptions; the relevant material appears in Table 5, not Table 3.
Circularity Check
No significant circularity: the survey's contributions are taxonomic and bibliographic, with no derivation or fitted prediction that reduces to its own inputs.
full rationale
This paper is a survey, not a derivation chain: it does not fit parameters to data, derive predictions from first principles, or claim a mathematical result built on its own conclusions. The central contributions are a taxonomy of generative methods, a catalog of benchmarks, and an evaluation framework assembled from established statistics and from explicit citations to external works. The AKE definition in Eq. (11) is internally questionable—using raw sorted values makes it a 1-Wasserstein distance on marginal quantiles rather than a Kendall-dependence measure as claimed—but that is a correctness defect, not circularity, because the metric is not equivalent to an input by construction and the survey does not use AKE to justify its own existence or conclusions. The self-citations (e.g., [14,25,26,27,28]) appear only as background examples of GAN applications and financial modeling; they are not load-bearing for the survey's stated contributions. The premise that extreme-event synthesis requires an extremeness-focused evaluation framework is a scope assumption, not a circular derivation. Therefore, no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (1)
- domain assumption The surveyed papers are accurately represented in this review.
Cite this review
Pith. "Pith review of Beyond the Norm: A Survey of Synthetic Data Generation for Rare Events." pith.science (2026). https://pith.science/paper/CK3AFJ6B
@misc{pith2026250606380,
author = {Pith},
title = {Pith review of: Beyond the Norm: A Survey of Synthetic Data Generation for Rare Events},
year = {2026},
howpublished = {\url{https://pith.science/paper/CK3AFJ6B}},
note = {Machine review of arXiv:2506.06380}
}
read the original abstract
Extreme events, such as market crashes, natural disasters, and pandemics, are rare but catastrophic, often triggering cascading failures across interconnected systems. Accurate prediction and early warning can help minimize losses and improve preparedness. While data-driven methods offer powerful capabilities for extreme event modeling, they require abundant training data, yet extreme event data is inherently scarce, creating a fundamental challenge. Synthetic data generation has emerged as a powerful solution. However, existing surveys focus on general data with privacy preservation emphasis, rather than extreme events' unique performance requirements. This survey provides the first overview of synthetic data generation for extreme events. We systematically review generative modeling techniques and large language models, particularly those enhanced by statistical theory as well as specialized training and sampling mechanisms to capture heavy-tailed distributions. We summarize benchmark datasets and introduce a tailored evaluation framework covering statistical, dependence, visual, and task-oriented metrics. A central contribution is our in-depth analysis of each metric's applicability in extremeness and domain-specific adaptations, providing actionable guidance for model evaluation in extreme settings. We categorize key application domains and identify underexplored areas like behavioral finance, wildfires, earthquakes, windstorms, and infectious outbreaks. Finally, we outline open challenges, providing a structured foundation for advancing synthetic rare-event research.
Figures
Reference graph
Works this paper leans on
-
[2]
Michaël Allouche, Stéphane Girard, and Emmanuel Gobet. 2022. EV-GAN: Simulation of extreme events with ReLU neural networks. Journal of Machine Learning Research 23, 150 (2022), 1–39
work page 2022
-
[21]
Stéphane Girard, Emmanuel Gobet, and Jean Pachebat. 2024. Deep generative modeling of multivariate dependent extremes. (2024)
work page 2024
-
[38]
Mohan Jiang, Yaxin Liang, Siyuan Han, Kunyuan Ma, Yuan Chen, and Zhen Xu. 2024. Leveraging Generative Adversarial Networks for Addressing Data Imbalance in Financial Market Supervision. arXiv preprint arXiv:2412.15222 (2024)
arXiv 2024
-
[1]
Hervé Abdi. 2007. The Kendall rank correlation coefficient. Encyclopedia of measurement and statistics (2007)
work page 2007
-
[3]
Martin Arjovsky, Soumith Chintala, and Léon Bottou. 2017. Wasserstein generative adversarial networks. In International conference on machine learning. PMLR, 214–223
work page 2017
-
[4]
Samuel A Assefa, Danial Dervovic, Mahmoud Mahfouz, Robert E Tillman, Prashant Reddy, and Manuela Veloso. 2020. Generating synthetic data in finance: opportunities, challenges and pitfalls. In Proceedings of the First ACM International Conference on AI in Finance
work page 2020
-
[5]
Azizjon Azimi, Bonu Boboeva, Ilyas Varshavskiy, Shuhrat Khalilbekov, Akhlitdin Nizamitdinov, Najima Noyoftova, and Sergey Shulgin. 2024. zGAN: An Outlier-focused Generative Adversarial Network For Realistic Synthetic Data Generation. arXiv preprint arXiv:2410.20808 (2024)
work page Pith review arXiv 2024
-
[6]
Seth Bassetti, Brian Hutchinson, Claudia Tebaldi, and Ben Kravitz. 2024. DiffESM: Conditional emulation of temperature and precipitation in Earth system models with 3D diffusion models. Journal of Advances in Modeling Earth Systems (2024)
work page 2024
Show all 78 references
-
[7]
Siddharth Bhatia, Arjit Jain, and Bryan Hooi. 2021. Exgan: Adversarial generation of extreme samples. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 6750–6758
2021
-
[8]
Younes Boulaguiem, Jakob Zscheischler, Edoardo Vignotto, Karin van der Wiel, and Sebastian Engelke. 2022. Modeling and simulating spatial extremes by combining extreme value theory with generative adversarial networks. Environmental Data Science 1 (2022), e5
2022
-
[9]
Enrique Castillo and Ali S Hadi. 1997. Fitting the generalized Pareto distribution to data. J. Amer. Statist. Assoc. (1997)
1997
-
[10]
Ling Chen. 2024. Risk Management with Feature-Enriched Generative Adversarial Networks (FE-GAN). arXiv preprint arXiv:2411.15519 (2024)
2024 arXiv
-
[11]
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009. Pearson correlation coefficient. Noise reduction in speech processing (2009), 1–4
2009
-
[12]
Rama Cont, Mihai Cucuringu, Renyuan Xu, and Chao Zhang. 2022. Tail-gan: Learning to simulate tail risk scenarios. arXiv preprint arXiv:2203.01664 (2022)
2022 arXiv
-
[13]
Jonathan D Cryer. 2008. Time series analysis. Springer
2008
-
[14]
Ankan Dash, Jingyi Gu, and Guiling Wang. 2024. HI-GAN: Hierarchical Inpainting GAN with Auxiliary Inputs for Combined RGB and Depth Inpainting. arXiv preprint arXiv:2402.10334 (2024). Manuscript submitted to ACM 34 Jingyi Gu, Xuan Zhang, and Guiling Wang
2024 arXiv
-
[15]
Anthony C Davison, Simone A Padoan, and Mathieu Ribatet. 2012. Statistical modeling of spatial extremes. (2012)
2012
-
[16]
Vivek Dhakal, Anna Maria Feit, Per Ola Kristensson, and Antti Oulasvirta. 2018. Observations on typing from 136 million keystrokes. In Proceedings of the 2018 CHI conference on human factors in computing systems
2018
-
[17]
J-L Dufresne, M-A Foujols, Sébastien Denvil, Arnaud Caubel, Olivier Marti, Olivier Aumont, Yves Balkanski, Slimane Bekki, Hugo Bellenger, Rachid Benshila, et al. 2013. Climate change projections using the IPSL-CM5 Earth System Model: from CMIP3 to CMIP5. Climate dynamics 40 (2...
2013
-
[18]
Ana Ferreira and Laurens De Haan. 2015. On the block maxima method in extreme value theory: PWM estimators. (2015)
2015
-
[19]
Alvaro Figueira and Bruno Vaz. 2022. Survey on synthetic data generation, evaluation methods and GANs. Mathematics 10, 15 (2022), 2733
2022
-
[20]
Tim RL Fry. 1989. Univariate and multivariate Burr distributions: a survey. (1989)
1989
-
[22]
Aldren Gonzales, Guruprabha Guruswamy, and Scott R Smith. 2023. Synthetic data in health care: A narrative review. PLOS Digital Health (2023)
2023
-
[23]
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. Advances in neural information processing systems 27 (2014)
2014
-
[24]
Mahesh Kumar Goyal. 2023. Synthetic Data Revolutionizes Rare Disease Research: How Large Language Models and Generative AI are Overcoming Data Scarcity and Privacy Challenges. https://doi.org/10.17762/ijritcc.v11i11.11411
2023 doi
-
[25]
Jingyi Gu, Fadi P Deek, and Guiling Wang. 2023. Stock broad-index trend patterns learning via domain knowledge informed generative network. arXiv preprint arXiv:2302.14164 (2023)
2023 arXiv
-
[26]
Jingyi Gu, Wenlu Du, AM Muntasir Rahman, and Guiling Wang. 2023. Margin Trader: A Reinforcement Learning Framework for Portfolio Management with Margin and Constraints. In Proceedings of the Fourth ACM International Conference on AI in Finance . 610–618
2023
-
[27]
Jingyi Gu, Wenlu Du, and Guiling Wang. 2025. RAGIC: Risk-Aware Generative Framework for Stock Interval Construction. IEEE Transactions on Knowledge and Data Engineering (2025)
2025
-
[28]
Jingyi Gu, Sarvesh Shukla, Junyi Ye, Ajim Uddin, and Guiling Wang. 2023. Deep learning model with sentiment score and weekend effect in stock price prediction. SN Business & Economics 3, 7 (2023), 119
2023
-
[29]
Laurens Haan and Ana Ferreira. 2006. Extreme value theory: an introduction . Vol. 3. Springer
2006
-
[30]
Wilco Hazeleger, X Wang, Camiel Severijns, S Ştefănescu, R Bintanja, Andreas Sterl, Klaus Wyser, T Semmler, S Yang, B Van den Hurk, et al. 2012. EC-Earth V2. 2: description and validation of a new seamless earth system prediction model. Climate dynamics 39 (2012), 2611–2629
2012
-
[31]
Hans Hersbach, Bill Bell, Paul Berrisford, Shoji Hirahara, András Horányi, Joaquín Muñoz-Sabater, Julien Nicolas, Carole Peubey, Raluca Radu, Dinand Schepers, et al. 2020. The ERA5 global reanalysis. Quarterly journal of the royal meteorological society (2020)
2020
-
[32]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems (2017)
2017
-
[33]
Bruce M Hill. 1974. The rank-frequency form of Zipf’s law. J. Amer. Statist. Assoc. 69, 348 (1974), 1017–1026
1974
-
[34]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems (2020)
2020
-
[35]
Tao Hong, Pierre Pinson, Shu Fan, Hamidreza Zareipour, Alberto Troccoli, and Rob J Hyndman. 2016. Probabilistic energy forecasting: Global energy forecasting competition 2014 and beyond. , 896–913 pages
2016
-
[36]
GUO Hongxia, LI Yuan, CHEN Lingxuan, WANG Ziqiang, MA Qian, and LIU Yiming. 2023. An Improved Generative Adversarial Network for Extreme Scenarios Generation. In 2023 IEEE 7th Conference on Energy Internet and Energy System Integration (EI2) . IEEE, 1472–1477
2023
-
[37]
Todd Huster, Jeremy Cohen, Zinan Lin, Kevin Chan, Charles Kamhoua, Nandi O Leslie, Cho-Yu Jason Chiang, and Vyas Sekar. 2021. Pareto gan: Extending the representational power of gans to heavy-tailed distributions. In International Conference on Machine Learning . PMLR, 4523–4532
2021
-
[39]
Leskovec Jure, Lang Kevin, Dasgupta Anirban, and Mahoney Michael. 2008. Community structure in large networks: Natural cluster sizes and the absence of large well-defined clusters. Internet Mathematics (2008)
2008
-
[40]
Divas Karimanzira. 2024. Mass Conservative Time-Series GAN for Synthetic Extreme Flood-Event Generation: Impact on Probabilistic Forecasting Models. Stats 7, 3 (2024), 808–826
2024
-
[41]
Richard W Katz, Marc B Parlange, and Philippe Naveau. 2002. Statistics of extremes in hydrology. Advances in water resources (2002)
2002
-
[42]
Jennifer E Kay, Clara Deser, A Phillips, A Mai, Cecile Hannay, Gary Strand, Julie Michelle Arblaster, SC Bates, Gokhan Danabasoglu, James Edwards, et al. 2015. The Community Earth System Model (CESM) large ensemble project: A community resource for studying climate change in t...
2015
-
[43]
Tae Kyun Kim. 2015. T test as a parametric statistic. Korean journal of anesthesiology 68, 6 (2015), 540–546
2015
-
[44]
Diederik P Kingma, Max Welling, et al. 2013. Auto-encoding variational bayes
2013
-
[45]
Konstantin Klemmer, Sudipan Saha, Matthias Kahl, Tianlin Xu, and Xiao Xiang Zhu. 2021. Generative modeling of spatio-temporal weather patterns with extreme event conditioning. arXiv preprint arXiv:2104.12469 (2021)
2021 arXiv
-
[46]
Nicolas Lafon, Philippe Naveau, and Ronan Fablet. 2023. A VAE approach to sample multivariate extremes. arXiv preprint arXiv:2306.10987 (2023)
2023
-
[47]
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther. 2016. Autoencoding beyond pixels using a learned similarity metric. In International conference on machine learning . PMLR, 1558–1566
2016
-
[48]
Malcolm R Leadbetter. 1991. On a basis for ‘Peaks over Threshold’modeling. Statistics & Probability Letters (1991). Manuscript submitted to ACM Beyond the Norm: A Survey of Synthetic Data Generation for Rare Events 35
1991
-
[49]
Jure Leskovec, Kevin J Lang, Anirban Dasgupta, and Michael W Mahoney. 2009. Community structure in large networks: Natural cluster sizes and the absence of large well-defined clusters. Internet Mathematics (2009)
2009
-
[50]
Jun Li, Dingcheng Li, Ping Li, and Gennady Samorodnitsky. 2024. Generalized Pareto GAN: Generating Extremes of Distributions. In2024 International Joint Conference on Neural Networks (IJCNN) . IEEE, 1–8
2024
-
[51]
Yanchun Li, Peng Li, Tianmeng Yang, Zelong Chen, Xuan Song, Yumin Zhao, Jicheng Liu, and Wei Feng. 2023. A C-DCGAN-based method for generating extreme risk scenarios of high percentage new energy systems. In 2023 3rd International Conference on Intelligent Power and Systems (I...
2023
-
[52]
Zexiang Li. 2024. Improved Conditional Generative Adversarial Network-based Approach for Extreme Scenario Generation in Renewable Energy Sources. In 2024 IEEE 6th International Conference on Civil A viation Safety and Information Technology (ICCASIT) . IEEE, 1391–1398
2024
-
[53]
Andrzej Maćkiewicz and Waldemar Ratajczak. 1993. Principal components analysis (PCA). Computers & Geosciences (1993)
1993
-
[54]
John I Marden. 2004. Positions and QQ plots. Statist. Sci. (2004)
2004
-
[55]
Frank J Massey Jr. 1951. The Kolmogorov-Smirnov test for goodness of fit. Journal of the American statistical Association (1951)
1951
-
[56]
Michael Meiser and Ingo Zinnikus. 2024. A Survey on the Use of Synthetic Data for Enhancing Key Aspects of Trustworthy AI in the Energy Domain: Challenges and Opportunities. Energies 17, 9 (2024), 1992
2024
-
[57]
Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784 (2014)
2014 arXiv
-
[58]
Jannes Münchmeyer, Dino Bindi, Ulf Leser, and Frederik Tilmann. 2021. Earthquake magnitude and location estimation from real time seismic waveforms with a transformer network. Geophysical Journal International (2021)
2021
-
[59]
Alison Peard and Jim Hall. 2023. Combining deep generative models with extreme value theory for synthetic hazard simulation: a multivariate and spatially coherent approach. arXiv preprint arXiv:2311.18521 (2023)
2023 arXiv
-
[60]
Alexandra Puchko, Robert Link, Brian Hutchinson, Ben Kravitz, and Abigail Snyder. 2020. Deepclimgan: A high-resolution climate data generator. arXiv preprint arXiv:2011.11705 (2020)
2020 arXiv
-
[61]
Sarva T Pulla, Hakan Yasarer, and Lance D Yarbrough. 2024. Synthetic Time Series Data in Groundwater Analytics: Challenges, Insights, and Applications. Water 16, 7 (2024), 949
2024
-
[62]
Alec Radford, Luke Metz, and Soumith Chintala. 2015. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434 (2015)
2015 arXiv
-
[63]
Oliver Alvarado Rodriguez, Zhihui Du, Joseph Patchett, Fuhuan Li, and David A Bader. 2022. Arachne: An Arkouda package for large-scale graph analytics. In 2022 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE
2022
-
[64]
Erika Schiappapietra and John Douglas. 2020. Modelling the spatial correlation of earthquake ground motion: Insights from the literature, data from the 2016–2017 Central Italy earthquake sequence and ground-motion simulations. Earth-science reviews 203 (2020), 103139
2020
-
[65]
Richard L Smith. 1990. Max-stable processes and spatial extremes. Unpublished manuscript 205 (1990), 1–32
1990
-
[66]
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015. Learning structured output representation using deep conditional generative models. Advances in neural information processing systems 28 (2015)
2015
-
[67]
Casper Kaae Sønderby, Tapani Raiko, Lars Maaløe, Søren Kaae Sønderby, and Ole Winther. 2016. Ladder variational autoencoders. Advances in neural information processing systems 29 (2016)
2016
-
[68]
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition
2016
-
[69]
Jakub Tomczak and Max Welling. 2018. VAE with a VampPrior. InInternational conference on artificial intelligence and statistics . PMLR, 1214–1223
2018
-
[70]
Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research (2008)
2008
-
[71]
K Van der Wiel, N Wanders, FM Selten, and MFP Bierkens. 2019. Added value of large ensemble simulations for assessing extreme river discharge in a 2 C warmer world. Geophysical Research Letters (2019)
2019
-
[72]
Lifang Wang, Xiaodong Guo, Jianchao Zeng, and Yi Hong. 2010. Using gumbel copula and empirical marginal distribution in estimation of distribution algorithm. In Third International Workshop on Advanced Computational Intelligence . IEEE, 583–587
2010
-
[73]
M Waseem, N Mani, G Andiego, and M Usman. 2017. A review of criteria of fit for hydrological models.International Research Journal of Engineering and Technology (IRJET) 4, 11 (2017), 1765–1772
2017
-
[74]
Yuki Yamagishi, Kazumi Saito, Kazuro Hirahara, and Naonori Ueda. 2021. Spatio-temporal clustering of earthquakes based on distribution of magnitudes. Applied Network Science 6 (2021), 1–17
2021
-
[75]
Derong Yi, Mingfeng Yu, Qiang Wang, Hao Tian, Leibao Wang, Yongqian Yan, Chenghuang Wu, Bo Hu, and Chunyan Li. 2024. Method for wind–solar–load extreme scenario generation based on an improved InfoGAN. Applied Sciences (2024)
2024
-
[76]
Jinsung Yoon, Daniel Jarrett, and Mihaela Van der Schaar. 2019. Time-series generative adversarial networks. Advances in neural information processing systems 32 (2019)
2019
-
[77]
Guang Zhao, Xihaier Luo, Shinjae Yoo, and Nicole D Jackson. 2024. GAN-based Extreme Conditional Distribution Estimation for Renewable Energy Systems. In 2024 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm) . IEEE
2024
-
[78]
Jun-Yan Zhu, Philipp Krähenbühl, Eli Shechtman, and Alexei A Efros. 2016. Generative visual manipulation on the natural image manifold. In ECCV. Springer. Manuscript submitted to ACM
2016
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.