Pith. sign in

REVIEW 5 major objections 4 minor 33 references

PMU Data Feature Considerations for Realistic, Synthetic Data Generation

T0 review · 5 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Synthetic power-grid data should include real PMU flaws: outliers, dropouts, and ambient oscillations.

desk verdict Useful empirical catalog of PMU data quirks from the PJM dataset, but the realism claim is asserted, not tested. read the letter →

arxiv 1908.05244 v1 pith:FU43PU5A submitted 2019-08-14 eess.SY cs.SY

classification eess.SYcs.SY
keywords phasormeasurementunitsynchrophasorsyntheticdatagenerationPMUqualityanomaliesambientoscillationsmissingpowersystemmeasurements
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that synthetic phasor measurement unit (PMU) data generated from power-system simulations are unrealistically clean, and that researchers who test grid analytics on such data will draw conclusions that do not transfer to field conditions. Using 30-minute records from 123 field PMUs in a public dataset, it measures how often real data contain outliers, missing samples from dropped packets, noise, and persistent low-frequency ambient oscillations. It concludes that synthetic data generation should deliberately include these features, roughly 1% outliers, random dropout events, and ambient modes near 0.3 and 0.5 Hz, to match the statistical character of industry measurements. The paper is best read as an evidence-based recipe for making simulation data behave like real synchrophasor data, rather than a proposal of any specific generation algorithm.

What carries the argument

The load-bearing structure is a variance decomposition, $\sigma_{\Delta V_M}^2 = \sigma_{\Delta V}^2 + \sigma_{\eta}^2 + \sigma_e^2$, separating total measured voltage variability into grid dynamics, measurement noise, and data-anomaly components. Alongside this, the paper uses a dropout rate $\rho$ and a maximum gap size $\chi$ to quantify missing data, an autocorrelation function to confirm the nearly independent noise structure, and a matrix pencil modal analysis on moving windows to identify ambient low-frequency modes. These quantities convert the qualitative idea of realism into measurable targets that a synthetic data generator can aim to hit.

What would settle it

Collect an independent set of PMU records from a different utility or time period and compute the same statistics; if most PMUs show no missing samples, outlier fractions well below 1%, and no spectral peaks near 0.3 and 0.5 Hz, the paper's feature targets are not universal. Alternatively, generate synthetic data with and without the recommended features and compare each to a held-out field record on the same variability, dropout, and modal statistics: the recommendation stands only if the feature-injected version is measurably closer to the field record.

Watch

Extended reading notes

Core claim

Real PMU measurements from the public 123-PMU dataset are not clean signals: they contain measurement noise with signal-to-noise ratios around 41-47 dB, roughly 1% outlier samples in voltage magnitude and angle, frequent dropout events (38% of PMUs had at least one missing value, and ten PMUs had gaps longer than 3 seconds), and persistent electromechanical ambient modes near 0.3 and 0.5 Hz that appear in every analysis window. Simulation outputs typically lack all of these features. The paper therefore establishes that the variability in true PMU data can be decomposed into three additive components, grid dynamics, measurement noise, and data anomalies, and that a realistic synthetic generator should reproduce each component along with missing-data statistics and ambient modal content.

Load-bearing premise

The recommendations rest on one public 30-minute dataset from 123 PMUs on one grid; if that dataset is not typical of field PMU data, the specific outlier rates, dropout rates, and modal frequencies it prescribes would not generalize to other systems.

Editorial extensions

If this is right

  • Synthetic PMU benchmarks for oscillation detection, state estimation, or event classification should be stress-tested with injections of roughly 1% outliers and random missing samples, or their reported accuracy may not hold on field data.
  • A simulator that reproduces only base power-flow dynamics will understate measurement variability by omitting the noise and anomaly variance terms; adding them brings synthetic variability closer to observed values such as the roughly $10^{-4}$ average voltage variance reported.
  • The persistent 0.3 and 0.5 Hz modes characterize the interconnection studied; synthetic grids meant to emulate a real system can use such persistent modal content as an acceptance criterion for ambient dynamics.
  • Missing-data handling becomes part of the experimental pipeline: because 38% of PMUs had at least one gap, synthetic datasets without dropout cannot exercise the imputation and bad-data rejection logic that field engineers must run.
  • Researchers can use the paper's variance decomposition as a checklist for synthetic data validation, comparing each component separately rather than only the overall signal shape.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural but untested extension is to model dropout events as spatially correlated across nearby PMUs, since packet losses often share communication paths; the paper reports per-PMU rates only and does not address joint dropout patterns.
  • One could benchmark the realism claim directly by training analytics on feature-injected synthetic data and measuring their performance on held-out field PMU records versus analytics trained on clean simulation data; the paper does not run that end-to-end comparison.
  • The 0.3 and 0.5 Hz ambient modes are likely signatures of one interconnection, not universal constants; synthetic generators targeting other grids should derive their own modal content from local field data rather than reusing these frequencies.
  • The variance-decomposition targets could be turned into a generative test: draw synthetic noise to match the observed autocorrelation behavior (near zero after lag 1), add outliers and dropouts at the measured rates, and statistically compare the resulting distributions with field data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper analyzes features of PMU measurements from the public PJM dataset (123 PMUs, one 30-minute window) to inform the generation of realistic synthetic PMU data. It proposes a variance decomposition that adds an anomaly term to the standard noise model, quantifies grid-dynamics variability and SNR on selected segments, reports outlier percentages in voltage magnitude and angle, computes dropout rates and maximum gap sizes for missing samples, and identifies low-frequency ambient oscillation modes via matrix-pencil analysis. The paper concludes that synthetic PMU data should include data anomalies, ambient oscillation content, and random missing-data samples in order to improve realism.

Significance. If the descriptive statistics are taken at face value, the paper provides a useful feature catalog for synthetic PMU data generation, with concrete prevalence numbers for outliers and missing data from a real industry dataset. The use of a public dataset and the clear presentation of per-PMU statistics (e.g., dropout rates and gap sizes in Fig. 10) are strengths. However, the central claim that injecting these features improves realism is not directly tested: no synthetic data is generated and no comparison is made between synthetic data with and without the proposed features. Several load-bearing methodological details, including the outlier detection rule, the selection of the representative segment, and the modal-analysis significance criterion, are unspecified. The paper is therefore best viewed as a motivating empirical study rather than a validated methodology.

major comments (5)
  1. [Abstract and Section VI] The central conclusion, stated in the Abstract and in Section VI, is that including anomalies, ambient oscillations, and missing data samples 'helps to improve the realism' of synthetic PMU data. This claim is never tested in the manuscript: no synthetic dataset is generated, and no comparison is made between synthetic data with and without the proposed features against a defined measure of realism. Section VI repeats the recommendation as a conclusion rather than demonstrating it. To support the claim, the authors should either add a validation experiment (e.g., generate synthetic data with and without the features and compare distributional statistics or downstream-task performance) or explicitly reframe the contribution as a feature-prevalence catalog to guide future synthetic-data generation.
  2. [Section III.A (Fig. 2)] The segment selected for estimating grid-dynamics variability is described as 'observed to possess predominantly noiseless and error-free data samples.' This is a hand-picked, noiseless segment, so the resulting average variability of 10^-4 cannot be treated as a representative characterization of grid-dynamics variability without an objective selection criterion. The authors should either define how the segment was chosen or use the full distribution from all 123 PMUs (as in Fig. 4) and explain why the single selected segment is representative.
  3. [Section III.C (Fig. 8)] The outlier percentages in Fig. 8 are load-bearing for the recommendation to inject anomaly features, yet the outlier detection rule is not specified. The text states that samples are 'outside their local vicinities,' but the local window size, the outlier threshold, and the treatment of non-stationarity (including the angle-difference transformation (4)) are not defined. Without this algorithmic detail, the reported average of 1% and the maximum above 7% cannot be reproduced or meaningfully interpreted. Please specify the detection method and all parameters used.
  4. [Section V] The ambient-oscillation analysis uses a matrix-pencil method with a 10-second moving window and 5-second step, but the manuscript does not specify how modes are declared 'significant,' how many windows were analyzed, whether the frequency data come from one PMU or all 123 PMUs, or how the claimed persistent 0.3 Hz and 0.5 Hz modes are identified across windows. These details are necessary to support the strong claim that the two modes 'appear in all data windows.' Please provide the mode-selection criterion and, where possible, the distribution of detected modes across windows and PMUs.
  5. [Sections II and IV] All statistics in the paper are computed from a single 30-minute window of the PJM public PMU dataset. The manuscript implicitly assumes this dataset is representative of typical industry PMU data, and this assumption is load-bearing if the reported dropout rates, outlier percentages, and modal content are intended to set generation parameters for other systems. The authors should acknowledge this limitation explicitly and either include additional datasets or time periods or present the results as case-study statistics for this particular system.
minor comments (4)
  1. [Figures 11 and 12] The captions of Fig. 11 and Fig. 12 both read 'First-order, stationary, 30-min voltage angles,' which is inconsistent with the content: Fig. 11 shows frequency measurements and Fig. 12 shows mode-analysis results. Please correct the captions.
  2. [Section III.B] The conclusion that the noise signal has 'strong i.i.d' features is based on a single autocorrelation function from one noise signal and an average ACF of 0.06 over lags. This is not a statistically rigorous whiteness test; a Ljung-Box test or confidence bounds under the null hypothesis would be more appropriate, and the sample size used for the ACF should be stated.
  3. [Section IV] The dropout-rate and gap-size statistics would be more reproducible if the authors stated how missing samples were identified (e.g., detection of NaN or sentinel values such as 9999) and whether any preprocessing was applied to distinguish actual packet drops from bad data.
  4. [References] Several figures are credited to the first author's dissertation [13]; for reproducibility, the data-processing steps used to generate these figures should be described in the paper rather than deferred to an unpublished dissertation. Reference [25] should also include an access date for the online resource.

Circularity Check

0 steps flagged · score 2.0 of 10

No substantive circularity; empirical feature survey relies on external PJM data, with a minor self-citation that is not load-bearing.

full rationale

The paper's derivation chain is descriptive rather than inferential. Equation (1) is imported from an external source [26], and equation (2) is a definitional extension that adds an anomaly variance term sigma_e^2; no fitted parameter is later renamed as a prediction. All reported statistics (outlier percentages, dropout rates, gap sizes, and ambient modes) are computed directly from the public PJM dataset [25], an external benchmark, rather than from any parameter fitted by the authors. The self-citation [13] is used only to source illustrative figures (Figs. 1, 4, and 9), and those figures are not the load-bearing evidence for the paper's conclusions; the quantitative claims are computed from PJM data in the present paper. The central recommendation that synthetic data should include anomalies, missing samples, and ambient oscillation content is a design suggestion, not a derived prediction: the paper does not demonstrate that including these features improves realism, but that is a correctness/validation gap, not a circular reduction. Accordingly, no step in the paper reduces to its own inputs by construction, and the score reflects only the presence of a minor non-load-bearing self-citation.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The central claim rests on statistical observations from one public dataset, so no fitted model parameters are used. The main assumptions are about representativeness of the dataset and the validity of the statistical decompositions and modal analysis.

free parameters (4)
  • Outlier detection rule = unspecified
    Outlier percentages in Fig. 8 are computed with an undefined rule for 'outside their local vicinities', which determines the reported rates.
  • Median filter order = 90 samples
    Order-90 moving-window median filter used for noise extraction in Section III.B; a hand-chosen setting that affects the noise statistics.
  • Average variability window size = 5 samples
    Non-overlapping window of 5 samples used to compute the average variability value of 10^-4 in Section III.A; a hand-chosen window length.
  • Matrix pencil analysis window = 10 sec window, 5 sec step
    Parameters for modal analysis in Section V; these hand-chosen settings affect which modes are detected.
assumptions (4)
  • domain assumption PJM dataset is representative of industry PMU data
    All statistics are drawn from this single dataset; its representativeness is not discussed (Section II).
  • domain assumption Variance components add independently (Eq. 2)
    Equation (2) assumes anomaly variance sigma_e^2 adds to dynamics and noise variances with no interaction, an untested modeling choice.
  • domain assumption PMU measurement noise is approximately Gaussian with zero mean
    Section III.B uses zero-mean normal statistics despite citing [27] which questions the Gaussian assumption.
  • standard math Matrix pencil correctly identifies ambient modes
    The paper uses a standard modal analysis technique and interprets results without validation on this data.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PMU Data Feature Considerations for Realistic, Synthetic Data Generation." pith.science (2026). https://pith.science/paper/FU43PU5A

@misc{pith2026190805244,
  author       = {Pith},
  title        = {Pith review of: PMU Data Feature Considerations for Realistic, Synthetic Data Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FU43PU5A}},
  note         = {Machine review of arXiv:1908.05244}
}
read the original abstract

It is critical that the qualities and features of synthetically-generated, PMU measurements used for grid analysis matches those of measurements obtained from field-based PMUs. This ensures that analysis results generated by researchers during grid studies replicate those outcomes typically expected by engineers in real-life situations. In this paper, essential features associated with industry PMU-derived data measurements are analyzed for input considerations in the generation of vast amounts of synthetic power system data. Inherent variabilities in PMU data as a result of the random dynamics in power system operations, oscillatory contents, and the prevalence of bad data are presented. Statistical results show that in the generation of large datasets of synthetic, grid measurements, an inclusion of different data anomalies, ambient oscillation contents, and random cases of missing data samples due to packet drops helps to improve the realism of experimental data used in power systems analysis.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

33 extracted references · 33 canonical work pages

  1. [13]

    Data analytics and wide -area visualization associated with power systems using phasor measurements,

    I. Idehen, "Data analytics and wide -area visualization associated with power systems using phasor measurements," Ph.D dissertation, Dept. Elect. Eng., Texas A&M Univ., College Station, TX, USA, 2019

  2. [1]

    A. G. Phadke and J. S. Thorp, Synchronized phasor measurements and their applications. Springer, 2008

  3. [2]

    The Wide World of Wide -area Measurement,

    A. G. Phadke and R. M. d. Moraes, "The Wide World of Wide -area Measurement," IEEE Power and Energy Magazine, vol. 6, no. 5, pp. 52- 65, 2008

  4. [3]

    National Academies of Sciences and Medi cine, Analytic research foundations for the next -generation electric grid

    E. National Academies of Sciences and Medi cine, Analytic research foundations for the next -generation electric grid . National Academies Press, 2016

  5. [4]

    National Academies of Sciences and Medicine, Enhancing the Resilience of the Nation's Electricity System

    E. National Academies of Sciences and Medicine, Enhancing the Resilience of the Nation's Electricity System . National Academies Press, 2017

  6. [5]

    Definition and classification of power system stability IEEE/CIGRE joint task force on stability terms and definitions,

    P. Kundur et al., "Definition and classification of power system stability IEEE/CIGRE joint task force on stability terms and definitions," IEEE Transactions on Power Systems, vol. 19, no. 3, pp. 1387-1401, 2004

  7. [6]

    Toward a smart grid: power delivery for the 21st century,

    S. M. Amin and B. F. Wollenberg, "Toward a smart grid: power delivery for the 21st century," IEEE Power and Energy Magazine, vol. 3, no. 5, pp. 34-41, 2005

  8. [7]

    PMU Data Quality: A Framework for the Attributes of PMU Data Quality and Quality Impacts to Synchrophasor Applications,

    P. NASPI, "PMU Data Quality: A Framework for the Attributes of PMU Data Quality and Quality Impacts to Synchrophasor Applications," 2017

Show all 33 references
  1. [8]

    Characterizing and quantifying noise in PMU data,

    M. Brown, M. Biswal, S. Brahma, S. J. Ranade, and H. Cao, "Characterizing and quantifying noise in PMU data," in 2016 IEEE Power and Energy Society General Meeting (PESGM), 2016, pp. 1-5

  2. [9]

    The time skew problem in PMU measurements,

    Q. Zhang, V. Vittal, G. Heydt, Y. Chakhchoukh, N. Logic, and S. Sturgill, "The time skew problem in PMU measurements," in 2012 IEEE Power and Energy Society General Meeting, 2012, pp. 1-6

  3. [10]

    Data quality issues for synchrophasor applications Part I: a review,

    C. Huang et al., "Data quality issues for synchrophasor applications Part I: a review," Journal of Modern Power Systems and Clean Energy, vol. 4, no. 3, pp. 342-352, 2016

  4. [11]

    The Impact of Time Synchronization Deviation on the Performance of Synchrophasor Measurements and Wide Area Damping Control,

    T. Bi, J. Guo, K. Xu, L. Zhang, and Q. Yang, "The Impact of Time Synchronization Deviation on the Performance of Synchrophasor Measurements and Wide Area Damping Control," IEEE Transactions on Smart Grid, vol. 8, no. 4, pp. 1545-1552, 2017

  5. [12]

    Synchrophasor time skew: Formulation, detection and correction,

    Q. F. Zhang and V. M. Venkatasubramanian, "Synchrophasor time skew: Formulation, detection and correction," in 2014 North American Power Symposium (NAPS), 2014, pp. 1-6

  6. [14]

    Comparing the Topological and Electrical Structure of the North American Electric Power Infrastructure,

    E. Cotilla -Sanchez, P. D. H. Hines, C. Barrows, and S . Blumsack, "Comparing the Topological and Electrical Structure of the North American Electric Power Infrastructure," IEEE Systems Journal, vol. 6, no. 4, pp. 616-626, 2012

  7. [15]

    The power grid as a complex network: a survey,

    G. A. Pagani and M. Aiello, "The power grid as a complex network: a survey," Physica A: Statistical Mechanics and its Applications, vol. 392, no. 11, pp. 2688-2700, 2013

  8. [16]

    Creation of synthetic electric grid models for transient stability studies,

    T. Xu, A. B. Birchfield, K. S. Shetye, and T. J. Overbye, "Creation of synthetic electric grid models for transient stability studies," in The 10th Bulk Powe r Systems Dynamics and Control Symposium (IREP 2017) , 2017

  9. [17]

    T. ECEN. Electric Grid Test Case Repository [Online]. Available: https://electricgrids.engr.tamu.edu/electric-grid-test-cases/

  10. [18]

    Grid Structural Characteristics as Validation Criteria for Synthetic Networks,

    A. B. Birchfield, T. Xu, K. M. Gegner, K. S. Shetye, and T. J. Overbye, "Grid Structural Characteristics as Validation Criteria for Synthetic Networks," IEEE Transactions on Power Systems, vol. 32, no. 4, pp. 3258-3265, 2017

  11. [19]

    A synthetic data generator for clustering and outlier analysis,

    Y. Pei and O. Zaïane, "A synthetic data generator for clustering and outlier analysis," Department of Computing science, University of Alberta, edmonton, AB, Canada, 2006

  12. [20]

    Synthetic Generation of High-Dimensional Datasets,

    G. Albuquerque, T. Lowe, and M. Magno r, "Synthetic Generation of High-Dimensional Datasets," IEEE Transactions on Visualization and Computer Graphics, vol. 17, no. 12, pp. 2317-2324, 2011

  13. [21]

    Learning from imbalanced data sets with boosting and data generation: the databoost-im approach,

    H. Guo and H. L. Viktor, "Learning from imbalanced data sets with boosting and data generation: the databoost-im approach," ACM Sigkdd Explorations Newsletter, vol. 6, no. 1, pp. 30-39, 2004

  14. [22]

    Synthetic Dynamic PMU Data Generation: A Generative Adversarial Network Approach,

    X. Zheng, B. Wang, and L. Xie, "Synthetic Dynamic PMU Data Generation: A Generative Adversarial Network Approach," arXiv preprint arXiv:1812.03203, 2018

  15. [23]

    Online Detection of Low -Quality Synchrophasor Measurements: A Data-Driven Approach,

    M. Wu and L. Xie, "Online Detection of Low -Quality Synchrophasor Measurements: A Data-Driven Approach," IEEE Transactions on Power Systems, vol. 32, no. 4, pp. 2817-2827, 2017

  16. [24]

    Generation of synthetic data sets for evaluating the accuracy of knowledge discovery systems,

    D. R. Jeske et al., "Generation of synthetic data sets for evaluating the accuracy of knowledge discovery systems," in Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining, 2005, pp. 756-762: ACM

  17. [25]

    (2018, 08 -Nov-2018)

    PJM. (2018, 08 -Nov-2018). Private generic PMU data . Available: https://www.pjm.com/markets-and-operations/advanced-tech- pilots/private-generic-pmu-data.aspx

  18. [26]

    Identifying U seful Statistical Indicators of Proximity to Instability in Stochastic Power Systems,

    G. Ghanavati, P. D. H. Hines, and T. I. Lakoba, "Identifying U seful Statistical Indicators of Proximity to Instability in Stochastic Power Systems," IEEE Transactions on Power Systems, vol. 31, no. 2, pp. 1360- 1368, 2016

  19. [27]

    Assessing Gaussian Assumption of PMU Measurement Error Using Field Data,

    S. Wang, J. Zhao, Z. Huang, and R. Diao, "Assessing Gaussian Assumption of PMU Measurement Error Using Field Data," IEEE Transactions on Power Delivery, vol. 33, no. 6, pp. 3233-3236, 2018

  20. [28]

    Preprocessing synchronized phasor measurement data for spectral analysis of electromechanical oscillations in the Nordic Grid,

    L. Vanfretti, S. Bengtsson, and J. O. Gjerde, "Preprocessing synchronized phasor measurement data for spectral analysis of electromechanical oscillations in the Nordic Grid," International Transactions on Electrical Energy Systems, vol. 25, no. 2, pp. 348-358, 2015

  21. [29]

    G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time series analysis: forecasting and control. John Wiley & Sons, 2015

  22. [30]

    P. J. Brockwell, R. A. Davis, and M. V. Calder, Introduction to time series and forecasting. Springer, 2002

  23. [31]

    Trudnowski

    D. Trudnowski. Properties of the Dominant Inter -Area Modes in the WECC Interconnect [Online]. Available: https://www.wecc.biz/Reliability/WECCmodesPaper130113Trudnowski .pdf

  24. [32]

    1996 System Disturbances [Online]

    NERC. 1996 System Disturbances [Online]. Available: https://www.nerc.com/pa/rrm/ea/System%20Disturbance%20Reports%2 0DL/1996SystemDisturbance.pdf

  25. [33]

    Initial results in electromechanical mode identification from ambient data,

    J. W. Pierre, D. J. Trudnowski, and M. K. Donnelly, "Initial results in electromechanical mode identification from ambient data," IEEE Transactions on Power Systems, vol. 12, no. 3, pp. 1245-1251, 1997

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.