{"id":"be480968-b520-4194-ae6c-0d52110a5edb","arxiv_id":"1908.05244","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Real PMU measurements contain missing samples, outliers, and ambient oscillations that synthetic data should replicate, according to statistics from 123 public PMUs.","lead":"This paper analyzes real power-grid phasor measurement unit (PMU) data from a public dataset to quantify features like missing samples, outliers, and low-frequency oscillations. It argues that synthetic PMU data should include these features so that research results better match real-world outcomes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that adding anomalies, missing samples, and ambient modes improves synthetic-PMU realism is asserted, not tested; Section VI concludes without any comparison of synthetic data with and without these features.","rationale":"The reader identified the representativeness of the PJM dataset as the weakest assumption. That is a real generalization concern, but the more load-bearing weakness is that the central claim is not empirically validated at all. The paper measures features in one dataset and then asserts that including them improves realism; realism is never operationalized, no synthetic data is produced, and no comparison is performed. Thus the conclusion in Section VI overreaches even under a perfectly representative dataset. The descriptive statistics are useful and the conditional verdict is appropriate. My concern reinforces that condition rather than moving it; the paper should either run the proposed comparison or soften the claim to a recommendation for feature consideration in synthetic-data pipelines.","tokens_in":8167,"tokens_out":4210,"duration_ms":43472,"concrete_test":"Generate synthetic PMU voltage magnitude, angle, and frequency series from the ACTIVSg2000 test case in three variants: (A) raw simulation output; (B) A with outlier and dropout injection at the Section IV rates; (C) B with 0.3 Hz and 0.5 Hz ambient modes superimposed. Compare each variant to held-out real PMU data from a different PJM time period or another public dataset using a pre-registered distributional metric (e.g., Wasserstein distance on first-differenced angle and frequency, and KS statistic on outlier/dropout counts) plus a downstream mode-estimation error. If C does not beat A on the chosen metric, the central claim that these inclusions improve realism is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, stated in the Abstract and Section VI, is that including data anomalies, ambient oscillation content, and random missing-data samples improves the realism of synthetic PMU measurements. The evidence offered is descriptive: one 30-minute PJM dataset shows roughly 1% outlier samples, 38% of PMUs with at least one missing value, and persistent modes near 0.3 and 0.5 Hz. These observations establish prevalence, not benefit. Realism is never defined or measured, no synthetic dataset is generated, and no comparison is made between synthetic data with and without the proposed features. Section VI repeats the recommendation as a conclusion rather than demonstrating it. In addition, the statistics come from one time window of one utility's public data, with no error bars or cross-validation; Section II's implicit assumption that PJM data represent industry PMU data generally is load-bearing if the specific rates and modes are meant to transfer. Even granting representativeness, the central claim remains untested. The paper is a useful feature catalog, but the realism claim needs a validation experiment or softer wording.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes features of PMU measurements from the public PJM dataset (123 PMUs, one 30-minute window) to inform the generation of realistic synthetic PMU data. It proposes a variance decomposition that adds an anomaly term to the standard noise model, quantifies grid-dynamics variability and SNR on selected segments, reports outlier percentages in voltage magnitude and angle, computes dropout rates and maximum gap sizes for missing samples, and identifies low-frequency ambient oscillation modes via matrix-pencil analysis. The paper concludes that synthetic PMU data should include data anomalies, ambient oscillation content, and random missing-data samples in order to improve realism.","tokens_in":8374,"tokens_out":4237,"duration_ms":42881,"significance":"If the descriptive statistics are taken at face value, the paper provides a useful feature catalog for synthetic PMU data generation, with concrete prevalence numbers for outliers and missing data from a real industry dataset. The use of a public dataset and the clear presentation of per-PMU statistics (e.g., dropout rates and gap sizes in Fig. 10) are strengths. However, the central claim that injecting these features improves realism is not directly tested: no synthetic data is generated and no comparison is made between synthetic data with and without the proposed features. Several load-bearing methodological details, including the outlier detection rule, the selection of the representative segment, and the modal-analysis significance criterion, are unspecified. The paper is therefore best viewed as a motivating empirical study rather than a validated methodology.","major_comments":[{"comment":"The central conclusion, stated in the Abstract and in Section VI, is that including anomalies, ambient oscillations, and missing data samples 'helps to improve the realism' of synthetic PMU data. This claim is never tested in the manuscript: no synthetic dataset is generated, and no comparison is made between synthetic data with and without the proposed features against a defined measure of realism. Section VI repeats the recommendation as a conclusion rather than demonstrating it. To support the claim, the authors should either add a validation experiment (e.g., generate synthetic data with and without the features and compare distributional statistics or downstream-task performance) or explicitly reframe the contribution as a feature-prevalence catalog to guide future synthetic-data generation.","section":"Abstract and Section VI"},{"comment":"The segment selected for estimating grid-dynamics variability is described as 'observed to possess predominantly noiseless and error-free data samples.' This is a hand-picked, noiseless segment, so the resulting average variability of 10^-4 cannot be treated as a representative characterization of grid-dynamics variability without an objective selection criterion. The authors should either define how the segment was chosen or use the full distribution from all 123 PMUs (as in Fig. 4) and explain why the single selected segment is representative.","section":"Section III.A (Fig. 2)"},{"comment":"The outlier percentages in Fig. 8 are load-bearing for the recommendation to inject anomaly features, yet the outlier detection rule is not specified. The text states that samples are 'outside their local vicinities,' but the local window size, the outlier threshold, and the treatment of non-stationarity (including the angle-difference transformation (4)) are not defined. Without this algorithmic detail, the reported average of 1% and the maximum above 7% cannot be reproduced or meaningfully interpreted. Please specify the detection method and all parameters used.","section":"Section III.C (Fig. 8)"},{"comment":"The ambient-oscillation analysis uses a matrix-pencil method with a 10-second moving window and 5-second step, but the manuscript does not specify how modes are declared 'significant,' how many windows were analyzed, whether the frequency data come from one PMU or all 123 PMUs, or how the claimed persistent 0.3 Hz and 0.5 Hz modes are identified across windows. These details are necessary to support the strong claim that the two modes 'appear in all data windows.' Please provide the mode-selection criterion and, where possible, the distribution of detected modes across windows and PMUs.","section":"Section V"},{"comment":"All statistics in the paper are computed from a single 30-minute window of the PJM public PMU dataset. The manuscript implicitly assumes this dataset is representative of typical industry PMU data, and this assumption is load-bearing if the reported dropout rates, outlier percentages, and modal content are intended to set generation parameters for other systems. The authors should acknowledge this limitation explicitly and either include additional datasets or time periods or present the results as case-study statistics for this particular system.","section":"Sections II and IV"}],"minor_comments":[{"comment":"The captions of Fig. 11 and Fig. 12 both read 'First-order, stationary, 30-min voltage angles,' which is inconsistent with the content: Fig. 11 shows frequency measurements and Fig. 12 shows mode-analysis results. Please correct the captions.","section":"Figures 11 and 12"},{"comment":"The conclusion that the noise signal has 'strong i.i.d' features is based on a single autocorrelation function from one noise signal and an average ACF of 0.06 over lags. This is not a statistically rigorous whiteness test; a Ljung-Box test or confidence bounds under the null hypothesis would be more appropriate, and the sample size used for the ACF should be stated.","section":"Section III.B"},{"comment":"The dropout-rate and gap-size statistics would be more reproducible if the authors stated how missing samples were identified (e.g., detection of NaN or sentinel values such as 9999) and whether any preprocessing was applied to distinguish actual packet drops from bad data.","section":"Section IV"},{"comment":"Several figures are credited to the first author's dissertation [13]; for reproducibility, the data-processing steps used to generate these figures should be described in the paper rather than deferred to an unpublished dissertation. Reference [25] should also include an access date for the online resource.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a concise, data-descriptive study whose main gap is the unvalidated realism-improvement claim. The descriptive statistics are valuable, but either a validation experiment or a softened conclusion is needed to make the central claim defensible. The manuscript fits the scope of the venue, but the current framing overstates what the evidence supports."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a short empirical paper: the authors take 30 minutes of public PJM PMU data from 123 units and quantify the features synthetic data usually lacks — noise levels, outlier percentages, missing-sample rates and gap sizes, and ambient oscillation modes. The measured statistics are new for this dataset, and the dropout/gap distribution (38% of PMUs with at least one missing sample, ten with gaps over 3 seconds) plus the consistent ~0.3 and ~0.5 Hz modes are genuinely useful reference numbers for anyone building synthetic PMU generators. The paper is honest about being a feature catalog; it does not claim to have built a generator.\n\nThe soft spots are real but not fatal. The central sentence in the abstract and conclusion — that including these features \"helps to improve the realism\" of synthetic data — is never tested. Realism is not defined, no synthetic dataset is generated, and no comparison with and without features is made. That should be softened to a hypothesis or backed with a small validation experiment. Also: the 5-second segment in Section III.A is hand-picked as noiseless, the outlier detection rule is not specified, and the i.i.d. conclusion rests on an average ACF of 0.06, which is weak evidence. The matrix-pencil window choice is given but without sensitivity analysis. All statistics come from one 30-minute window from one utility, so there are no error bars and the representativeness assumption is load-bearing if the specific rates are meant to transfer. There are typos and mislabeled captions (Figs. 11 and 12 both say \"First-order, stationary, 30-min voltage angles\"), which suggests light proofreading.\n\nThe citation pattern is fine: some figures come from the first author's dissertation [13], but the core statistics appear to be computed directly from the public dataset, so I don't see a problem there.\n\nOverall, this is a useful data note, not a breakthrough. If the authors soften the realism claim or test it, it becomes a solid conference paper. I'd send it to peer review rather than desk reject — the descriptive statistics deserve an archival record, and the claims are checkable.\n\nRecommendation: engage, but ask for either a validation experiment or rewording of the central claim.","headline":"Useful empirical catalog of PMU data quirks from the PJM dataset, but the realism claim is asserted, not tested.","tokens_in":8883,"tokens_out":1967,"would_cite":true,"duration_ms":19272,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Synthetic power-grid data should include real PMU flaws: outliers, dropouts, and ambient oscillations.","keywords":["phasor measurement unit","synchrophasor","synthetic data generation","PMU data quality","data anomalies","ambient oscillations","missing data","power system measurements"],"falsifier":"Collect an independent set of PMU records from a different utility or time period and compute the same statistics; if most PMUs show no missing samples, outlier fractions well below 1%, and no spectral peaks near 0.3 and 0.5 Hz, the paper's feature targets are not universal. Alternatively, generate synthetic data with and without the recommended features and compare each to a held-out field record on the same variability, dropout, and modal statistics: the recommendation stands only if the feature-injected version is measurably closer to the field record.","tokens_in":7973,"feed_emoji":"⚡","tokens_out":5566,"duration_ms":53648,"temperature":0.7,"pith_summary":"This paper argues that synthetic phasor measurement unit (PMU) data generated from power-system simulations are unrealistically clean, and that researchers who test grid analytics on such data will draw conclusions that do not transfer to field conditions. Using 30-minute records from 123 field PMUs in a public dataset, it measures how often real data contain outliers, missing samples from dropped packets, noise, and persistent low-frequency ambient oscillations. It concludes that synthetic data generation should deliberately include these features, roughly 1% outliers, random dropout events, and ambient modes near 0.3 and 0.5 Hz, to match the statistical character of industry measurements. The paper is best read as an evidence-based recipe for making simulation data behave like real synchrophasor data, rather than a proposal of any specific generation algorithm.","feed_headline":"38% of real PMUs drop data—synthetic grids should too","feed_subtitle":"Field-PMU analysis finds outliers, packet-drop gaps, and ambient 0.3/0.5 Hz modes that clean simulations lack.","key_machinery":"The load-bearing structure is a variance decomposition, $\\sigma_{\\Delta V_M}^2 = \\sigma_{\\Delta V}^2 + \\sigma_{\\eta}^2 + \\sigma_e^2$, separating total measured voltage variability into grid dynamics, measurement noise, and data-anomaly components. Alongside this, the paper uses a dropout rate $\\rho$ and a maximum gap size $\\chi$ to quantify missing data, an autocorrelation function to confirm the nearly independent noise structure, and a matrix pencil modal analysis on moving windows to identify ambient low-frequency modes. These quantities convert the qualitative idea of realism into measurable targets that a synthetic data generator can aim to hit.","core_discovery":"Real PMU measurements from the public 123-PMU dataset are not clean signals: they contain measurement noise with signal-to-noise ratios around 41-47 dB, roughly 1% outlier samples in voltage magnitude and angle, frequent dropout events (38% of PMUs had at least one missing value, and ten PMUs had gaps longer than 3 seconds), and persistent electromechanical ambient modes near 0.3 and 0.5 Hz that appear in every analysis window. Simulation outputs typically lack all of these features. The paper therefore establishes that the variability in true PMU data can be decomposed into three additive components, grid dynamics, measurement noise, and data anomalies, and that a realistic synthetic generator should reproduce each component along with missing-data statistics and ambient modal content.","pith_inferences":["A natural but untested extension is to model dropout events as spatially correlated across nearby PMUs, since packet losses often share communication paths; the paper reports per-PMU rates only and does not address joint dropout patterns.","One could benchmark the realism claim directly by training analytics on feature-injected synthetic data and measuring their performance on held-out field PMU records versus analytics trained on clean simulation data; the paper does not run that end-to-end comparison.","The 0.3 and 0.5 Hz ambient modes are likely signatures of one interconnection, not universal constants; synthetic generators targeting other grids should derive their own modal content from local field data rather than reusing these frequencies.","The variance-decomposition targets could be turned into a generative test: draw synthetic noise to match the observed autocorrelation behavior (near zero after lag 1), add outliers and dropouts at the measured rates, and statistically compare the resulting distributions with field data."],"forward_implications":["Synthetic PMU benchmarks for oscillation detection, state estimation, or event classification should be stress-tested with injections of roughly 1% outliers and random missing samples, or their reported accuracy may not hold on field data.","A simulator that reproduces only base power-flow dynamics will understate measurement variability by omitting the noise and anomaly variance terms; adding them brings synthetic variability closer to observed values such as the roughly $10^{-4}$ average voltage variance reported.","The persistent 0.3 and 0.5 Hz modes characterize the interconnection studied; synthetic grids meant to emulate a real system can use such persistent modal content as an acceptance criterion for ambient dynamics.","Missing-data handling becomes part of the experimental pipeline: because 38% of PMUs had at least one gap, synthetic datasets without dropout cannot exercise the imputation and bad-data rejection logic that field engineers must run.","Researchers can use the paper's variance decomposition as a checklist for synthetic data validation, comparing each component separately rather than only the overall signal shape."],"supporting_citations":[{"why":"Supplies the public 123-PMU field dataset from which all outlier, dropout, and ambient-mode statistics are computed.","marker":"[25]"},{"why":"Establishes the typical 41-47 dB SNR range for power-system measurements, used to classify signal quality and infer the noise variability component.","marker":"[8]"},{"why":"Provides the base variance relation $\\sigma_{\\Delta V_M}^2 = \\sigma_{\\Delta V}^2 + \\sigma_\\eta^2$ that the paper extends with an anomaly variance term.","marker":"[26]"},{"why":"Defines the dropout-rate and gap-size measures used to quantify missing PMU data samples.","marker":"[7]"},{"why":"Supports the interpretation of persistent low-frequency modes as ambient electromechanical inter-area modes in an interconnected system.","marker":"[31]"},{"why":"Prior related work referenced for real-versus-simulated voltage comparisons and for voltage-angle outlier observations.","marker":"[13]"}],"fun_headline_variants":["Synthetic PMU data should mimic real noise, gaps, and modes","To fake grid data realistically, include PMU noise, dropouts, and modes","Real PMU flaws: noise, 0.3 Hz modes, and 38% dropouts—simulate all","Include 0.3 Hz modes and noise in synthetic PMU data for realism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The recommendations rest on one public 30-minute dataset from 123 PMUs on one grid; if that dataset is not typical of field PMU data, the specific outlier rates, dropout rates, and modal frequencies it prescribes would not generalize to other systems.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic PMU data should mimic real noise, gaps, and modes","To fake grid data realistically, include PMU noise, dropouts, and modes","Real PMU flaws: noise, 0.3 Hz modes, and 38% dropouts—simulate all","Include 0.3 Hz modes and noise in synthetic PMU data for realism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001423,"raw_usage":{"total_tokens":5693,"prompt_tokens":845,"completion_tokens":4848,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":4755}},"tokens_in":461,"tokens_out":4848,"duration_ms":33286,"temperature":1.0,"reasoning_tokens":4755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:19:07.488439+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect an independent set of PMU records from a different utility or time period and compute the same statistics; if most PMUs show no missing samples, outlier fractions well below 1%, and no spectral peaks near 0.3 and 0.5 Hz, the paper's feature targets are not universal. Alternatively, generate synthetic data with and without the recommended features and compare each to a held-out field record on the same variability, dropout, and modal statistics: the recommendation stands only if the feature-injected version is measurably closer to the field record.","supporting_citations":[{"cited_title":"(2018, 08 -Nov-2018)","cited_arxiv_id":null,"evidence_quote":"Supplies the public 123-PMU field dataset from which all outlier, dropout, and ambient-mode statistics are computed."},{"cited_title":"Characterizing and quantifying noise in PMU data,","cited_arxiv_id":null,"evidence_quote":"Establishes the typical 41-47 dB SNR range for power-system measurements, used to classify signal quality and infer the noise variability component."},{"cited_title":"Identifying U seful Statistical Indicators of Proximity to Instability in Stochastic Power Systems,","cited_arxiv_id":null,"evidence_quote":"Provides the base variance relation $\\sigma_{\\Delta V_M}^2 = \\sigma_{\\Delta V}^2 + \\sigma_\\eta^2$ that the paper extends with an anomaly variance term."},{"cited_title":"PMU Data Quality: A Framework for the Attributes of PMU Data Quality and Quality Impacts to Synchrophasor Applications,","cited_arxiv_id":null,"evidence_quote":"Defines the dropout-rate and gap-size measures used to quantify missing PMU data samples."},{"cited_title":"Trudnowski","cited_arxiv_id":null,"evidence_quote":"Supports the interpretation of persistent low-frequency modes as ambient electromechanical inter-area modes in an interconnected system."},{"cited_title":"Data analytics and wide -area visualization associated with power systems using phasor measurements,","cited_arxiv_id":null,"evidence_quote":"Prior related work referenced for real-versus-simulated voltage comparisons and for voltage-angle outlier observations."}],"review_version":1}