{"id":"7ddc2fcc-9955-462c-b09b-b49266a8a629","arxiv_id":"1908.03823","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Across 13 measured grids, single-peak frequency distributions are approximately normal, and island grids generally show larger frequency swings than mainland grids.","lead":"This paper measures the power frequency of 13 grids worldwide using FNET/GridEye sensors and compares how steady each grid's frequency is. It finds that grids with a single-peak frequency distribution behave roughly like a normal (bell-shaped) curve, while others show multiple peaks.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'almost normal' claim rests on an unfalsifiable visual comparison: the Gaussian in Fig. 8 uses the same data's mean and standard deviation, and no goodness-of-fit test, tail metric, or unimodality criterion is reported.","rationale":"Good faith reading: the paper is a descriptive observational survey, and its raw measurements, if accessible, would be valuable; the observation that island grids tend to have larger frequency deviations is plausible. However, the strongest interpretive claim—that single-peak distributions are approximately Gaussian—is the one the paper highlights as 'a meaningful insight,' and it is the least supported. The reader's weakest assumption (one FDR per grid) is real, but it concerns external validity; the normality claim can be tested on the reported data themselves, and the paper provides no quantitative basis for it. A second, independent issue is the Egypt-NERC comparison in Section III, which compares a standard deviation to a frequency threshold; that is dimensionally confused and should be corrected, but it does not affect the single/multi-peak classification. The absence of public data (historical data require privilege authorization) prevents independent replication, so the concrete test above would need to be run by the authors or by a reviewer with authorized access. This supports the reader's CONDITIONAL verdict rather than ACCEPT or REJECT.","tokens_in":6886,"tokens_out":4983,"duration_ms":55552,"concrete_test":"For each of the seven grids labeled single-peak (EI, WECC, Hawaii, Germany, Japan, Australia, Egypt), recompute the empirical distribution from the same FDR data and: (1) apply Hartigan's dip test for unimodality and repeat the single/multi-peak classification over bin widths of 0.001, 0.005, and 0.01 Hz plus a kernel density estimate; (2) fit the Gaussian, Laplace, and Student-t distributions by maximum likelihood and compare AIC/BIC and the largest absolute difference between empirical and fitted CDFs (Kolmogorov-Smirnov distance). If the dip test confirms unimodality but a non-Gaussian distribution has lower AIC or the KS distance exceeds a pre-specified tolerance (e.g., 0.01) for the Gaussian, the 'follow normal' claim should be downgraded to 'unimodal, approximately symmetric' and the fit quality quantified. If the Gaussian is competitive on all three grids, the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central insight—that single-peak frequency distributions almost follow the normal distribution—is asserted from Section III and Fig. 8, where a Gaussian with the sample mean and standard deviation is overlaid on the same histogram. This is not a falsifiable fit: with those parameters the normal curve is anchored to the data's center, and 'almost' is never defined. No Anderson-Darling or Kolmogorov-Smirnov statistic, quantile-quantile plot, tail comparison, or alternative distribution (e.g., Laplace, Student-t) is reported. Because each FDR records 10 samples/s for three months, the samples are huge; classical normality tests would almost surely reject exact Gaussianity, so the meaningful question is whether the Gaussian captures the distribution within a practically relevant tolerance—exactly what the paper does not quantify. The single-peak/multi-peak classification in Fig. 7 is likewise made by eye, with no bin-width sensitivity analysis and no unimodality test, so the set of seven grids whose distributions are claimed Gaussian is not robust. The one-FDR-per-grid representativeness issue (Section II) is a further limitation, but the primary weakness is internal to the analysis: the key claim cannot be evaluated or falsified from the reported evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an observational statistical study of power frequency in 13 worldwide power grids using FNET/GridEye distribution-level frequency disturbance recorder measurements sampled at 10 Hz over roughly three months. For each grid it computes mean frequency and standard deviation, compares mainland versus island behavior, and classifies the empirical frequency distributions into single-peak and multi-peak families. It then selects EI, Egypt, and Japan for more detailed daily statistics, and claims that the single-peak distributions 'almost follow the normal distribution,' illustrating this by overlaying Gaussian curves on the histograms of seven grids. The paper concludes that the observed statistical characteristics can aid grid operators and future frequency regulation studies.","tokens_in":7072,"tokens_out":2105,"duration_ms":24257,"significance":"The dataset assembled here is valuable: simultaneous multi-grid frequency observations across 13 systems are rare, and the reported means and standard deviations provide a useful descriptive baseline for cross-grid frequency behavior. If the single-peak/multi-peak classification and the Gaussian approximation for single-peak grids were rigorously established, the paper would provide a simple empirical model of interest to frequency-control and reserve-sizing studies. However, the central normality claim is currently supported only by an in-sample visual overlay, and the classification is made by eye; as presented, the scientific contribution is descriptive rather than confirmatory. The paper would be strengthened materially by adding quantitative distributional tests, unimodality criteria, and sensitivity analysis.","major_comments":[{"comment":"The headline claim that single-peak frequency distributions 'almost follow the normal distribution' is not supported by the evidence reported. The red normal curves in Fig. 8 are constructed from the sample mean and standard deviation of the very same histogram data, so the agreement is an in-sample property, not an independent test. No goodness-of-fit statistic (e.g., Kolmogorov-Smirnov, Anderson-Darling), tail-quantile comparison, or alternative distribution benchmark (e.g., Laplace, Student-t) is provided, and the qualifier 'almost' is never defined. Since each grid contributes millions of samples (10 samples/s for three months), exact Gaussianity would be rejected by any classical test; the practically relevant question is whether the Gaussian approximation holds within some stated tolerance, and the manuscript does not quantify that tolerance.","section":"Section III, Figs. 7 and 8"},{"comment":"The single-peak versus multi-peak classification of the 13 grids is based on visual inspection of one histogram per grid, with no unimodality test, no bin-width sensitivity analysis, and no quantitative criterion for what constitutes a peak. For borderline cases such as Egypt (Fig. 7(g)), whose histogram appears skewed and heavy-tailed, the inclusion in the 'single-peak normal' family is not self-evident; for ERCOT and the multi-peak group, the visible secondary modes could be artifacts of binning or of a sensor-specific event. Without a reproducible classification rule, the claim that the 13 grids divide into these two families is not robust.","section":"Section III, Fig. 7 and Table I"},{"comment":"The cross-grid comparison assumes that one FDR per country represents the frequency behavior of the entire interconnection, justified only by Refs. [29]-[30]. For large systems such as EI and WECC, a single distribution-level sensor can be influenced by local load or generation events, and the manuscript provides no validation against other sensors on the same grid or against control-area frequency statistics. Consequently, the reported mean, standard deviation, and the single/multi-peak classification could reflect local sensor behavior rather than grid-wide frequency status, which limits the generality of the cross-grid conclusions.","section":"Section II, last paragraph"}],"minor_comments":[{"comment":"The manuscript states that three months of data are retrieved for the 13 grids, but the daily analysis in Fig. 6 is described as using a one-month period; the discrepancy should be clarified.","section":"Section III, Fig. 6"},{"comment":"The phrase 'due to its clean, low-cost and inexhaustible features' refers to renewable energy sources but uses a singular pronoun; please revise for grammatical consistency.","section":"Introduction, paragraph 1"},{"comment":"Reference [2] appears to be a URL fragment appended to the National Conference of State Legislatures reference without proper formatting, and several references (e.g., [13]) are not cited in the text; the reference list should be checked for completeness and consistency.","section":"References"},{"comment":"The sentence 'The frequencies in all mainland power grids have smaller standard deviations, comparing to the frequencies in island power grids' is contradicted by the paper's own observation that Egypt, a mainland grid, has the highest standard deviation; rephrase to acknowledge the exceptions.","section":"Section II, paragraph 2"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a useful observational dataset, but the central 'almost normal' claim is under-supported as written. The revision path is clear: add quantitative distributional tests and a defined tolerance, make the single/multi-peak classification reproducible, and address the single-sensor representativeness issue. I do not see a need for rejection, but the current evidence does not justify acceptance at a journal level."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is the 13-grid frequency survey: same measurement network, same time window, same preprocessing, and a clean table of means and standard deviations. The single-peak/multi-peak taxonomy is a genuinely new observation for FNET/GridEye data, and it gives operators a quick way to think about frequency behavior on small islands versus large interconnections.\n\nThe headline claim—that single-peak distributions are 'almost normal'—is not established by the evidence in the paper. The red normal curves in Fig. 8 are generated from the sample mean and standard deviation of the same histograms, so the overlay is an in-sample description, not a fit test. No Anderson-Darling, KS, QQ plot, or tail metric is reported; 'almost' is never defined. With 10 samples per second for three months, exact normality is almost certainly false; what matters is whether the Gaussian is close enough for reserve sizing or control studies, and the paper doesn't quantify that tolerance. Also, the single-peak versus multi-peak classification is made by eye; there is no bin-width sensitivity or unimodality test. So the central interpretive claim is falsifiable in principle but isn't actually tested here.\n\nTwo other soft spots. First, one FDR per grid is thin for EI or WECC; the paper cites earlier work saying intra-system frequency differences are small, but a single distribution-level sensor can still be biased by local events. That caveat weakens the cross-grid comparison, though it's a minor issue for a survey. Second, the Egypt comparison with NERC load-shedding thresholds is dimensionally confused—standard deviation is not the same quantity as a frequency setpoint—and should be fixed or removed.\n\nThe data are gated behind privilege authorization, which limits reproducibility, but the descriptive statistics themselves are straightforward summaries of the measurements. I'm fairly confident the raw numbers are what they say they are.\n\nWho should read this: anyone needing a quick empirical reference for frequency dispersion across a range of grid sizes, and researchers working on frequency-distribution models. It deserves a serious referee, because the dataset is rare and the survey structure is useful, but the normality claim needs proper testing and the one-sensor-per-grid limitation needs a sensitivity check before the paper is accepted.\n\nRecommendation: send to peer review with a request for goodness-of-fit analysis, bin-width sensitivity, and correction of the Egypt/NERC passage.","headline":"A worthwhile 13-grid frequency survey whose 'almost normal' claim is a same-data overlay, not a tested result.","tokens_in":7650,"tokens_out":2488,"would_cite":false,"duration_ms":26347,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A bell curve describes frequency in seven of 13 power grids.","keywords":["worldwide power grids","frequency observations","statistical analysis","PMU application","big data","FNET/GridEye","normal distribution","frequency disturbance recorder"],"falsifier":"Take two or more FDRs located in different regions of the same interconnection (for example, the U.S. Eastern Interconnection) and compute their three-month frequency histograms separately; if the single- versus multi-peak classification or the fitted Gaussian mean and standard deviation differ markedly between the sensors, then the paper's per-grid statistics reflect local sensor behavior rather than grid-wide frequency status.","tokens_in":6617,"feed_emoji":"⚡","tokens_out":8581,"duration_ms":76206,"temperature":0.7,"pith_summary":"Using three months of distribution-level frequency data from the FNET/GridEye wide-area monitoring system, this paper describes the statistical behavior of 13 power grids across five continents. It divides them into two families: seven grids with single-peak frequency histograms and six with multi-peak histograms. For the single-peak family, the paper claims the measured frequency distribution is well approximated by a normal distribution built from the same data's mean and standard deviation. If this holds, frequency behavior in those grids can be summarized by two numbers, and the study offers a simple statistical typology for comparing grid operations worldwide.","feed_headline":"A bell curve describes frequency in seven of 13 power grids","feed_subtitle":"One sensor per grid, three months of data: grids split into single-peak and multi-peak families.","key_machinery":"The key machinery is the frequency disturbance recorder (FDR), a low-cost distribution-level sensor that streams GPS-synchronized frequency measurements at ten samples per second to FNET/GridEye servers. Three months of data from one FDR per grid are cleaned, converted to per-unit frequency, and summarized as empirical probability density functions. The paper then fits a two-parameter normal distribution to the single-peak histograms using the sample mean and standard deviation, using the visual match between the fitted curve and the histogram as evidence. Daily means and standard deviations for EI, Egypt, and Japan are additionally used to reveal operating status such as Egypt's persistent under-frequency operation.","core_discovery":"The central discovery is that for seven of the thirteen grids – the U.S. Eastern Interconnection, WECC, Hawaii, Germany, Japan, Australia, and Egypt – the empirical probability density function of frequency closely matches a Gaussian curve whose parameters are the empirically computed mean and standard deviation. The remaining six grids – ERCOT, Saudi Arabia, Northern Ireland, Ireland, England, and Bahamas – show multi-peak histograms that no single normal distribution can represent. The paper further shows that mainland grids generally have smaller frequency standard deviations than island grids, with Egypt as the largest-deviation outlier and Hawaii as a low-deviation island, and that most grids operate within NERC load-shedding thresholds.","pith_inferences":["If the Gaussian behavior is persistent, mean and standard deviation could serve as simple health metrics: a shift in the mean or widening of the standard deviation after a policy change would be detectable with these two numbers alone.","The six multi-peak grids invite a mixture-of-Gaussians or heavy-tailed follow-up analysis, which the paper does not pursue.","The single-sensor assumption is directly testable: comparing two FDRs on the same interconnection would show whether the reported statistics are grid properties or sensor properties.","Because FNET/GridEye already deploys hundreds of FDRs in dozens of countries, the same histogram procedure could produce a global map of frequency statistics far beyond the 13 grids studied here."],"forward_implications":["For single-peak grids, a normal distribution with the observed mean and standard deviation provides a workable empirical model, so frequency can be summarized by two numbers without a more complex model.","The single-peak/multi-peak taxonomy gives a simple way to compare frequency statistics across countries and could serve as a baseline for tracking how renewable penetration changes frequency behavior.","Mainland grids in this sample generally exhibit smaller frequency standard deviations than island grids, suggesting that interconnection size dampens frequency fluctuation; Egypt is an exception with the highest deviation, and Hawaii is an island with unusually low deviation.","Most studied grids keep their frequency deviations within NERC's under-frequency load-shedding thresholds, while Egypt's standard deviation approaches that limit, indicating a stressed system over the observation period."],"supporting_citations":[{"why":"Describes the FNET/GridEye distribution-level wide-area monitoring system and the FDR hardware that supplies all frequency measurements analyzed in the paper.","marker":"[23]"},{"why":"Supports the assumption that frequency differences within a single interconnection are small, the basis for selecting one FDR per grid.","marker":"[29]"},{"why":"Documents the FNET implementation and reinforces the same single-sensor representativeness assumption.","marker":"[30]"},{"why":"Sets the NERC under-frequency load-shedding thresholds against which the paper compares observed frequency deviations.","marker":"[31]"},{"why":"Earlier FNET/GridEye study of worldwide grid frequency that the paper extends and that validates the observation approach.","marker":"[15]"}],"fun_headline_variants":["Seven of 13 grids show bell-curve frequency","Frequency distributions split: 7 normal, 6 not","Power grid frequency: mainland vs island patterns","Gaussian fit seen in 7 of 13 power grids","Frequency stats differ by grid type, not just region"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire cross-grid comparison rests on the assumption that a single frequency disturbance recorder per country captures the frequency behavior of the whole interconnection, with no validation against other sensors on the same grid.","fun_headline_variants_meta":{"raw":{"variants":["Seven of 13 grids show bell-curve frequency","Frequency distributions split: 7 normal, 6 not","Power grid frequency: mainland vs island patterns","Gaussian fit seen in 7 of 13 power grids","Frequency stats differ by grid type, not just region"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1213,"prompt_tokens":892,"completion_tokens":321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":244}},"tokens_in":508,"tokens_out":321,"duration_ms":4409,"temperature":1.0,"reasoning_tokens":244,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:01:02.917098+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two or more FDRs located in different regions of the same interconnection (for example, the U.S. Eastern Interconnection) and compute their three-month frequency histograms separately; if the single- versus multi-peak classification or the fitted Gaussian mean and standard deviation differ markedly between the sensors, then the paper's per-grid statistics reflect local sensor behavior rather than grid-wide frequency status.","supporting_citations":[{"cited_title":"A Distribution Level Wide Area Monitoring System for the Electric Power Grid –FNET/GridEye,","cited_arxiv_id":null,"evidence_quote":"Describes the FNET/GridEye distribution-level wide-area monitoring system and the FDR hardware that supplies all frequency measurements analyzed in the paper."},{"cited_title":"Internet based frequency monitoring network (FNET),","cited_arxiv_id":null,"evidence_quote":"Supports the assumption that frequency differences within a single interconnection are small, the basis for selecting one FDR per grid."},{"cited_title":"Power system frequency moni toring network (FNET) implementation,","cited_arxiv_id":null,"evidence_quote":"Documents the FNET implementation and reinforces the same single-sensor representativeness assumption."},{"cited_title":"Automatic Under frequency Load Shedding Requirements","cited_arxiv_id":null,"evidence_quote":"Sets the NERC under-frequency load-shedding thresholds against which the paper compares observed frequency deviations."},{"cited_title":"Observation of inertial frequency response of main power grids worldwide using FNET/GridEye,","cited_arxiv_id":null,"evidence_quote":"Earlier FNET/GridEye study of worldwide grid frequency that the paper extends and that validates the observation approach."}],"review_version":1}