{"id":"2e0ae3ce-6e73-4afd-8990-db7f9414abea","arxiv_id":"2412.08065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A survey benchmarking five open-source power system simulators that support grid-forming inverters, with timing and capability comparisons for machine learning data generation.","lead":"This paper compares five open-source power grid simulators that can model grid-forming inverters, the technology that lets solar and battery plants help keep the grid stable. It serves as a reference guide for engineers who use simulator-generated data to train machine learning models for power systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Table II timing ranking is not backed by a documented measurement; Section IV says only ANDES and PSID.jl were simulated, yet Dynaomegao's 0.34s/374s appears without hardware, solver, or fidelity details.","rationale":"The reader's CONDITIONAL verdict is appropriate: the survey's qualitative content reliably summarizes five open-source tools and their GFM model support, and the ANDES-vs-PSID.jl case study provides a useful consistency check. The quantitative ranking, however, is the main basis for the paper's ML-focused guidance, and it is under-supported. My concern is more specific than the reader's weakest_assumption: it is not merely that hardware/solver details are missing, but that Section IV explicitly limits the performed simulations to ANDES and PSID.jl, while Table II also lists Dynaomegao. That discrepancy makes the Dynaomegao timing entry effectively unverifiable from the manuscript. A concrete reproducibility test would settle whether the ranking survives. If the Dynaomegao number cannot be reproduced or was generated at a different fidelity, the paper should present that entry as non-comparable, which would not overturn the survey's qualitative value but would require a revised conclusion about relative ML suitability. Since the reader already recommended CONDITIONAL and this concern reinforces that conditional status rather than moving it to acceptance or rejection, the verdict should remain unchanged.","tokens_in":10600,"tokens_out":2654,"duration_ms":29226,"concrete_test":"Request or independently reproduce the benchmark artifact: run the same modified IEEE 14-bus case (generator 5 as SG, GFM-Droop, and GFM-VSM; generation trip at 1s; duration 10s; step 0.01s) on one machine with the same compiler/runtime versions for ANDES, PSID.jl, and Dynaomegao. Use Dynaomegao's DynaSwing QSP mode, not DynaWave quasi-EMT, and record hardware, solver, parallel workers, and warm-up. If Dynaomegao's 0.34s/374s cannot be reproduced within a factor of two, or if the original number was not measured under these conditions, Table II should be revised to label that entry as reported-by-developer or removed from the direct comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central actionable claim is that Table II quantifies simulator suitability for ML data generation via simulation time, with PSID.jl fastest, ANDES competitive under parallelization, and Dynaomegao in between. That ranking is load-bearing for the paper's conclusion that ML practitioners can use these numbers for practical guidance. The weakest point is the provenance of the Dynaomegao timing. Section IV explicitly states: \"Since ANDES and PSID.jl provide detailed manuals and programming guidance, TDSs were performed on these two simulators\" using a modified IEEE 14-bus system. No Dynaomegao simulation is described, yet Table II reports Dynaomegao at 0.34s per run and 374s per 1000 samples. The paper gives no indication whether that number was measured by the authors, taken from documentation, or produced under the same conditions. The table caption declares \"Simulation time (1/1000 samples)\" with step size 0.01s and duration 10s, but omits hardware, solver algorithm, warm-up, parallelization, scenario-sampling procedure, and whether Dynaomegao used DynaSwing (QSP) or DynaWave (quasi-EMT). If Dynaomegao's entry came from a quasi-EMT model while ANDES and PSID.jl used positive-sequence QSP, the comparison mixes fidelities and the ranking is not apples-to-apples. The reader's weakest_assumption correctly flags under-specification; the sharper issue is that one of the three timed entries appears to have no described measurement at all. This does not invalidate the qualitative survey, but it does mean the quantitative guidance in Table II is not currently supported as reported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript surveys five open-source power system dynamic simulators (ANDES, PowerSimulationsDynamics.jl, Dynaωo, OpenDSS, and GridLAB-D) that support grid-forming (GFM) inverter models, with an emphasis on their suitability for machine-learning data generation. It categorizes time-domain simulation approaches, reviews GFM control structures, and presents feature comparisons in Tables I and II. The authors report time-domain simulation case studies on a modified IEEE 14-bus system for ANDES and PSID.jl with generator 5 configured as SG, GFM-droop, and GFM-VSM. The central quantitative contribution is Table II's simulation-time comparison, which ranks PSID.jl fastest (0.05 s/31.1 s per run/1,000 samples), Dynaωo intermediate (0.34 s/374 s), and ANDES slowest (1.27 s/224 s), and the paper concludes that these numbers provide practical guidance for ML practitioners.","tokens_in":10939,"tokens_out":8791,"duration_ms":78750,"significance":"The survey addresses a timely need: open-source simulators with GFM capability are evolving quickly, and ML researchers require guidance on which tools support flexible scenario generation and fast data production. The paper's strengths are organizational: it consolidates GFM model availability (e.g., REGCA1/REGCV1/REGF1-3 in ANDES, GFM-droop/VSM/dVOC/matching in PSID.jl and Dynaωo), interface features, and data formats from project documentation, and it provides a concrete two-simulator case study. If the timing comparison can be made reproducible and properly scoped, the survey would be a useful reference. At present, however, the headline numerical ranking is not supported by the evidence reported in the manuscript.","major_comments":[{"comment":"The Dynaωo timing entries (0.34 s/374 s) in Table II are not supported by any described measurement. Section IV states that 'TDSs were performed on these two simulators' (ANDES and PSID.jl) and no Dynaωo simulation is described anywhere in the paper. Please state explicitly whether the Dynaωo values were measured by the authors, taken from project documentation, or estimated, and under what conditions. Without this provenance, the ranking implied by Table II—PSID.jl fastest, Dynaωo intermediate, ANDES slowest—is not established.","section":"Section IV, Table II"},{"comment":"The wall-clock-time comparison is under-specified to the point of not being reproducible. The caption reports only step size (0.01 s) and duration (10 s); it omits hardware, solver algorithm and tolerances, parallelization method, warm-up, the procedure used to generate 1,000 samples, and, for Dynaωo, whether the timing used DynaSwing (QSP) or DynaWave (quasi-EMT). If Dynaωo's entry reflects a quasi-EMT simulation while ANDES and PSID.jl use positive-sequence QSP, the entries are not directly comparable. Please provide a full measurement protocol or restrict the quantitative claim to the two simulators actually benchmarked.","section":"Section IV, Table II"},{"comment":"The relationship between the single-run and 1,000-sample timings is internally inconsistent. The entries imply average per-sample times of 0.224 s for ANDES, 0.031 s for PSID.jl, and 0.374 s for Dynaωo, which are not consistent with the stated single-run times (1.27 s, 0.05 s, and 0.34 s, respectively). If the 1,000-sample runs are simply 1,000 independent simulations, the averages should be close to the single-run times; the large deviations indicate undocumented parallelization, warm-starting, or different sampling procedures. Please define how the 1,000-sample values were generated and confirm that the same procedure was used for all simulators.","section":"Table II"}],"minor_comments":[{"comment":"The text says the modified 14-bus system structure and results are shown in Figs. 3, 4, and 5, but Fig. 3 is the DL/RL framework diagram and no 14-bus system structure figure appears; please renumber the figures or insert the missing diagram.","section":"Section IV"},{"comment":"The header 'Simulation time (1/1000 samples)' is ambiguous; please state explicitly that the entries are single-run time / total wall-clock time for 1,000 samples.","section":"Table II"},{"comment":"The symbol '/' is used with two different meanings: 'does not mention this function' in Table I and 'not applicable' in Table II; please define the symbols separately for each table.","section":"Tables I and II"},{"comment":"The statement that 'due to Python's slower speed compared to Julia and C++, ANDES has the longest simulation time' is a causal attribution that is not established by the reported measurements; solver implementation, model complexity, and numerical settings are also relevant.","section":"Table II, Section III-A"},{"comment":"The claim that the ANDES and PSID.jl results are 'very similar' is qualitative; please add a quantitative comparison (e.g., maximum or RMS trajectory error) and specify the GFM control parameters used, so that the case study is reproducible.","section":"Section IV"},{"comment":"No simulator version numbers or documentation access dates are given; since capabilities and model lists change between releases, please state the versions reviewed.","section":"Tables I and II"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the provenance of the Dynaωo timing value. I would be comfortable with a major revision that either (a) reports a reproducible measurement for Dynaωo under identical conditions, or (b) removes Dynaωo from the timing comparison and clearly labels the remaining ranking as a two-simulator benchmark. The paper would also benefit from a data/code availability statement for the IEEE 14-bus case study."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a serviceable survey of five open-source simulators with GFM support, and the qualitative tables are genuinely useful. But the quantitative heart of the paper—Table II's timing comparison—is not currently supported, and the paper shouldn't be published as-is.\n\nWhat's new: it collects current GFM model support for ANDES, PSID.jl, Dynaωo, OpenDSS, and GridLAB-D, and it positions them for ML data generation, which the older surveys don't do. The case study comparing ANDES and PSID.jl on a modified 14-bus system is a concrete sanity check and shows the two simulators produce similar trajectories. That's real work.\n\nWhat's soft: the timing table is the main problem. Section IV only describes running ANDES and PSID.jl, yet Table II reports Dynaωo at 0.34s per run and 374s per 1,000 samples, with no hint of whether that was measured, taken from documentation, or what fidelity was used. Dynaωo has both a QSP module (DynaSwing) and a quasi-EMT module (DynaWave), and mixing those with ANDES/PSID's QSP would make the comparison meaningless. Even for ANDES and PSID.jl, no hardware, solver settings, parallelization, or variance are reported, so the 0.05s vs 1.27s headline is a fragile result. 'Unknown' for Dynaωo's data format sits oddly next to a precise timing entry. The tool-selection criteria are also unstated, which is a minor but real omission for a survey. No code or data was shipped, which would have made the timing verifiable.\n\nThat said, the qualitative content is mostly fine. The feature tables look consistent with each tool's documentation, the GFM model taxonomy is accurately sketched, and the conclusion that PSID.jl is fast and ANDES is parallelizable is plausible even if the numbers aren't rigorous. [21] is a peripheral self-citation and not a problem.\n\nWho's this for: researchers starting to generate ML training data from open-source dynamic simulators. They'll get a good map of options, and they should treat Table II with suspicion until the authors document the measurements.\n\nRecommendation: send it to review. The survey fills a small but real gap, and the timing issue is fixable—either report full measurement conditions for all three tools or drop Dynaωo from the timing comparison. A serious referee should catch this before it becomes a widely cited number.","headline":"A useful but uneven survey of open-source GFM-capable dynamic simulators; the qualitative comparison is worth publishing after the timing claims are either substantiated or removed.","tokens_in":11464,"tokens_out":2132,"would_cite":false,"duration_ms":21018,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"According to the paper's IEEE 14-bus benchmark, PSID.jl generates 1,000 ML training samples in 31.1 seconds, 7–12× faster than the other timed simulators.","keywords":["open-source power system dynamic simulator","grid-forming inverter","time-domain simulation","machine-learning data generation","quasi-static phasor simulation","ANDES","PowerSimulationsDynamics.jl","IEEE 14-bus system"],"falsifier":"Re-run the 1,000-sample benchmark on the same IEEE 14-bus case on identical hardware with each simulator's documented solver and parallelization settings; if PowerSimulationsDynamics.jl no longer produces the shortest wall-clock time by the reported margin, the paper's ranking fails. A weaker check: repeat the comparison on a larger bus system or with a randomized fault scan and see whether the ordering changes.","tokens_in":10408,"feed_emoji":"⚡","tokens_out":11407,"duration_ms":93565,"temperature":0.7,"pith_summary":"This paper sets out to tell machine-learning researchers which open-source power system dynamic simulator to use when generating training data for tasks like transient stability prediction and reinforcement learning, now that grid-forming (GFM) inverters—units that establish voltage and frequency rather than tracking the grid—need to be modeled. It surveys five tools—ANDES, PowerSimulationsDynamics.jl, Dynaωo, OpenDSS, and GridLAB-D—and compares them on modeling capability, GFM control support, programmability, and measured simulation speed. Its headline empirical result is a benchmark on the IEEE 14-bus system: PowerSimulationsDynamics.jl completes one 10-second dynamic simulation in 0.05 seconds and 1,000 Monte Carlo samples in 31.1 seconds, while ANDES takes 1.27 seconds and 224 seconds and Dynaωo takes 0.34 seconds and 374 seconds. The paper frames these numbers as practical guidance for ML data generation, and cross-checks ANDES and PSID.jl on a modified 14-bus case with GFM droop and virtual synchronous machine controls. If the comparison is fair, the practical takeaway is that simulator choice can change the cost of an ML dataset by an order of magnitude.","feed_headline":"Fastest open-source grid simulator for AI data: PSID.jl","feed_subtitle":"A 14-bus benchmark puts PSID.jl at 0.05 s per run and 31 s per 1,000 samples versus minutes for rivals.","key_machinery":"The load-bearing device is the two-table comparison. Table I catalogs base functions—programming language, unbalanced modeling, power flow, small-signal stability analysis, TDS type, renewable and inverter model library, and parallel-computing support—so a reader can see what each tool can simulate. Table II catalogs ML-facing features—manual quality, ease of code-based parameter modification and result retrieval, external data formats, and simulation time per one run and per 1,000 samples—and supplies the quantitative ranking. The numerical anchor is a quasi-static phasor (QSP) simulation of a modified IEEE 14-bus system, a fast phasor-domain approximation of transmission network dynamics, run with a 0.01 s step size and a 10 s horizon, with generator 4 tripped at 1 s while generator 5 is set successively as a synchronous generator, a GFM droop controller, and a GFM virtual synchronous machine. The GFM model families (droop, VSM, dVOC, and matching control) are the shared objects whose support determines each simulator's relevance to this workload.","core_discovery":"The paper's central claim is that, viewed through the needs of machine-learning data generation, the five surveyed open-source simulators differ substantially and measurably, and that the differences are captured in two comparison tables plus a small numerical case study. The decisive numbers are the simulation-time rows of Table II: for a modified IEEE 14-bus system with step size 0.01 s and a 10 s horizon, PowerSimulationsDynamics.jl takes 0.05 s per run and 31.1 s for 1,000 samples, compared with ANDES at 1.27 s per run and 224 s per 1,000 samples and Dynaωo at 0.34 s per run and 374 s per 1,000 samples; OpenDSS and GridLAB-D are marked not applicable because they target distribution systems. The paper frames this as practical guidance for ML professionals, and supports it with the observation that ANDES and PSID.jl produce similar GFM-Droop and GFM-VSM trajectories on the same 14-bus test case, suggesting that the speedup does not come from modeling a different system.","pith_inferences":["Because the paper does not report hardware, DAE solver choice, or parallelization settings, the cleanest reading is that the ranking is a demonstration that large speedups exist, not a guarantee on every machine; users should re-time on their own cluster before scaling up.","A natural extension the paper leaves implicit is to benchmark the same tools on larger systems and on fault scans that sample many locations, since ML datasets are rarely built from a single trip scenario.","The comparison highlights an unresolved trade-off: QSP-based tools are fastest, but ML models that must learn converter switching or fast electromagnetic transient dynamics would need Dynaωo's quasi-EMT mode, which is much slower at 374 s per 1,000 samples; the paper does not quantify that trade-off.","If the ranking holds across system sizes, the choice of simulator is itself a practical hyperparameter of any ML pipeline, and reporting it alongside model performance would make power-system ML studies more reproducible."],"forward_implications":["For transmission-level machine-learning data generation at the scale of a 14-bus-like system, PowerSimulationsDynamics.jl cuts single-run cost by roughly 6–25× and 1,000-sample cost by roughly 7–12× compared with the other timed simulators, assuming the benchmark conditions hold.","ANDES, despite its slower single run, remains in contention for 1,000-sample batches when multi-CPU parallelization is used, as the paper explicitly notes.","OpenDSS and GridLAB-D are positioned for distribution-system studies rather than the transmission QSP workloads typical of bulk-system ML datasets, so the timing comparison intentionally excludes them.","The agreement between ANDES and PSID.jl trajectories under GFM droop and GFM VSM control supports using either of the two as a dynamical reference while picking the faster one for bulk data generation."],"supporting_citations":[{"why":"Supplies the TDS classification into quasi-static phasor and EMT simulations that frames the whole comparison.","marker":"[1]"},{"why":"Establishes that machine-learning power-system research depends on simulator-generated data.","marker":"[2]"},{"why":"Documents the reinforcement-learning side of the simulator interaction loop.","marker":"[3]"},{"why":"Source for ANDES' capabilities, GFM model library, and the hybrid symbolic-numeric modeling approach.","marker":"[4]"},{"why":"Source for OpenDSS' dynamic modes, customization via DLLs, and external interface features.","marker":"[11]"},{"why":"Source for GridLAB-D's distribution-oriented simulation capabilities and GFM droop support.","marker":"[12]"},{"why":"Provides the GFM control taxonomy (droop, VSM, matching, VOC/dVOC) used to compare model support.","marker":"[17]"},{"why":"Supplies the characterization of GFM inverters as controllable voltage sources that motivates the survey.","marker":"[19]"},{"why":"Source for PowerSimulationsDynamics.jl's model library and reported simulation performance.","marker":"[22]"},{"why":"Source for Dynaωo's submodules and GFM model set, including quasi-EMT simulation.","marker":"[23]"}],"fun_headline_variants":["PSID.jl leads open-source grid sims for ML data speed","Grid sim speed test: PSID.jl 31s vs ANDES 224s for 1k samples","Why PSID.jl is the top pick for ML grid simulations","Open-source sim shootout: PSID.jl wins for AI data","For ML on power grids, PSID.jl runs 1000 samples in 31s"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the timing comparison in Table II is fair and representative: the simulators ran on comparable hardware with comparable solvers, the IEEE 14-bus case mirrors typical machine-learning data-generation workloads, and wall-clock time is the right criterion for suitability.","fun_headline_variants_meta":{"raw":{"variants":["PSID.jl leads open-source grid sims for ML data speed","Grid sim speed test: PSID.jl 31s vs ANDES 224s for 1k samples","Why PSID.jl is the top pick for ML grid simulations","Open-source sim shootout: PSID.jl wins for AI data","For ML on power grids, PSID.jl runs 1000 samples in 31s"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000975,"raw_usage":{"total_tokens":4105,"prompt_tokens":872,"completion_tokens":3233,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":3124}},"tokens_in":488,"tokens_out":3233,"duration_ms":21259,"temperature":1.0,"reasoning_tokens":3124,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:14:58.653259+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 1,000-sample benchmark on the same IEEE 14-bus case on identical hardware with each simulator's documented solver and parallelization settings; if PowerSimulationsDynamics.jl no longer produces the shortest wall-clock time by the reported margin, the paper's ranking fails. A weaker check: repeat the comparison on a larger bus system or with a randomized fault scan and see whether the ordering changes.","supporting_citations":[{"cited_title":"Revisiting power systems time-domain simulation methods and models,","cited_arxiv_id":null,"evidence_quote":"Supplies the TDS classification into quasi-static phasor and EMT simulations that frames the whole comparison."},{"cited_title":"Machine learning driven smart electric power systems: Current trends and new perspectives,","cited_arxiv_id":null,"evidence_quote":"Establishes that machine-learning power-system research depends on simulator-generated data."},{"cited_title":"Reinforcement learning for selective key applications in power systems: Recent advances and future challenges,","cited_arxiv_id":null,"evidence_quote":"Documents the reinforcement-learning side of the simulator interaction loop."},{"cited_title":"Hybrid symbolic-numeric framework for power system modeling and analysis,","cited_arxiv_id":null,"evidence_quote":"Source for ANDES' capabilities, GFM model library, and the hybrid symbolic-numeric modeling approach."},{"cited_title":"An open source platform for collaborating on smart grid research,","cited_arxiv_id":null,"evidence_quote":"Source for OpenDSS' dynamic modes, customization via DLLs, and external interface features."},{"cited_title":"GridLAB-D: an agent-based simulation framework for smart grids,","cited_arxiv_id":null,"evidence_quote":"Source for GridLAB-D's distribution-oriented simulation capabilities and GFM droop support."},{"cited_title":"Frequency stability of synchronous machines and grid-forming power converters,","cited_arxiv_id":null,"evidence_quote":"Provides the GFM control taxonomy (droop, VSM, matching, VOC/dVOC) used to compare model support."},{"cited_title":"Grid-forming inverters: Are they the key for high renewable penetration?","cited_arxiv_id":null,"evidence_quote":"Supplies the characterization of GFM inverters as controllable voltage sources that motivates the survey."},{"cited_title":"Towards an open-source solution using modelica for time-domain simulation of power systems,","cited_arxiv_id":null,"evidence_quote":"Source for Dynaωo's submodules and GFM model set, including quasi-EMT simulation."}],"review_version":1}