{"id":"3288a2e6-4531-4725-b1bd-f3b4991665a8","arxiv_id":"2501.09513","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A boundary-focused dataset generation method combining separating hyperplanes and directed walks improves the generalization of decision-tree classifiers for dynamic security assessment on two test systems.","lead":"The paper presents a toolbox that generates power system operating points close to the security boundary by combining two previously published sampling algorithms. The authors show that decision trees trained on such boundary-focused datasets classify both near-boundary and distant operating points more reliably than models trained with naive or importance sampling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal role of boundary-adjacent samples is not isolated from other dataset differences; controlled resampling is needed before the central claim is supported.","rationale":"The reader identified the directed walk's smoothness assumption as the weakest point. That is a legitimate method-reliability concern, but the paper's own figures show the walk achieves 79-85% HIC samples on the two test systems, so the assumption held well enough for the reported experiments. The more load-bearing issue for the central claim is the evaluation design: the compared datasets differ in multiple ways beyond HIC share, including input-space reduction, projection of infeasible points, and class balance. Table III quantifies how strongly balance alone affects F1, so without controlling for these factors the observed performance advantage cannot be causally attributed to boundary sampling. The proposed concrete test isolates the HIC share while holding other dataset properties fixed, which directly addresses the causal claim. The reader's concern and the one raised here are complementary; hence partial agreement. The verdict remains conditional: the paper's conclusion is plausible but needs a controlled HIC-share experiment to be accepted as robust.","tokens_in":22340,"tokens_out":12990,"duration_ms":121210,"concrete_test":"Construct resampled versions of the proposed method's dataset (or the LHC dataset) that match the original in size and secure/insecure balance but vary the HIC share from roughly 10% to 85% by dropping or adding HIC-region points from the same source. Train the identical CART configuration (max depth 5, ccp-alpha 0.01) on each version and evaluate on the held-out LHC and importance test sets. If F1 increases monotonically with HIC share, the central claim gains support; a flat or non-monotonic relationship would show that boundary concentration is not the active ingredient.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that datasets with a substantial share of OPs near the security boundary significantly enhance the performance of data-driven DSA tools. The supporting evidence in Tables I and II compares decision trees trained on datasets that differ not only in HIC share but also in other properties: the proposed method restricts sampling to a convex polytope constructed via OBBT and separating hyperplanes, uses a different infeasible-point projection, and its directed walk induces path-dependent clustering. Table III shows that rebalancing a proposed-method dataset alone changes F1 scores by roughly 0.2 to 0.3, demonstrating that dataset balance is a major confound. Because the three training datasets differ along multiple axes, the observed generalization advantages of the proposed method's decision tree cannot be attributed specifically to the 'substantial share of OPs near the security boundary' emphasized in the conclusion. Furthermore, the boundary test set is generated by the same algorithm, so it is not an independent probe of boundary coverage. This is load-bearing because the paper's main causal claim is inferred from a comparison across methods rather than from a controlled manipulation of the HIC share.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a dataset generation toolbox for dynamic security assessment (DSA) that combines optimization-based bound tightening, separating hyperplanes, and a directed-walk algorithm to produce operating points (OPs) near the small-signal stability and AC-feasibility boundary. The authors compare decision trees trained on datasets from their method, Latin Hypercube (LHC) sampling, and importance sampling on the PGLib-OPF 39-bus and 162-bus systems, reporting F1 scores on test sets from each sampling method plus an additional boundary-focused test set. The central claim, stated in the conclusion, is that datasets with a substantial share of boundary-adjacent OPs significantly enhance the performance of data-driven DSA tools, provided the dataset is balanced between secure and insecure points.","tokens_in":22594,"tokens_out":6570,"duration_ms":64872,"significance":"If the central claim were established, the work would be practically important: DSA dataset generation is a recognized bottleneck, and the authors contribute a modular, publicly available toolbox and a reproducible comparison on two standard test systems. The misclassification analysis in Figs. 9 and 10, showing that errors concentrate near the security boundary, is a useful diagnostic that supports the broader motivation. However, the experiments as designed do not isolate the effect of boundary-adjacent samples from other dataset differences, so the main causal claim is not yet supported by the evidence presented.","major_comments":[{"comment":"The central claim that a 'substantial share of OPs near the security boundary' improves DSA performance is inferred from comparisons of decision trees trained on three pipelines that differ along multiple axes. The proposed method restricts sampling to an OBBT/hyperplane polytope (Section III-A), projects infeasible points through a different optimization (Section III-B), and produces path-dependent clusters via directed walks (Section III-C), while the LHC and importance benchmarks use their own sampling and projection mechanisms. The observed F1 differences therefore cannot be attributed specifically to the HIC-region share. Table III compounds this concern: rebalancing a proposed-method dataset changes F1 scores by roughly 0.2 to 0.3, showing that dataset balance is a major confound that is not controlled for in Tables I and II. A controlled experiment that varies the HIC share while holding the rest of the pipeline fixed (for example, resampling a single base dataset into variants with 0%, 20%, 50%, and 80% HIC points) is necessary before the conclusion in Section VI can be supported.","section":"IV-D, Tables I-III"},{"comment":"The 'Boundary' test set is generated by the proposed method, as stated in Section IV-D2 ('This test set contains OPs generated by the proposed method which lie near the security boundary'). This makes it a distributionally aligned test set rather than an independent probe of boundary classification; the proposed method's higher boundary F1 could reflect training/test distribution overlap rather than better classification of the true boundary. The authors should construct a boundary test set independently of the proposed method, for example by densely sampling a narrow damping-ratio band with LHC or importance sampling, or by evaluating on a hold-out set generated by a separate boundary-characterization procedure.","section":"IV-D2"},{"comment":"The directed-walk algorithm assumes that the damping ratio of the least-damped mode is smooth enough in generator-active-power space that a local finite-difference gradient, combined with step-size reduction and a 1 MW discretization, reliably walks toward the security boundary without missing disconnected boundary components. No evidence is provided for this smoothness or connectivity assumption, and the paper does not quantify how completely the proposed method covers the boundary. Because the method's advertised ability to 'capture' the boundary depends on this assumption, the authors should either provide a quantitative coverage analysis (for example, the fraction of LHC boundary points within a small distance of a proposed-method HIC point, across multiple random initializations) or explicitly discuss the assumption as a scope condition.","section":"III-C1, Eqs. (12)-(13)"}],"minor_comments":[{"comment":"Tables I and II are typeset in a way that separates row labels from data (for example, 'Proposed MethodTraining Testing' appears as disconnected text), making the tables difficult to read; the row and column structure should be redrawn.","section":"Tables I and II"},{"comment":"Figures 4-8 show axis labels as unicode placeholders (for example, '/uni00000012/uni00000010/uni00000012') in the preprint; vector graphics with proper generator names are needed.","section":"Figures 4-8"},{"comment":"Section IV-D4 says 'The LHC and importance sampling datasets used for testing are shown in Fig. 3,' but Fig. 3 reports percentages of feasible, stable, secure, and HIC-region samples, not the test sets themselves; the text should be clarified.","section":"IV-D4"},{"comment":"The step-size distance thresholds d1, d2, and d3 in Eq. (12) are not reported; only the values of epsilon_1 through epsilon_4 are given, so the directed-walk parameters are not fully reproducible.","section":"III-C1, Eq. (12)"},{"comment":"Figure 1 refers to 'Section IV.A' and subsequent sections, but the corresponding method descriptions are in Section III; the cross-references should be corrected.","section":"Fig. 1"},{"comment":"There are typos such as 'missclassificaton' in Section IV-D3 and 'anayzed' in Section IV-D; a proofreading pass is needed.","section":"IV-D3 and elsewhere"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant problem and provides a useful open-source toolbox, but the experimental design does not yet isolate the variable named in the central claim. The confound identified in the major comments is load-bearing, so I cannot recommend acceptance in the current form. The paper would be substantially strengthened by a controlled resampling experiment and an independent boundary test set."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this paper delivers a genuinely useful public toolbox for boundary-focused DSA dataset generation, and its benchmark results are the first systematic evidence I know of that such sampling helps decision trees generalize. But the headline claim—that a substantial share of boundary-adjacent OPs enhances performance—is not actually isolated from other differences between the datasets, so the causal story is weaker than the conclusion states.\n\nWhat's new: the toolbox combining separating hyperplanes and directed walks, publicly available, and the systematic comparison on PGLib 39- and 162-bus systems. The decision tree trained on the proposed dataset generalizes well to LHC and importance-sampling test sets, which is a meaningful result: a tree trained mostly near the boundary still handles points far from it. The misclassification analysis is nice—errors concentrate near the boundary, which supports the intuition. And the balancing experiment (Table III) shows balance matters a lot; that's a useful practical note.\n\nSoft spots, in proportion. The stress-test concern about confounds is fair. The proposed datasets differ from the benchmarks not just in HIC share but also in being restricted to a convex polytope built via OBBT and separating hyperplanes, in using a different infeasible-point projection, and in having path-dependent clustering from directed walks. Table III shows that rebalancing alone shifts F1 by 0.2–0.3, so dataset composition is a real confound. The conclusion's causal wording goes beyond what a three-way comparison can support. The boundary test set is also generated by the same method, so it is not an independent probe.\n\nMinor but worth noting: no confidence intervals; hyperparameters like β and the ε scalars were tuned on pre-tests; only two small systems; and the directed-walk smoothness assumption (finite-difference gradient in generator-space) is plausible but not validated for fragmented boundaries—the figures actually hint at a dispersed boundary in the 39-bus case.\n\nWho this is for: power-systems researchers doing data-driven DSA or synthetic dataset generation. They'll get a working toolbox, a clear baseline, and a reasonable set of experiments. The paper deserves a serious referee, and I'd send it out with a request to temper the causal claim and ideally add a controlled resampling ablation—e.g., reweight LHC or importance datasets to match the proposed method's HIC share and see if the gains persist. That would turn a suggestive result into a convincing one.","headline":"A useful public toolbox and a suggestive but not yet airtight case that boundary-focused sampling improves DSA classifiers.","tokens_in":23122,"tokens_out":3133,"would_cite":true,"duration_ms":29929,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the quality of a data-driven dynamic security assessment tool is set by how well its training data covers the security boundary, and it presents a toolbox that concentrates samples there while keeping secure and…","keywords":["dynamic security assessment","dataset generation","security boundary","small-signal stability","directed walks","separating hyperplanes","decision trees","operating point sampling"],"falsifier":"On a two-generator slice of the 39-bus system, exhaustively grid the active-power plane at 1 MW resolution, compute the true high-information-content region from the damping ratio, and check whether the directed walks launched from all initialization points visit every connected component; if any boundary component is never entered, the method's claim to comprehensive boundary capture is falsified.","tokens_in":22133,"feed_emoji":"⚡","tokens_out":10028,"duration_ms":85660,"temperature":0.7,"pith_summary":"This paper tries to establish that synthetic training data for dynamic security assessment is only as good as its coverage of the security boundary, the region where operating points switch from secure to insecure. It proposes a dataset-generation method that concentrates samples inside a narrow band around the small-signal stability boundary, while balancing secure and insecure labels, through bound tightening, separating hyperplanes, and directed walks. On 39-bus and 162-bus test systems, a decision tree trained with 85% and 79% of its samples inside that band generalizes across test sets produced by naive and importance sampling, and misclassifications concentrate near the boundary. The conclusion is that a large share of boundary-adjacent operating points improves data-driven dynamic security assessment, and that dataset balance matters as much as boundary proximity.","feed_headline":"Near-boundary training data lifts grid security classifier accuracy","feed_subtitle":"Walking generator setpoints to the security boundary yields training data that beats naive and importance sampling.","key_machinery":"The central mechanism is the directed walk in generator active-power space, which uses the finite-difference gradient of the damping ratio of the least-damped mode as a direction of travel and a step size that shrinks as the operating point approaches the boundary (Eqs. 12 and 13). The walks start from feasible dispatches produced by a space-reduction pipeline: optimization-based bound tightening tightens the input bounds, separating hyperplanes cut away provably infeasible volumes, and Hit-and-Run sampling proposes candidates inside the remaining convex polytope. When a walk enters the high-information-content region, the algorithm sweeps a surrounding 1 MW grid to collect neighboring operating points, turning a few feasible dispatches into a dense, balanced boundary dataset.","core_discovery":"The paper's central claim is that operating points close to the security boundary carry most of the discriminating information a classifier needs, and that a dataset deliberately enriched in such points, while balanced between secure and insecure labels, makes a data-driven dynamic security assessment tool accurate near the boundary without sacrificing accuracy far from it. The authors build such datasets by first shrinking the feasible operating region with optimization-based bound tightening and separating hyperplane infeasibility certificates, then walking feasible dispatches toward the boundary using a finite-difference gradient of the least-damped mode's damping ratio, and finally collecting points inside a margin of $\\pm 0.25\\%$ damping around the boundary, defined here as AC feasibility plus a minimum damping of $\\zeta_{\\min}=3\\%$. In their case studies, decision trees trained on these datasets show F1-scores that generalize across test sets from other sampling methods, while trees trained on uniform or importance-sampled data degrade on boundary test sets. The stated conclusion is that the high-information-content region must be explicitly targeted and the dataset kept balanced for data-driven dynamic security assessment to work well.","pith_inferences":["The paper's unimodal importance-sampling benchmark suggests a testable hypothesis: a mixture or copula-based importance distribution fitted to multiple boundary components would recover much of the gap to the proposed method, since the visible failure is a single Gaussian's inability to follow a curved or dispersed boundary.","Because the directed walk finds the boundary component nearest each initialization point, the number and placement of feasible initialization dispatches acts as a tunable coverage knob; random restarts or wider polytope sampling would make boundary coverage explicit rather than incidental.","The 1 MW discretization inside the high-information-content region sets a resolution floor for the decision boundary the classifier can learn; reducing the grid spacing would sharpen boundary accuracy at a higher sampling cost.","The boundary-region density could be reused as an active-learning acquisition function, proposing new dispatches where the current classifier's margin is smallest and potentially lowering the simulation budget while keeping boundary coverage high."],"forward_implications":["A decision tree trained with 85% of its samples inside the high-information-content region on the 39-bus system, and 79% on the 162-bus system, still classifies far-from-boundary operating points accurately, so boundary-focused training data does not trade away far-field accuracy.","Misclassified operating points cluster inside the high-information-content region even for the boundary-trained tree, which means residual errors live where the label flips and are not cured by adding more far-field samples.","Balancing secure and insecure labels in the proposed-method dataset raises cross-test F1-scores, for example from 0.59 to 0.85 on the 162-bus boundary test set, so boundary enrichment alone is not enough.","The Latin hypercube tree outperforms the importance tree on the proposed method's test set for the 39-bus system, indicating that a unimodal Gaussian importance distribution can miss a fragmented boundary.","Because the directed walk needs only a scalar distance-to-boundary measure, the same toolbox extends to voltage stability or transient stability by swapping the damping-ratio sensitivity for another index."],"supporting_citations":[{"why":"provides the directed walk algorithm that moves operating points toward the security boundary and is extended here to sample the high-information-content region.","marker":"[14]"},{"why":"supplies the separating hyperplane infeasibility certificates and polytope construction that reduce the input search space before the walks.","marker":"[17]"},{"why":"defines the optimization-based bound tightening used to tighten voltage and angle bounds and sharpen the convex relaxation.","marker":"[19]"},{"why":"introduces the quadratic convex relaxation used to certify feasibility and to project infeasible dispatches onto the feasible region.","marker":"[22]"},{"why":"provides the Hit-and-Run sampling method used to draw candidate points uniformly inside the reduced convex polytope.","marker":"[24]"},{"why":"supplies the 39-bus and 162-bus test systems on which the case studies benchmark the proposed method against naive and importance sampling.","marker":"[25]"},{"why":"provides the generator, voltage regulator, governor, and power system stabilizer models used for small-signal stability assessment.","marker":"[26]"},{"why":"defines the classification and regression tree algorithm used to evaluate the datasets in the case studies.","marker":"[33]"}],"fun_headline_variants":["Boundary-focused sampling boosts grid security classifier accuracy","Near-boundary data key to dynamic security assessment accuracy","Sampling at security boundary lifts DSA model performance","Grid security classifiers improve with boundary-enriched data","Balanced data near security boundary sharpens DSA predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the system's most fragile oscillation mode changes smoothly enough as generator outputs are adjusted that a step-by-step gradient walk, with shrinking step sizes and a 1 MW grid, reliably reaches the security boundary and does not skip over fragmented parts of it.","fun_headline_variants_meta":{"raw":{"variants":["Boundary-focused sampling boosts grid security classifier accuracy","Near-boundary data key to dynamic security assessment accuracy","Sampling at security boundary lifts DSA model performance","Grid security classifiers improve with boundary-enriched data","Balanced data near security boundary sharpens DSA predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000152,"raw_usage":{"total_tokens":1199,"prompt_tokens":936,"completion_tokens":263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":188}},"tokens_in":552,"tokens_out":263,"duration_ms":3712,"temperature":1.0,"reasoning_tokens":188,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:56:31.078033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a two-generator slice of the 39-bus system, exhaustively grid the active-power plane at 1 MW resolution, compute the true high-information-content region from the damping ratio, and check whether the directed walks launched from all initialization points visit every connected component; if any boundary component is never entered, the method's claim to comprehensive boundary capture is falsified.","supporting_citations":[{"cited_title":"Efficient database generation for data-driven security assessment of power sys- tems,","cited_arxiv_id":null,"evidence_quote":"provides the directed walk algorithm that moves operating points toward the security boundary and is extended here to sample the high-information-content region."},{"cited_title":"Strengthening convex relaxations with bound tightening for power network optimization,","cited_arxiv_id":null,"evidence_quote":"defines the optimization-based bound tightening used to tighten voltage and angle bounds and sharpen the convex relaxation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the Hit-and-Run sampling method used to draw candidate points uniformly inside the reduced convex polytope."},{"cited_title":"Powersystems. jl—a power system data management package for large scale modeling,","cited_arxiv_id":null,"evidence_quote":"provides the generator, voltage regulator, governor, and power system stabilizer models used for small-signal stability assessment."},{"cited_title":"Classification and regression trees,","cited_arxiv_id":null,"evidence_quote":"defines the classification and regression tree algorithm used to evaluate the datasets in the case studies."}],"review_version":1}