{"id":"a586f47b-fda2-4a34-8cf4-8fdc50de0add","arxiv_id":"2510.26307","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes HGNN-based cybersecurity anomaly detection by node/edge/subgraph anomalies and static/dynamic graphs, but the central taxonomy tables are absent from the text.","lead":"This paper is a literature survey of heterogeneous graph neural network (HGNN) methods for detecting cybersecurity anomalies, classifying approaches by anomaly type and graph dynamics. A generalist might read it to get a structured map of a fragmented research area, but the promised taxonomy tables are missing and several claims are internally inconsistent.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central contribution — the taxonomy and comparative Tables 1–6 — is announced but not present in the manuscript; without those tables, the survey's core artifact cannot be inspected or verified.","rationale":"I read the paper as an attempt to provide a field-reference survey whose value depends on a reliable taxonomy and on comparative tables that let readers see how models, datasets, and metrics relate. The most load-bearing condition is that those comparative artifacts actually exist and are accurate. The manuscript does not satisfy that condition: Tables 1–6 are referenced in the text but only captions appear, so the classification and comparisons cannot be checked. This is more immediate than the reader's stated weakest assumption about representative sampling: even if the selected models were representative, the organizational output that would demonstrate this is absent. The reader's rationale already notes the missing artifact, so my agreement is partial rather than full. I do not see this as a disagreement with the field, and I am not questioning the authors' intent; the issue is purely that the central evidence is not in the manuscript. There is no formal verification, code release, or independent reproduction that could substitute. Therefore the REJECT verdict remains appropriate, though the concern could be resolved by adding the missing tables in a revision.","tokens_in":22336,"tokens_out":2382,"duration_ms":25558,"concrete_test":"Inspect the submitted PDF directly: locate 'Table 1' through 'Table 6' and record whether each has a visible tabular body (rows/columns) or only a caption line. Also count the reference list entries and compare to the §7 claim of 89 representative papers. If the tables contain no data, the central survey artifact is missing and the REJECT verdict stands; if the tables are present with substantive rows, the reader should reconsider.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The stated contribution is a structured taxonomy and comparative analysis that lets readers see how methods differ. The manuscript repeatedly points to Tables 1–6 (e.g., 'Table 1 presents the proposed taxonomy...'; 'Table 2 presents this taxonomy...'), but no table bodies appear in the text — only captions. Thus the load-bearing artifact is missing. The prose around methods is too high-level to substitute: it asserts that DOMINANT, HeCo, AddGraph, etc. fall into categories, but the actual mapping, supervision/architecture columns, and dataset/metric comparisons are not delivered. Relatedly, §7 claims 'reviewed over 100+ studies and cited 89 representative papers,' while the bibliography contains 86 numbered references with duplicates (e.g., [10] and [68] both cite HGT; [23] and [73] both cite UNSW-NB15). This does not by itself invalidate the prose, but it undermines the survey's reliability as a field reference and makes the 'comprehensive' claim unverifiable. The taxonomy may be salvageable, but in this version the central deliverable is absent, so the paper cannot function as the structured reference it claims to be.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of heterogeneous graph neural network (HGNN) methods for anomaly detection in cybersecurity. It proposes a taxonomy organized by anomaly granularity (node-level, edge-level, subgraph-level) and graph dynamics (static versus dynamic), discusses representative models in Section 3, maps them to four cybersecurity application domains in Section 4, reviews evaluation metrics and benchmark datasets in Section 5, and lists open challenges in Section 6. The stated contribution is to provide a structured, comparative reference for a fragmented literature, with Tables 1–6 as the vehicle for that comparison.","tokens_in":22595,"tokens_out":9670,"duration_ms":96051,"significance":"A reliable survey of HGNN-based anomaly detection for cybersecurity would fill a genuine gap: the field is fragmented, lacks standardized benchmarks, and is difficult to navigate. The paper's prose identifies relevant organizing dimensions—structural versus semantic versus temporal anomaly evidence, label scarcity, class imbalance, and evaluation inconsistency—and the high-level taxonomy is sensible. It does not offer new models, proofs, or code; its value is entirely organizational. That value is currently unrealized because the tables and accurate method-to-reference mappings are not present, so readers cannot verify the survey's central claims.","major_comments":[{"comment":"Tables 1–6 are captions only; no table body appears. The text repeatedly refers to these tables as the location of the proposed taxonomy and the comparative analyses (e.g., §3: 'Table 1 presents the proposed taxonomy...' and 'Table 2 presents this taxonomy...'; §5: 'Table 5 summarizes the core evaluation metrics' and 'Table 6 summarizes the key characteristics of each benchmark dataset'). Without the table bodies, readers cannot inspect which models are assigned to which category, what supervision/temporal/architecture features are compared, or how datasets are characterized. This is not a presentational issue; it is the paper's central deliverable, and it is missing.","section":"§3 and §5, Tables 1–6"},{"comment":"The conclusion states 'We reviewed over 100+ studies and cited 89 representative papers.' The bibliography contains 86 numbered references, and at least two are duplicates: [10] and [68] are both Hu et al., 'Heterogeneous Graph Transformer,' and [23] and [73] are both Moustafa & Slay, 'UNSW-NB15.' The unique cited count is therefore ≤84. No systematic search or inclusion criteria are reported, so the 'comprehensive' claim is unverifiable and the 89/100+ statement is internally inconsistent.","section":"§7 and References"},{"comment":"Several dataset citations are wrong. In §4.1, 'CERT [72]' points to Brown et al., 'Recurrent Neural Network Attention Mechanisms for Interpretable System Log Anomaly Detection,' not to the CERT insider-threat dataset; CERT appears correctly as [71] (and [22]). In §5.2.2, CTU-13 is cited as [76], but [76] is the DARPA/Unicorn paper; CTU-13 is [75]. These mismatches make the dataset-to-model mappings unreliable.","section":"§4.1, §5.2.2, References"},{"comment":"Several model descriptions are not supported by the cited references. SemiGNN is introduced in §3.1 as a fraud-detection model with hierarchical attention over metapath subgraphs, but [39] is 'SemiGNN-PPI,' a protein–protein interaction prediction paper. Bi-GCN in §3.2 is described as a bidirectional message-passing rumor-detection model, but [53] is 'Bi-GCN: Binary Graph Convolutional Network' for image classification. AANE in §3.2 is said to be an adversarial autoencoder, but [55] is 'Accelerated Attributed Network Embedding.' HON-GAT in §3.3 is described as a motif-based subgraph anomaly detector, but [61] is 'HONGAT: Graph Attention Networks in the Presence of High-Order Neighbors' for node classification. These are not citation typos; they misassign methods within the taxonomy that is the survey's main contribution.","section":"§3.1–§3.3"},{"comment":"The 'CAIDA APT simulation dataset [77]' is cited to 'APT datasets and attack modeling for automated detection methods: A review' (Stojanović et al.), which is a survey, not a dataset. The sentence 'OCAN has been evaluated on the CERT Insider Threat Dataset [71]' attributes an evaluation to OCAN that the cited OCAN paper (one-class adversarial nets) does not report; CERT is not among that paper's benchmarks. These errors further reduce confidence in the application mapping in Section 4.","section":"§4.4"}],"minor_comments":[{"comment":"The same sentences are repeated verbatim, and an incomplete placeholder 'such as ....' appears in §2.1.2. The schema example in §2.1.3 is also stated twice.","section":"§2.1.2, §2.1.3"},{"comment":"'the released by Los Alamos National Laboratory' is a grammatical fragment with a missing dataset name.","section":"§4.1"},{"comment":"Duplicates [10]=[68] and [23]=[73] should be merged. Several references have incomplete metadata, e.g., [41], [60], and [86].","section":"References"},{"comment":"The text uses 'HON-GAT' while reference [61] is 'HONGAT'; use one consistent acronym.","section":"§3.3"},{"comment":"'eBay transaction graph [85]' cites the eBay product dataset, which is not a transaction graph with fraud labels; the actual dataset should be identified or the sentence revised.","section":"§5.2.2"},{"comment":"The abstract and contributions list call the taxonomy 'novel,' but §2.4 explicitly says the graph-centric taxonomy is adopted from prior surveys [6,7,17]. Please soften or justify the novelty claim.","section":"Abstract, §1, §2.4"}],"recommendation":"major_revision","confidential_remarks":"I considered rejecting on the grounds that Tables 1–6 are absent and are the survey's core artifact. I recommend major revision rather than rejection because the organizational framework is coherent and the defects—supplying the table contents, correcting the method-to-reference mappings, and fixing the count inconsistency—are identifiable and potentially repairable. If the authors cannot supply the actual table contents, or if spot-checks of additional references reveal the same systematic mismatch, the paper should not proceed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper announces a structured taxonomy and a set of comparative tables as its main contribution, but the tables are not in the manuscript. Only their captions appear. That makes the survey unusable as a reference, and I think the right call is to send it back rather than send it to referees.\n\nWhat it does well: the prose organizes the field into node-, edge-, and subgraph-level anomaly detection and maps representative HGNN models to cybersecurity applications (insider threats, network intrusion, access fraud, APTs). The method summaries are mostly accurate, and the discussion of evaluation metrics and datasets, while high-level, is a reasonable starting point for someone new to the area. The taxonomy itself is explicitly inherited from earlier graph anomaly detection surveys, so there's no new conceptual contribution — but the synthesis of the cybersecurity-applied HGNN literature is useful.\n\nThe problems are load-bearing. Tables 1–6 are invoked throughout as the place where the taxonomy, model comparisons, supervision types, and dataset characteristics are laid out, but the table bodies are absent. The reader can't check which models are in which category, what their supervision signals are, or how they compare on any metric. That's not a nitpick; it's the central deliverable of a survey like this.\n\nThe citation and consistency issues are smaller but in the same direction. Section 7 claims 'over 100+ studies' while 86 references are listed. Some references are mis-cited (CTU-13 appears as [76] instead of [75]; CERT appears as [72] when the dataset is [71]). There are duplicated passages in Section 2 and a few malformed sentences that suggest hasty editing.\n\nNone of this invalidates the prose summaries, but together it means the paper cannot serve as the dependable structured reference it claims to be. A reader might use it for orientation, but they'd have to do their own verification for anything load-bearing.\n\nWho is this for? Junior researchers or practitioners who want a quick map of the HGNN-anomaly-detection-in-cybersecurity landscape. That's a real need. But this version doesn't meet it.\n\nRecommendation: desk reject with an invitation to resubmit once the tables are actually included, the reference list is cleaned up, and the '100+ studies' claim is made consistent with the bibliography. If those things happen, it could be a genuinely useful survey and worth peer review. As it stands, I wouldn't send it to referees.","headline":"Useful prose orientation, but the core comparative tables are missing — as a reference it currently doesn't work.","tokens_in":23037,"tokens_out":3424,"would_cite":false,"duration_ms":33409,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that the scattered field of HGNN-based cybersecurity anomaly detection can be organized by anomaly type (node, edge, subgraph) and by graph dynamics (static vs. temporal), and that the field's main bottleneck is the absen","keywords":["heterogeneous graph neural networks","anomaly detection","cybersecurity","taxonomy","insider threat detection","network intrusion detection","temporal graphs","benchmark datasets"],"falsifier":"A systematic literature search of security and graph-learning venues from 2019 to 2025 that finds a substantial share of HGNN-based cyber anomaly detection papers fitting none of the three anomaly-granularity classes, or requiring category boundaries to be redrawn, would refute the survey's completeness claim. Likewise, checking a model placed in the 'temporal' family and finding it handles no timestamp information would refute the taxonomy's accuracy.","tokens_in":22239,"feed_emoji":"🛡️","tokens_out":3688,"duration_ms":35651,"temperature":0.7,"pith_summary":"The paper tries to bring order to a fragmented research area: using heterogeneous graph neural networks (HGNNs) to detect cyber anomalies. Its central proposal is a taxonomy that classifies methods by the type of graph element they flag as anomalous—node, edge, or subgraph—and by whether they operate on static or dynamic graphs. It then maps representative models to four cybersecurity settings (insider threats, network intrusion, access-log fraud, advanced persistent threats) and reviews the benchmark datasets and metrics the field uses. The payoff the authors intend is a structured foundation: a reader can see what approaches exist, where they fit, what they are evaluated on, and where the gaps are. The survey's value stands on whether its chosen set of 89 representative papers and its categories actually cover the field.","feed_headline":"New taxonomy sorts graph-based cyber anomaly detectors","feed_subtitle":"Heterogeneous-GNN methods are classified by what they flag and whether they model time; benchmarks remain the bottleneck.","key_machinery":"The load-bearing object is the taxonomy itself: a two-axis classification (anomaly granularity × graph dynamics) with sub-groupings by learning strategy, presented through comparative tables that map each representative model to its supervision type, temporal capability, and architectural mechanism. The taxonomy does the argument's work by converting a scattered set of papers into comparable slots, which is also what lets the authors locate gaps such as the lack of unified static-temporal-semantic reasoning and the scarcity of realistic temporal benchmarks. The second supporting piece is the benchmark/metrics summary (AUROC, AUPRC, Precision@K, Recall@K, NDCG@K, F1), which grounds the claim","core_discovery":"On its own terms, the survey claims that the diverse HGNN-based anomaly detection literature can be organized along two axes—the granularity of the anomaly target (node, edge, or subgraph) and the presence of temporal dynamics (static vs. dynamic). Under this framing it groups node-level methods into reconstruction-based, attention/inconsistency-aware, contrastive, semi-supervised, and temporal families; edge-level methods into reconstruction, temporal, hierarchical, and adversarial/generative families; and subgraph-level methods into metapath/metagraph, structural-invariance/motif, temporal, one-class, knowledge-distillation, and hybrid families. It then shows how representative models such","pith_inferences":["A reader can use the node/edge/subgraph axis as a selection guide, but the taxonomy may be underdetermined: many real anomalies span levels (an anomalous node inside a clique), so a strictly categorical assignment may obscure the multi-resolution nature of attacks.","The survey's identification of 'static benchmarks' as a bottleneck suggests a concrete testable prediction: a dynamically annotated, heterogeneous benchmark suite would reshuffle current performance rankings of models like DOMINANT versus temporal HGNNs.","The absence of a systematic search protocol opens the door to extending this survey into a full systematic review; doing so would either validate or revise the claimed 89-paper coverage.","Practitioners could operationalize the gap analysis: a deployable system should prioritize handling temporal drift and partial observability, which the survey repeatedly singles out as missing from current evaluations."],"forward_implications":["If the taxonomy is right, new HGNN anomaly detection papers can be positioned in one of the categories, making cross-method comparison easier.","The claimed gaps imply that unified frameworks jointly modeling structure, semantics, and time are the field's open frontier.","The evaluation review implies that reported results on Cora, Amazon, CICIDS2017, and similar datasets are not directly comparable; adopting standardized benchmarks would change which models look best.","The application mapping implies that models from other domains (e.g., rumor detection) transfer to insider threat and APT detection, an insight available to subsequent research.","The survey's open-challenges section implies that scalability, interpretability, and streaming deployment should be part of future evaluation, not just detection accuracy."],"fun_headline_variants":["Survey: HGNNs for cyber anomaly detection, sorted by target and time","Taxonomy of heterogeneous GNNs for cyber anomaly detection, with gaps","Heterogeneous GNN survey: two-axis taxonomy, benchmark critique","Structured survey of graph neural networks for cybersecurity anomalies"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The survey assumes that its hand-selected set of 89 representative papers and its node/edge/subgraph categories faithfully and completely cover the field of HGNN-based cyber anomaly detection; if that sample is unrepresentative, the taxonomy and the derived gap analysis lose validity.","fun_headline_variants_meta":{"raw":{"variants":["Survey: HGNNs for cyber anomaly detection, sorted by target and time","Taxonomy of heterogeneous GNNs for cyber anomaly detection, with gaps","Heterogeneous GNN survey: two-axis taxonomy, benchmark critique","Structured survey of graph neural networks for cybersecurity anomalies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000245,"raw_usage":{"total_tokens":1373,"prompt_tokens":742,"completion_tokens":631,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":556}},"tokens_in":486,"tokens_out":631,"duration_ms":6361,"temperature":1.0,"reasoning_tokens":556,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:13:27.194758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search of security and graph-learning venues from 2019 to 2025 that finds a substantial share of HGNN-based cyber anomaly detection papers fitting none of the three anomaly-granularity classes, or requiring category boundaries to be redrawn, would refute the survey's completeness claim. Likewise, checking a model placed in the 'temporal' family and finding it handles no timestamp information would refute the taxonomy's accuracy.","supporting_citations":[],"review_version":1}