{"id":"f1e55247-817d-4bbb-b20a-3ff84ea4795e","arxiv_id":"2507.16541","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A data-centric taxonomy for Federated Graph Learning that classifies 79 studies by data characteristics and data utilization, plus a discussion of integration with pre-trained large models.","lead":"This survey reorganizes research on Federated Graph Learning (FGL) into a two-level data-centric taxonomy: what the data looks like and how methods use it. It covers 79 FGL papers and maps them by graph type, data distribution, client visibility, training phase, and data-related challenge.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Orthogonality claim is undercut by nested category definitions: ego-graph is defined as a subgraph instance and knowledge graph as a heterogeneous graph, so unique Table II placements are not guaranteed.","rationale":"The reader's verdict identified table-level inconsistencies such as duplicated entries, mismatched references, and an undefined phase name. I agree those are real, but I see a more load-bearing root cause: the survey's own definitions make some taxonomy categories nested, not orthogonal. Ego-graph-oriented FGL is explicitly called a specialized instance of subgraph-oriented FGL, and knowledge graphs are explicitly called specialized heterogeneous graphs. With nested categories, unique placement is impossible without a priority rule, and the 'orthogonal criteria' claim in the abstract and contribution (b) cannot be sustained as written. The survey has genuine strengths: a broad corpus of 79 works, concrete worked equations for HGFL, FedSpray, FedStar, FedGL, and FedPUB, and a useful data-centric framing for organizing FGL research. The flaw is conceptual and definitional rather than a matter of isolated sloppiness, but it is also fixable: the authors could either present the criteria as overlapping facets or add explicit disjointness and priority rules for Table II. Since the central organizing idea remains reasonable and the necessary corrections are editorial and definitional, I would keep the reader's CONDITIONAL verdict rather than moving to rejection. My concern partially overlaps with the reader's weakest assumption about correct, unambiguous placement, but it identifies the cause at the level of category definitions rather than at the level of table curation.","tokens_in":37150,"tokens_out":6450,"duration_ms":70877,"concrete_test":"Independently re-annotate the 79 entries of Table II from the primary sources, blind to the survey's table, using only the definitions in Sec. IV-A and IV-C. For each duplicate (at minimum FedGNN [78], FeSoG [34], FL-GMT [87]), record the data format and visibility stated in the original paper. If FedGNN's user-item graph is bipartite and not heterogeneous, its Heterogeneous/Subgraph placement is unsupported; if the duplicates are instead forced by nested definitions, then the categories overlap and the orthogonality claim is false. Report the overlap matrix; any non-zero overlap among format or visibility categories settles the issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and contribution (b), is that the two-level taxonomy is built from 'three orthogonal criteria' and categorizes 'all notable FGL studies in a fine-grained fashion.' For that claim to hold, the values within each criterion must be distinct enough that a study can be placed without ambiguity. The paper's own definitions violate this. Sec. IV-C defines ego-graph-oriented FGL as 'a specialized instance of subgraph-oriented FGL,' and Sec. IV-A defines a knowledge graph as 'a specialized heterogeneous graph.' Consequently, the Data Format and Visibility criteria are nested rather than orthogonal: a knowledge graph can simultaneously satisfy 'Heterogeneous' and 'Knowledge Graph,' and an ego-graph setting can satisfy both 'Subgraph-oriented' and 'Ego-graph-oriented.' Table II then places FedGNN [78] and FeSoG [34] under both Heterogeneous/Subgraph and Bipartite/Ego-Graph, and FL-GMT [87] under both Knowledge-Graph/Subgraph and Bipartite/Subgraph. These are not isolated typos; they are the predictable result of categories that overlap by definition. Table III compounds the problem by introducing an undefined 'Global Training' phase and using 'Global-side' instead of the defined 'Server-Side' dimension, so the sequential and positional axes cannot be applied consistently. The survey does not state a priority rule or a multi-label convention for Table II, so a reader cannot tell whether a repeated entry means 'method belongs to two configurations' or 'the taxonomy is not a partition.' Since the stated contribution is a fine-grained, orthogonal mapping, this unresolved ambiguity is load-bearing. The surveyed corpus itself and the worked examples in Sec. V are useful, but the headline claim needs either relaxed wording ('overlapping facets' instead of 'orthogonal criteria') or explicit rules for disjoint categorization.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript surveys federated graph learning (FGL) from a data-centric viewpoint. It proposes a two-level taxonomy: Data Characteristics (data format, decentralization format, visibility) and Data Utilization (motivational challenge, positional dimension, sequential phase). The authors claim that the three criteria in each level are orthogonal and that the taxonomy categorizes all notable FGL studies in a fine-grained way. The survey additionally reviews FGL applications across six domains, discusses integration with pretrained large models, and outlines future directions. It covers 79 FGL works, a larger volume than the three earlier surveys it compares against (7, 37, and 61 works).","tokens_in":37414,"tokens_out":5172,"duration_ms":51433,"significance":"The data-centric reframing is timely, and a reliable taxonomy would be a genuinely useful reference for researchers who want to locate FGL methods by data format, decentralization, visibility, training phase, or data-centric challenge. The breadth of coverage, the application survey, and the PLM-FGL discussion are valuable. However, the paper's central contribution currently rests on an orthogonality claim that the paper's own definitions refute, and the catalog tables contain internal inconsistencies. There are no derivations, machine-checked proofs, or falsifiable predictions to assess; the value of the paper depends on the clarity, consistency, and verifiability of its taxonomy and survey entries.","major_comments":[{"comment":"The claim that the three criteria in each taxonomy level are 'orthogonal' is undercut by the paper's own definitions. Sec. IV-C states that ego-graph-oriented FGL is 'a specialized instance of subgraph-oriented FGL,' and Sec. IV-A defines a knowledge graph as 'a specialized heterogeneous graph.' Consequently, the Data Format and Visibility criteria are nested rather than orthogonal. Table II contains repeated entries that follow directly from this overlap: FedGNN [78] and FeSoG [34] appear in both the Heterogeneous/Horizontal/Subgraph row and the Bipartite/Horizontal/Ego-Graph row, and FL-GMT [87] appears in both the Knowledge-Graph/Horizontal/Subgraph row and the Bipartite/Horizontal/Subgraph row. Since Sec. IV-D states no priority rule and no multi-label convention, a reader cannot determine which configuration a repeated method belongs to; the claimed fine-grained categorization is therefore not reproducible.","section":"Abstract and Sec. IV"},{"comment":"The paper claims in the abstract and in contribution (b) that the taxonomy categorizes 'all notable FGL studies in a fine-grained fashion,' and Sec. IV-D says Table II 'displays all categories.' Table II, however, realizes only 11 of the 4 x 2 x 3 = 24 possible combinations of data format, decentralization format, and visibility. Missing configurations include Heterogeneous/Graph-oriented, Knowledge-Graph/Vertical/Subgraph-oriented, and Bipartite/Vertical/Subgraph-oriented. The paper gives no explanation for these empty cells, and it does not state whether they are impossible, unpopulated, or out of scope. Without such a statement, the comprehensiveness claim of the first-level taxonomy is unsupported.","section":"Sec. IV-D and Table II"},{"comment":"The positional and sequential dimensions are defined in Sec. V-A as 'Client-Side' and 'Server-Side' and in Sec. V-B as Initialization, Local Training, Global Aggregation, and Post-aggregation. Table III instead uses 'Global-side' in the Positional Dimensions column and introduces a phase called 'Global Training' in the Data Privacy row. Neither 'Global-side' nor 'Global Training' is defined in Sec. V. The undefined phase and renamed dimension mean that the three criteria cannot be applied consistently to classify methods, which undermines the second-level taxonomy's central claim.","section":"Sec. V and Table III"},{"comment":"Several reference and naming inconsistencies prevent verification of the catalog. FedHG+ is cited as [113] in Sec. V-D1 and as [116] in Sec. V-D6; FedHGN is cited as [74] in Sec. V-D8 but as [116] in Table III; FedHGL is cited as [118] in Sec. V-D2 and as [158] in Sec. VII-F; and the entry 'nFedGNN[106]' appears in Table III without appearing in the text. These entries need to be reconciled with the reference list and with the names used in the body before the survey's classifications can be checked by readers.","section":"Table III and Sec. V-D"}],"minor_comments":[{"comment":"The text contains several grammatical and typographical errors, including 'This survey propose,' 'remains unadapted to reorganize FGL research,' and 'reconcile the tradeoff.' A careful language edit would improve readability.","section":"Abstract and Sec. I"},{"comment":"The 'Organization of the Survey' paragraph lists Secs. II, IV, V, VI, VII, VIII, and IX but omits Sec. III (Comparison with Other FGL Surveys).","section":"Sec. I"},{"comment":"The heading 'Open-wrold Graph Learning' contains a typo and should read 'Open-world Graph Learning.'","section":"Sec. IX-A3"},{"comment":"Several references are duplicated: [12] duplicates [3], [119] duplicates [52], and [117] duplicates [34] and [120]. These should be consolidated to a single citation each.","section":"References"},{"comment":"The relationship of Sec. VI (Euclidean-oriented FGL) to the two-level taxonomy is not stated explicitly, so it remains unclear whether this section is an extension of the taxonomy or a separate dimension.","section":"Sec. VI"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should know two things right away: the data-centric taxonomy is genuinely new relative to the three prior FGL surveys, and the central orthogonality claim does not hold as written. The paper is worth engaging with, but it needs a careful revision before anyone relies on its tables.\n\nWhat is actually good: the two-level taxonomy (Data Characteristics, Data Utilization) is a fresh organizational lens. Earlier surveys sort by scenario or method; this one maps studies onto data format, decentralization, visibility, and training phase. The worked examples in Sec. V (FedSpray, FedTAD, FedGTA) are concrete and well-chosen, and the PLM-FGL discussion is a useful pointer to a young area. The bibliography is broad (79 works in the main table), and Table I honestly compares coverage against the three earlier surveys. This is a useful reference for someone entering FGL from a data-centric angle.\n\nWhere it gets soft: the 'three orthogonal criteria' is load-bearing, and it fails. Sec. IV-C defines ego-graph FGL as 'a specialized instance of subgraph-oriented FGL,' and Sec. IV-A defines knowledge graphs as 'specialized heterogeneous graphs.' Those are nested categories, not orthogonal axes. The consequence shows up in Table II: FedGNN and FeSoG appear under both Heterogeneous and Bipartite rows, and FL-GMT under both KG and Bip. The paper does not state a priority rule or a multi-label convention, so the reader cannot tell whether a repeated entry means 'belongs to two configurations' or 'the taxonomy is not a partition.' Table III adds a second issue: it lists an undefined 'Global Training' phase instead of the defined 'Global Aggregation,' and uses 'Global-side' rather than the defined 'Server-Side.' The reference mismatches (FedHG+ under both [113] and [116]) are minor by comparison but add noise.\n\nThe survey's value does not depend on perfect orthogonality. If the authors relax the claim to 'overlapping facets' or specify explicit disjointness rules, the taxonomy still works as a reference. A stated paper-selection methodology would also strengthen the comprehensiveness claim, but that is fixable.\n\nBottom line: this is a solid draft for a subfield reference, not a breakthrough. It deserves a serious referee, but the referee should send it back for a consistency pass on the taxonomy definitions and tables. I would not cite it in its current form, but I would not bury it either. Send it to peer review, conditional on revision.","headline":"Useful data-centric FGL taxonomy, but the orthogonality claim is undercut by nested categories and inconsistent tables.","tokens_in":38054,"tokens_out":3163,"would_cite":false,"duration_ms":28706,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes the first data-centric taxonomy for federated graph learning, organizing studies by data characteristics and data utilization.","keywords":["federated graph learning","data-centric AI","graph neural networks","federated learning","taxonomy","survey","graph data heterogeneity","pre-trained large models"],"falsifier":"Opening Tables II and III and checking for papers that appear in more than one configuration or under an undefined phase name would settle the taxonomy's claim to be fine-grained and unambiguous: FedGNN and FeSoG appear in multiple rows, FedHG+ is listed under both reference markers [113] and [116], and the sequential dimension names a 'Global Training' phase rather than the defined 'Global Aggregation' phase.","tokens_in":36932,"feed_emoji":"🗂️","tokens_out":3220,"duration_ms":35045,"temperature":0.7,"pith_summary":"Federated graph learning (FGL) trains graph models across decentralized data holders without sharing raw data, and most prior surveys organize the field by model architecture or deployment scenario. This survey argues that a data-centric organization is needed, because nearly all FGL challenges are data-related: heterogeneity, missing structure, imbalance, efficiency, and privacy. It claims to be the first to reorganize FGL studies through a two-level taxonomy, where each level is defined by three orthogonal criteria. If the placement of the 79 surveyed works into the resulting tables is correct, researchers gain a practical reference for locating methods by data format, decentralization, visibility, training phase, and the data-centric challenge they address. The survey also extends the data-centric lens to FGL's integration with pre-trained large models and to emerging directions like continual graph learning and graph unlearning.","feed_headline":"First data-centric taxonomy sorts 79 federated graph learning works","feed_subtitle":"Two levels, three orthogonal criteria each, let researchers find methods by data format, decentralization, visibility, and challenge.","key_machinery":"The central object is a two-level taxonomy with orthogonal criteria. The Data Characteristics level captures the structural and distributional properties of data: graph format, decentralization format, and visibility format. The Data Utilization level captures how methods use data during training: the motivational challenge, the positional dimension, and the sequential dimension. The taxonomy carries the argument by being the instrument through which all 79 surveyed FGL works are categorized and compared, and it is the survey's claimed novelty over earlier scenario- and methodology-based taxonomies.","core_discovery":"The central claim is that every notable FGL study can be characterized by an orthogonal combination of three data-property criteria and three data-utilization criteria. The Data Characteristics level classifies works by graph format (homogeneous, heterogeneous, knowledge, or bipartite), decentralization format (horizontal or vertical), and client-side visibility (graph-oriented, subgraph-oriented, or ego-graph-oriented). The Data Utilization level classifies works by the data-centric challenge they target (quality, quantity, collaboration, efficiency, or privacy), the position of the innovation (client-side or server-side), and the training phase in which it acts (initialization, local training, global aggregation, or post-aggregation). The survey presents the resulting combinations as Tables II and III and claims these provide a fine-grained and unambiguous map of the field.","pith_inferences":["A natural test of the taxonomy is whether the three criteria in each level are truly orthogonal; the tables suggest some correlation, since ego-graph visibility appears almost exclusively under horizontal decentralization.","The taxonomy's client-server framing leaves peer-to-peer and Euclidean-oriented FGL on the sidelines, so a further level or a separate section would be needed to integrate those settings into the same map.","A concrete extension would be a searchable decision tree built from Tables II and III, letting practitioners filter by data format plus challenge, which the paper does not itself provide.","If the identified table inconsistencies are corrected, the survey could serve as a shared benchmark reference for coverage of FGL methods by data-centric attributes."],"forward_implications":["A researcher can locate FGL methods by data format, partition type, visibility, challenge, client- or server-side focus, and training phase in a single reference.","The data-centric reframing aligns FGL with the broader data-centric graph machine learning movement, making methods comparable by the data problem they solve rather than by backbone architecture.","The survey identifies FGL's integration with pre-trained large models as an early-stage area with two paradigms: PLM-enhanced FGL and FGL-enhanced PLMs.","Future directions are mapped to continual graph learning, graph unlearning, open-world graph learning, multimodal graph learning, and explainable aggregation, each positioned as an underexplored FGL extension.","The taxonomy serves as a starting point for practitioners who want to map a real-world decentralized data problem to a specific class of FGL methods."],"supporting_citations":[{"why":"Supplies the data-centric graph machine learning perspective that the survey extends to the federated setting.","marker":"[5]"},{"why":"Provides the Metis graph partitioning method used to simulate decentralized FGL data in the survey's background.","marker":"[7]"},{"why":"Provides the Louvain community-detection partitioning method used to create local subgraphs in FGL simulations.","marker":"[8]"},{"why":"Defines the GCN backbone used in the representative FGL training procedure presented by the survey.","marker":"[12]"},{"why":"Defines FedAvg, the default aggregation strategy in the survey's conceptual training pipeline.","marker":"[13]"},{"why":"Earlier FGL survey that first categorized approaches by graph data distribution, serving as the baseline this survey reorganizes.","marker":"[17]"},{"why":"Earlier FGL survey with a scenario-driven taxonomy of training architectures, compared against the new data-centric view in Table I.","marker":"[18]"},{"why":"Earlier FGL survey clarifying how federated learning and graph learning interact, compared against the new data-centric view in Table I.","marker":"[19]"}],"fun_headline_variants":["Data-centric taxonomy maps federated graph learning","New taxonomy organizes 79 FGL works by data usage","Two-level framework classifies federated graph learning","FGL survey built on orthogonal data criteria"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy is useful only if every surveyed paper can be assigned to exactly one configuration in Tables II and III, and if the cited references actually support those placements.","fun_headline_variants_meta":{"raw":{"variants":["Data-centric taxonomy maps federated graph learning","New taxonomy organizes 79 FGL works by data usage","Two-level framework classifies federated graph learning","FGL survey built on orthogonal data criteria"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00047,"raw_usage":{"total_tokens":2327,"prompt_tokens":917,"completion_tokens":1410,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1351}},"tokens_in":533,"tokens_out":1410,"duration_ms":11146,"temperature":1.0,"reasoning_tokens":1351,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:07:13.393154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Opening Tables II and III and checking for papers that appear in more than one configuration or under an undefined phase name would settle the taxonomy's claim to be fine-grained and unambiguous: FedGNN and FeSoG appear in multiple rows, FedHG+ is listed under both reference markers [113] and [116], and the sequential dimension names a 'Global Training' phase rather than the defined 'Global Aggregation' phase.","supporting_citations":[],"review_version":1}