{"id":"903c24fe-0ea9-41ed-91da-958f8d8a0240","arxiv_id":"1908.08478","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The authors construct and analyze a 52,406-author co-authorship network of network scientists, revealing a growing community with a giant component covering 62.8% of researchers.","lead":"This paper maps the last 20 years of network science by building a co-authorship network of 52,406 researchers whose papers cite one of three landmark network papers. It shows how the field grew from a physics-dominated group into a more connected, interdisciplinary community.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Name disambiguation is under-validated: identical full names, especially for the large Chinese subcommunity, may be merged, potentially inflating the giant-component growth that underpins the main connectivity claim.","rationale":"The Reader's weakest assumption is the citation-based definition of a network scientist. The paper itself acknowledges that this is arbitrary and states it is a proxy; the title 'as seen through' makes the lens explicit. That assumption affects external validity but does not threaten the internal validity of the network constructed from the chosen papers. A more load-bearing problem is author name disambiguation, which the paper flags in Section III but then asserts without evidence to be negligible. The dataset has a large Chinese subpopulation, where names are short and ambiguous; merging distinct authors is likely to create false co-authorship edges. These false edges can artificially connect components, making the giant component larger and more connected over time, which is exactly the paper's central narrative. The paper provides no sensitivity analysis for this. The right verdict remains conditional, but the condition should include a disambiguation robustness check rather than only a re-evaluation of the seed-paper choice. The authors' own limitation statement is the strongest clue that this is the area to probe, and the paper's cited justification from earlier smaller studies does not transfer to a dataset where the largest community is majority Chinese.","tokens_in":8166,"tokens_out":10258,"duration_ms":108585,"concrete_test":"Reconstruct the co-authorship network using WoS author identifiers (ORCID/ResearcherID) for the subset of records that contain them, or apply a conservative disambiguation rule that merges identical full names only when the affiliation country/address also matches. Compare the largest-component fraction (reported as 62.8%), the temporal ratio curve of Fig. 10, and the top centrality rankings to the original. If the giant component fraction changes by more than a percentage point or the growth trend flattens, the connectivity conclusion is not robust to author disambiguation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central connectivity result depends on correctly identifying authors, but Section III acknowledges that identical full names are not distinguished, calls the issue 'mainly relevant for Asian authors,' and dismisses it as negligible by citing earlier studies. This dataset is not comparable: the largest community is 54% Chinese (Table III), and Chinese names in WoS are typically short pinyin strings with high collision rates. Merging any two distinct authors with the same full name introduces spurious edges between their disjoint co-author groups, inflating the giant component, raising clustering, and distorting the temporal connectivity trend in Fig. 10. The authors do not quantify the collision rate or test sensitivity. Because the paper's headline claim is that the network science community is 'diverse but not divided' and the giant component has grown to 62.8%, a few percent of false merges could be decisive. The burden is on the authors to show the effect is actually negligible in this dataset.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper constructs and analyzes the co-authorship network of 52,406 researchers who have at least one paper citing at least one of three seminal network science papers (Watts & Strogatz 1998, Barabási & Albert 1999, Girvan & Newman 2002). The authors characterize the papers themselves (research areas, journals, keywords, countries), then study the topology and dynamics of the co-authorship network: degree distribution, clustering, centralities, community structure, and the growth of the largest connected component over time. They find that the largest component has grown to 62.8% of the network, interpret this as evidence of a 'diverse but not divided' community, and report a correlation between centrality and citation counts. The anonymized data are made available in a GitHub repository.","tokens_in":8356,"tokens_out":4825,"duration_ms":49865,"significance":"If the data-construction choices hold up, this is a useful descriptive and reference contribution: it provides a large, openly available co-authorship dataset for a field-defining citation-based population, and it quantifies the community's evolution over 20 years. The authors are appropriately careful in places—they explicitly acknowledge that the citation-based definition of 'network science' is arbitrary, they share their data, and they rely on standard graph metrics. The central claims about growing connectivity and community diversity are falsifiable and important for the science-of-science literature. However, the validity of the headline connectivity results depends on author name disambiguation, and the paper's treatment of that issue is inadequate for this specific dataset.","major_comments":[{"comment":"The claim that the error from conflating distinct authors with identical full names is 'negligible, as also pointed out by Newman [22] and by Barabasi et al. [20]' is not supported for this dataset. Table III shows the largest community is 54% Chinese, and Web of Science full-name strings for Chinese authors are typically short pinyin with high collision rates. The cited prior works examined smaller or different datasets, so they cannot justify this conclusion here. Since the growth of the giant component (Fig. 10) is the central evidence for the 'diverse but not divided' claim, the authors should quantify the collision rate in their dataset or demonstrate sensitivity of the main results to name-merging choices (e.g., by re-running the analysis with conservative splitting heuristics or by validating against ORCID data). Without such a check, the headline connectivity numbers rest on an unvalidated assumption.","section":"Section III, Data collection and preparation"},{"comment":"The reported global clustering coefficient of 0.98 is surprisingly high and is presented without explanation or a clear definition. The paper also reports an average local clustering coefficient of 0.77 and an average degree of 12.56; a transitivity of 0.98 in a network with that density is extreme and not self-evidently plausible. The authors should state whether 'global clustering coefficient' means the fraction of closed triples among all triples, the average local clustering, or some other quantity, and they should verify the computation. If the value is correct, a brief discussion of why co-authorship networks achieve such high transitivity (e.g., due to large-author papers) would help; if it is an artifact of the definition, the text currently overstates the clustering.","section":"Section V, Analysis of the co-authorship network"},{"comment":"The authors acknowledge that the definition of a network science paper (citing at least one of three selected papers) is arbitrary, but they do not test how robust their conclusions are to this choice. In particular, the growth of the giant component, the high clustering, and the centrality–citation correlation could in principle depend on the set of root papers or the citation-threshold. I ask for at least a limited sensitivity analysis—for example, varying the set of seminal papers (e.g., dropping one of the three) or using a stricter citation requirement—to show that the main findings are not artifacts of the specific definition. This would materially strengthen the paper's claim to describe 'the network science community' rather than merely the selected citing population.","section":"Section II and Section V"}],"minor_comments":[{"comment":"The column headers for Table II are ambiguous: the text lists 'betweenness' and 'harmonic' centralities, but the table layout is unclear about which column corresponds to which measure. Please add explicit headers or break the table into separate columns with clear labels.","section":"Table II"},{"comment":"The text states there is 'a strong correlation' between centrality and citation count, but no correlation coefficient or statistical test is reported. Please provide the Pearson and/or Spearman correlation values, or otherwise quantify the strength of the association.","section":"Section V, Fig. 9"},{"comment":"The sentence 'the error introduced by this problem is negligible, as also pointed out by Newman [22] and by Barab'asi et al. [20]' should be rephrased: the cited works do not establish that the error is negligible in a dataset with the demographic composition of the present one; at minimum, the statement should be presented as an assumption rather than a conclusion.","section":"Section III"},{"comment":"There are minor typographical errors: 'world clouds' should be 'word clouds' (Section IV), 'measurues' should be 'measures' (Fig. 9 caption), and 'Phyics' appears in the Table III legend. Also, the phrase 'the authors of this article emerge as a maximal clique' (Section V) refers to the 388-author consortium paper; consider clarifying that the consortium members form a clique, not the six listed authors only.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a descriptive, tribute-style study with transparent data sharing, but the name-disambiguation issue is a genuine correctness risk for the central connectivity claims. The authors should be asked to address it with a sensitivity analysis or a more careful justification before the paper can be recommended for acceptance. The clustering coefficient anomaly should also be resolved. The paper seems otherwise well within the scope of the conference; the issues are fixable with additional analysis rather than being fundamental. I would not reject, but the current version is not yet fully convincing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is a careful, descriptive map of the network-science community built from citations to three seminal papers. The dataset is large and the authors make the constructed network publicly available. The main caveat is that the authors never validate their name-disambiguation step, and that is not a minor omission given the Chinese-name collision risk in this dataset.\n\nWhat is new: prior co-authorship analyses (Newman et al.) covered about 1,600 authors; this one covers 52,406. The paper documents the field's growth from physics to computer science, the increasing Chinese share, community composition, centrality rankings, and the claim that the community is 'diverse but not divided' as evidenced by the growing giant component. It is honest about the arbitrariness of the citation-based definition (Section II), and the GitHub data is a real plus.\n\nSoft spots, in proportion: The stress-test note is on target. The paper says identical full names are not distinguished and calls the error negligible, citing Newman and Barabasi et al. But those studies used different bibliographic environments. Here, the largest community is 54% Chinese, and Chinese full names in Web of Science are typically short pinyin strings with high collision rates. Merging distinct authors adds spurious edges, which can inflate the giant component, the clustering coefficient, and the temporal connectivity trend in Fig. 10. The authors need to show a robustness check (e.g., using author identifiers, or a simulated collision test) or soften the 'diverse but not divided' claim.\n\nThe global clustering coefficient of 0.98 is also surprisingly high and unexplained. Co-authorship networks do have high clustering from multi-author cliques, but 0.98 means almost every connected triple is a triangle, which is inconsistent with the reported average local clustering of 0.77. I would want the definition and a sanity check.\n\nThe claimed correlation between centrality and citations is supported by a scatter plot but no correlation coefficient; they should report the actual number.\n\nWho this is for: science-of-science researchers, scientometricians, and anyone needing a descriptive benchmark of the network-science community. It is not a methodological leap, but it is a well-documented empirical contribution.\n\nI would send this to peer review. The editor should ask for the robustness checks above. If the authors do the sensitivity analysis, the paper becomes a solid, citable reference.","headline":"A solid, well-documented descriptive map of the network-science community that needs a name-disambiguation robustness check before the connectivity claims are fully convincing.","tokens_in":8828,"tokens_out":2709,"would_cite":true,"duration_ms":25471,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["05C82","91D30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A co-authorship network built from papers citing three landmark studies shows network science becoming a single, connected community over two decades.","keywords":["co-authorship network","network science","scientometrics","science of science","community structure","centrality","citation analysis","complex networks"],"falsifier":"Recompute the giant component ratio and the centrality–citation correlation using a different definition of network science, for instance papers published in network-science-specific journals or papers citing a broader set of landmark works. If the giant component drops far below 62.8% or the centrality–citation correlation weakens substantially, the paper's portrait of a single, connected community is an artifact of the citation proxy.","tokens_in":7991,"feed_emoji":"🕸️","tokens_out":4414,"duration_ms":39849,"temperature":0.7,"pith_summary":"The paper attempts to show that the identity and evolution of the network science community can be read from the co-authorship graph of researchers who cite at least one of three founding papers: Watts–Strogatz, Barabási–Albert, and Girvan–Newman. Analyzing 29,528 such papers and 52,406 authors, it finds that this community has grown steadily more connected, with the largest connected component now containing 62.8% of all authors. It also reports a strong correlation between an author's centrality in the co-authorship network and the citation count of their network science papers. The value of the claim is that it turns an arbitrary citation-based proxy into a quantitative portrait of how an interdisciplinary field forms, merges, and gains cohesion.","feed_headline":"Two decades of network science: one community, 62.8 percent connected","feed_subtitle":"A map of 52,406 authors shows citations and collaboration centrality moving together as the field matures.","key_machinery":"The load-bearing construction is the co-authorship network of network scientists. A node is any author of a paper citing at least one of three seminal works, and an edge joins two authors who co-authored at least one such citing paper. This single definition supplies the corpus, the vertex set, and the edge set, so the entire analysis depends on it. On top of this network the paper uses Clauset–Newman–Moore greedy modularity maximization for communities, betweenness and harmonic centrality for author importance, and a country-level collaboration graph for international patterns.","core_discovery":"The central discovery is a structural portrait: network science, as delimited by citing one of three milestone papers, is not a fragmented collection of sub-disciplines but a single growing component. The largest connected component of the co-authorship network comprises 32,904 of 52,406 authors (62.8%), a fraction that increased over time. Community detection reveals ten large communities, with the largest (14,136 authors) dominated by Chinese physicists, and smaller communities that are more homogeneous in discipline and country. Centrality in the co-authorship network correlates strongly with citation counts, so position in the collaboration graph tracks scientific impact.","pith_inferences":["A testable extension would be to treat the three landmark papers as a fixed seed set and vary the citation distance (direct citers only versus two-hop citers) to see whether the giant component ratio is stable; if it collapses, the boundary of the field is much fuzzier than the paper suggests.","The same construction could be applied to other fields with identifiable founding papers, turning 'a community that cites X' into a general tool for mapping disciplinary emergence.","The centrality–citation correlation could be compared against a null model of random rewiring, which would separate genuine structural advantage from the mere fact that highly cited authors appear in many papers and therefore have high degree."],"forward_implications":["If the proxy is faithful, the field's cohesion has been increasing over time, and the 62.8% giant component ratio is the quantitative signature of that cohesion.","Centrality in the co-authorship network can serve as a proxy for scientific impact where citation data are unavailable or unreliable.","The community structure implies that network science is held together by a few interdisciplinary bridges rather than by uniformly dense collaboration, since the largest community is a Chinese physics-heavy cluster.","The spatiotemporal data imply that China and the US dominate production, and that international collaborations are concentrated among European countries and the US."],"supporting_citations":[{"why":"Barabasi and Albert's scale-free networks paper is one of the three seminal papers whose citing works define the network science corpus.","marker":"[2]"},{"why":"Watts and Strogatz's small-world networks paper is one of the three seed papers used to identify network scientists.","marker":"[3]"},{"why":"Girvan and Newman's community structure paper is the third seed paper defining the network science corpus.","marker":"[4]"},{"why":"Clauset, Newman, and Moore's greedy modularity maximization algorithm is used to detect the largest communities in the co-authorship network.","marker":"[28]"}],"fun_headline_variants":["62.8% of network scientists in one giant component","Two decades, 52k authors, one big component","Network scientists form a single 62.8% giant component","Co-authorship map: network science is one connected community","62.8% connected: one giant component in network science"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis assumes that 'network science paper' can be defined as any paper citing at least one of the three selected milestone papers, even though the authors admit this is arbitrary and will both miss real network science and include unrelated citing papers.","fun_headline_variants_meta":{"raw":{"variants":["62.8% of network scientists in one giant component","Two decades, 52k authors, one big component","Network scientists form a single 62.8% giant component","Co-authorship map: network science is one connected community","62.8% connected: one giant component in network science"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000785,"raw_usage":{"total_tokens":3442,"prompt_tokens":904,"completion_tokens":2538,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":2454}},"tokens_in":520,"tokens_out":2538,"duration_ms":19961,"temperature":1.0,"reasoning_tokens":2454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:38:03.181387+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the giant component ratio and the centrality–citation correlation using a different definition of network science, for instance papers published in network-science-specific journals or papers citing a broader set of landmark works. If the giant component drops far below 62.8% or the centrality–citation correlation weakens substantially, the paper's portrait of a single, connected community is an artifact of the citation proxy.","supporting_citations":[{"cited_title":"Collective dynamics of small-world networks,","cited_arxiv_id":null,"evidence_quote":"Watts and Strogatz's small-world networks paper is one of the three seed papers used to identify network scientists."},{"cited_title":"Community structure in social and biological networks,","cited_arxiv_id":null,"evidence_quote":"Girvan and Newman's community structure paper is the third seed paper defining the network science corpus."},{"cited_title":"Finding community structure in very large networks,","cited_arxiv_id":null,"evidence_quote":"Clauset, Newman, and Moore's greedy modularity maximization algorithm is used to detect the largest communities in the co-authorship network."}],"review_version":1}