{"id":"9c129767-e3a6-4dde-bfbd-05726560251e","arxiv_id":"2501.04015","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A descriptive network analysis of roughly 31,500 publications by Egyptian-affiliated authors finds a sparse, power-law citation network and identifies highly central papers and author collaborations.","lead":"This paper builds citation and co-authorship networks for papers by Egyptian-affiliated authors, using data scraped from Google Scholar and the Semantic Scholar API. It reports that the citation network is sparse and power-law distributed, and it flags possible citation manipulation in some connected components.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The name-based Semantic Scholar matching in §V-A Phase 2 is unvalidated author disambiguation; because every node and edge originates there, all network statistics and the citation-manipulation finding inherit this risk.","rationale":"I read the paper as a descriptive bibliometric case study whose value depends almost entirely on the correctness and representativeness of the AlGoNet dataset. The reader's weakest assumption correctly identifies the name-based API matching in §V-A Phase 2 as the point of maximum fragility. In good faith, I looked for a more fundamental or more specific flaw, but none is more load-bearing than this: if the author-resolution step is wrong, then the node set, edge set, degree distribution, centrality rankings, SCC analysis, and the conclusion about citation manipulation are all built on the wrong foundation. The paper does not release the dataset or code, and it contains internal count inconsistencies (31,508 vs 30,905 papers; 320,969 used as both citation count and node count), so the reader cannot independently verify even the most basic quantities. This is not a disagreement with the scientific consensus about power-law citation distributions; it is a correctness risk in the data-collection pipeline. I also note that the paper gives no indication that any validation was performed, despite the strong claim of a 'reliable and robust dataset' in the abstract. Because the issue is potentially repairable by releasing artifacts, disclosing author-ID resolution, and adding a validation study, the reader's CONDITIONAL verdict is appropriate. My read does not change that verdict, so I mark it UNCHANGED. The concrete test I propose would settle the concern by re-fetching a sampled subset using canonical author IDs and comparing the resulting network statistics to the reported ones.","tokens_in":9699,"tokens_out":3919,"duration_ms":34946,"concrete_test":"Reconstruct AlGoNet for a stratified random sample of 300 seed researchers, using Semantic Scholar author IDs recovered from each Google Scholar profile (or ORCID) instead of name-string API queries; manually verify at least 30 sampled author IDs against known publication lists. Then recompute the degree distribution and SCC count on the corrected sample and extrapolate. If more than approximately 10% of seed researchers map to the wrong author ID, or if the corrected degree distribution changes materially (for example, the power-law xmin or α shifts outside the reported values), the full-network claims are not robust. As a second check, audit whether the 62 non-trivial SCCs survive the corrected node set, since the citation-manipulation interpretation depends entirely on those components persisting after disambiguation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that AlGoNet (31,508 papers, 320,969 citations) is a reliable and robust dataset describing Egyptian-affiliated research. The load-bearing step is §V-A Phase 2: 'For each researcher, we sent an API request by name to retrieve data on their published papers.' The seed set is 13,027 names scraped from Google Scholar profiles at seven universities, and these names are matched to Semantic Scholar author records by name string only. No author-ID resolution, no homonym handling, no transliteration disambiguation, and no post-hoc validation is reported. If a seed researcher's name resolves to the wrong Semantic Scholar author, every paper fetched for that researcher is off-target; conversely, if the correct author publishes under variants, papers are silently omitted. Because the network is then constructed from the union of per-author paper sets plus their references, errors at this stage propagate to node membership, every degree and centrality score, the power-law fit (α=1.725958, xmin=158), the SCC count, and the pre-publication cross-citation conclusion in §V-C7 Case 2. The paper also contains internal count inconsistencies (31,508 papers vs 30,905 in the authorship analysis; 320,969 treated both as citations and as 'the number of nodes'), which makes the unvalidated matching impossible to audit without releasing the dataset and code. Thus the minimal condition for the central claim—that the measured network actually represents Egyptian-affiliated scholarship—is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper scrapes Google Scholar profiles at seven Egyptian universities to obtain 13,027 researcher names, then queries the Semantic Scholar API by name for each researcher to retrieve their publications and reference lists, producing a cleaned dataset called AlGoNet that the authors report contains 31,508 papers and 320,969 citations. From these data they build citation and co-authorship networks and report descriptive statistics: temporal distribution, degree distribution with a claimed power-law fit (α = 1.725958, xmin = 158), authorship patterns, clustering coefficient, density, centrality measures, strongly connected components, co-authorship edge weights, and connected components. The authors conclude that the dataset is reliable and robust, that a small number of highly cited papers dominate the network, and that 62 non-singleton strongly connected components indicate citation manipulation through pre-publication cross-citation, in addition to legitimate journal supplement structures.","tokens_in":9983,"tokens_out":6528,"duration_ms":50529,"significance":"If the AlGoNet dataset were validated and publicly released, the paper would provide a useful descriptive baseline for Egyptian-affiliated scholarly output and a concrete case study in citation-network analysis. The paper uses standard tools (NetworkX, the powerlaw library) and reports several specific quantitative findings, which is a positive feature. However, the central claim of a 'reliable and robust dataset' is not currently supported because the name-based author matching is unvalidated, the reported dataset counts are internally inconsistent, the power-law fit lacks statistical confidence measures, and the citation-manipulation interpretation is not backed by systematic evidence. The paper does not state any plan to release the dataset or code, further limiting reproducibility. The contribution is therefore mostly descriptive and would need substantial methodological strengthening to rise to the standard of a rigorous network-analysis study.","major_comments":[{"comment":"The load-bearing step of dataset construction is unvalidated author disambiguation. The paper states, 'For each researcher, we sent an API request by name to retrieve data on their published papers' (§V-A, Phase 2), and the seed set consists of 13,027 names scraped from Google Scholar profiles (§V-A, Phase 1). No author-ID resolution, homonym handling, transliteration disambiguation, or post-hoc validation is reported. Because all nodes and edges in both networks are derived from the union of per-author paper sets, any false match or missed name variant propagates into every reported statistic, including the power-law fit (§V-C2), the SCC counts (§V-C7), and the co-authorship centrality analysis (§V-D). The paper should either release the dataset with matched author IDs and a validation protocol, or carry out and report a sampling-based precision/recall check of the name-to-author mapping.","section":"Section V-A, Phase 2"},{"comment":"The reported dataset counts are internally inconsistent and conflate citations with nodes. The paper reports 31,508 papers and 320,969 citations in §V-A; §V-C7 then states that 'the total number of SCCs was expected to be 320,969, the number of nodes in the graph,' treating the citation count as the node count. Additionally, §V-C3's authorship-pattern analysis is based on 30,905 papers without explaining the discrepancy from 31,508. These inconsistencies must be reconciled before the dataset can be called reliable and robust, and they prevent an independent audit of the analysis.","section":"Section V-A and V-C7"},{"comment":"The power-law claim is not supported by the reported statistics. The paper lists α=1.7259583924156112 and xmin=158 but gives no uncertainty on α, no goodness-of-fit measure such as the Kolmogorov-Smirnov statistic or its p-value, and no comparison against alternative heavy-tailed distributions (e.g., log-normal, stretched exponential). The paper also does not specify whether the degree in the distribution is in-degree, out-degree, or total degree. Without these, the claim that 'the degree distribution appears to follow a power law' is an unsupported assertion rather than an empirical finding.","section":"Section V-C2"},{"comment":"The conclusion of citation manipulation through pre-publication cross-citation is not established. The paper identifies 62 strongly connected components of size greater than one and attributes them to two causes, with Case 2 asserted as 'adopted behavior by some authors' that is 'unethical.' However, the only concrete example given is Case 1 (journal supplements), and no evidence is provided for the remaining 61 components: common-authorship analysis, publication-date distributions, or a manual inspection of even a sample. SCCs of size greater than one can also arise from data artifacts such as duplicated paper records, citation errors in publishers' metadata, or symmetric references in corrigenda and replies. The authors should either systematically categorize all 62 components and report the distribution of causes, or revise the claim to note that the observed SCCs are consistent with multiple data-generating mechanisms.","section":"Section V-C7, Case 2"}],"minor_comments":[{"comment":"The sentence 'The power-law distribution can be expressed as P (x) = Cx −α' should use proper notation and specify that x is the degree; also 'x_min' should be typeset as a subscript.","section":"Section V-C2"},{"comment":"The paper would benefit from stating the definition of degree (in-degree, out-degree, or total) in Section V-C2, since Section V-C2 says 'degree of a node represents the number of citations it received' (in-degree) but Section V-C6 uses degree centrality that may be total degree.","section":"Section V-C2 and V-C6"},{"comment":"The statement 'the total number of SCCs was expected to be 320,969' should be phrased as 'the number of SCCs would be 320,969 if every node were a singleton SCC'; as written, it could be read as claiming that the number of SCCs equals the number of nodes regardless of non-singleton components.","section":"Section V-C7"},{"comment":"The sentence 'the highest degree obtained (=0.0015)' likely should say 'highest degree centrality'; the value 0.0015 does not match the reported degree of 682 for the full network, so the denominator used for normalization should be clarified.","section":"Section V-C6"},{"comment":"The paper does not state whether the AlGoNet dataset or the analysis code will be made available; a data-availability statement would substantially help reproducibility.","section":"General"},{"comment":"The literature review is very brief and does not discuss existing work on author disambiguation or citation-network cleaning, both of which are directly relevant to the methodology.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"This paper appears to be a well-intentioned applied project, and the AlGoNet dataset could be a useful resource if it is properly validated and released. The main concerns are all addressable in a revision: validate the author-matching step, reconcile the dataset counts, add statistical support for the power-law fit, and temper or support the SCC interpretation. I would encourage the editor to invite a revised version with those changes, since the central idea is not fundamentally unsound but the current evidence is insufficient to support the paper's claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: this is a bibliometric case study of Egyptian-affiliated research, and its main contribution is the AlGoNet dataset itself—a new list of 31k papers and 320k citations assembled from Google Scholar profiles at seven Egyptian universities matched to Semantic Scholar records. No code or data is released, and the pipeline is a standard scrape-API-NetworkX workflow, so the novelty is purely the population studied. If the dataset is reliable, it's a useful descriptive benchmark for policymakers and bibliometricians.\n\nThe paper does some things well. It's clearly written, the network statistics are standard and mostly sensible, and the authors are upfront about the data cleaning steps they did take. The power-law fit (alpha=1.73, xmin=158) is a plausible summary of the degree distribution, and the co-authorship analysis correctly notes that the giant component of 73k authors reflects massive international collaborations like global surgery trials.\n\nThe soft spots are load-bearing. First, the name-based matching of 13,027 Google Scholar names to Semantic Scholar author records is unvalidated. No homonym handling, no ID resolution, no post-hoc check. If the wrong author is matched, every degree, centrality score, and SCC claim inherits that error. That is not a minor issue; it's the foundation. Second, the numbers don't reconcile: 31,508 papers and 320,969 citations in one place, 30,905 papers in the authorship analysis, and then 320,969 treated as 'the number of nodes' in the SCC section. That conflation of citations with nodes is a red flag that the graph construction isn't audited. Third, the citation-manipulation finding in Section V-C7 Case 2 is speculative. The 62 non-singleton SCCs are 'unreasonable' only under the assumption that mutual citation can't happen naturally, and no systematic evidence links them to pre-publication cross-citation. The supplement explanation is fine, but the manipulation claim needs stronger support.\n\nThere's no independent test here—the power law is fit to the same data it describes, so treat it as a summary, not a prediction.\n\nIs this paper serious? Yes, the authors are doing real empirical work, but the unvalidated matching and numerical inconsistencies prevent me from trusting the conclusions as posted. It deserves a serious referee because the dataset could be valuable and the flaws are fixable: release the artifacts, fix the node counts, validate the matching, and soften the manipulation claim.\n\nFor you: worth a look if you care about bibliometrics or Egyptian research policy, but don't cite it until the numbers are repaired.","headline":"A genuinely new descriptive dataset for Egyptian scholarly output, but the unvalidated name-based author matching and internal number inconsistencies make the posted results unreliable until fixed.","tokens_in":10531,"tokens_out":2936,"would_cite":false,"duration_ms":25415,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Citation network of Egyptian papers reveals manipulation pattern","keywords":["graph analysis","network analysis","citation network","co-authorship network","authorship pattern","network metrics","Networkx","Egyptian authors"],"falsifier":"Take a random sample of about 500 of the 13,027 name-matched researcher records and manually verify, against the Google Scholar profile and the Semantic Scholar record, that the papers returned actually belong to the same person (same affiliation, co-authors, and publication history); if the mismatch rate is non-trivial, the AlGoNet node set is unreliable and the paper's statistics do not describe Egyptian-affiliated research as claimed.","tokens_in":9471,"feed_emoji":"🕸","tokens_out":6738,"duration_ms":53101,"temperature":0.7,"pith_summary":"Drawing on open bibliographic data, this paper attempts to establish a reliable network-level picture of Egyptian-affiliated research: a citation graph of 31,508 papers and 320,969 citations, and a co-authorship graph of their authors. The authors argue the cleaned dataset, called AlGoNet, is robust enough to support conclusions about influence, collaboration, and behavior. Their headline findings are a power-law citation degree distribution with exponent $\\alpha = 1.725958$, and the detection of 62 non-trivial strongly connected components, which they attribute to journal-supplement citation artifacts and to deliberate pre-publication cross-citation that inflates authors' counts. If the dataset is sound, these results give researchers and policymakers concrete evidence about where Egyptian research impact concentrates and how it can be distorted.","feed_headline":"Citation network of Egyptian papers reveals manipulation pattern","feed_subtitle":"AlGoNet maps 31,508 papers and 320,969 citations, exposing power-law impact and pre-publication cross-citing.","key_machinery":"The load-bearing object is the AlGoNet dataset, built in two phases: web-scraping Google Scholar profiles at seven Egyptian universities to collect 13,027 researcher names, then sending name-based queries to the Semantic Scholar API to retrieve each researcher's papers, references, citation counts, and co-author IDs. The analytical machinery is standard graph theory implemented in NetworkX — degree and eigenvector centrality, clustering coefficient, graph density, and strongly connected components (sets of nodes in which every node can reach every other via directed paths), computed with a standard SCC algorithm. The SCC decomposition is what makes the paper's distinctive claim possible: in a citation graph where edges should point only from newer to older papers, the 62 cycles of size greater than one become a red flag, and manual inspection turns them into evidence about citation manipulation.","core_discovery":"On the paper's own terms, the central discovery is that the AlGoNet dataset, assembled from 13,027 researcher names scraped from Google Scholar and matched by name into Semantic Scholar, yields a citation network whose structure is both typical and revealing. The degree distribution follows a power law with scaling exponent $\\alpha = 1.725958$ and minimum degree $x_{\\min} = 158$, consistent with the 'rich get richer' dynamics seen in many scientific fields. The network is sparse (density $D = 3.55 \\times 10^{-6}$) and weakly clustered ($C = 0.019$), and about two-thirds of papers have degree near zero. The most consequential finding is 62 strongly connected components of size greater than one in a graph where papers should only point backward in time; the authors inspect these cycles and find two causes, supplementary-issue articles citing one another and intentional pre-publication cross-citation among papers sharing authors and publication months. This last pattern, the authors argue, is an adopted behavior aimed at raising citation counts and author ranks.","pith_inferences":["Because the researcher-to-paper matching is name-based and unvalidated, the reliability of every network statistic rests on that matching; a small validation study could test it.","The same SCC-detection procedure could be turned into an automated early-warning indicator for citation manipulation in other national or institutional bibliographic datasets.","The power-law exponent and network metrics could be compared across countries to see whether the Egyptian network's structure is typical or distinctive, but the paper does not make that comparison.","The data collection method could be extended to universities beyond the seven selected to reduce coverage bias and improve generalizability."],"forward_implications":["The AlGoNet dataset can serve as a foundation for future studies of Egyptian research impact, collaboration, and policy decisions.","The power-law degree distribution implies a small number of Egyptian-authored papers account for most citations, so evaluations and policies targeting high-impact work should focus on that tail.","The 62 non-trivial strongly connected components are evidence that some authors inflate citation counts by cross-citing not-yet-published papers, a practice that can distort rank-based metrics.","The co-authorship network's single giant connected component of 73,690 authors indicates that one large, globally connected collaboration cluster dominates Egyptian research networks alongside many small isolated groups.","Multi-authored papers are the majority, confirming a strong collaborative trend in Egyptian research output."],"supporting_citations":[{"why":"Supplies the API method used to retrieve papers, references, citation counts, and co-author IDs.","marker":"[12]"},{"why":"Provides the empirical baseline that citation degree distributions follow a power law, which the paper's degree fit is compared against.","marker":"[13]"},{"why":"Supplies the algorithm for finding strongly connected components used to detect the 62 cycles.","marker":"[18]"},{"why":"Motivates the centrality measures applied to the co-authorship network.","marker":"[10]"},{"why":"Presents the metadata extraction approach the paper's web-scraping-plus-API method builds on.","marker":"[11]"}],"fun_headline_variants":["Egyptian citation network hides 62 cross-citing cycles","Pre-publication cross-citing boosts Egyptian author ranks","Citation analysis exposes manipulation in Egyptian papers","Power-law and cycles: Egyptian citation network revealed","AlGoNet finds suspicious pre-publication citations in Egypt"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis rests on the assumption that the 13,027 researcher names scraped from Google Scholar were matched, by name alone and without validation, to the correct author records in Semantic Scholar; if a substantial share of those matches point to the wrong researchers, every network statistic and every conclusion about citation manipulation is built on the wrong nodes.","fun_headline_variants_meta":{"raw":{"variants":["Egyptian citation network hides 62 cross-citing cycles","Pre-publication cross-citing boosts Egyptian author ranks","Citation analysis exposes manipulation in Egyptian papers","Power-law and cycles: Egyptian citation network revealed","AlGoNet finds suspicious pre-publication citations in Egypt"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1311,"prompt_tokens":925,"completion_tokens":386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":541,"tokens_out":386,"duration_ms":3656,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:24:29.476314+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of about 500 of the 13,027 name-matched researcher records and manually verify, against the Google Scholar profile and the Semantic Scholar record, that the papers returned actually belong to the same person (same affiliation, co-authors, and publication history); if the mismatch rate is non-trivial, the AlGoNet node set is unreliable and the paper's statistics do not describe Egyptian-affiliated research as claimed.","supporting_citations":[{"cited_title":"Semantic scholar documentation,","cited_arxiv_id":null,"evidence_quote":"Supplies the API method used to retrieve papers, references, citation counts, and co-author IDs."},{"cited_title":"Case study–centrality measure analysis on co-authorship network,","cited_arxiv_id":null,"evidence_quote":"Motivates the centrality measures applied to the co-authorship network."},{"cited_title":"Rule Based Metadata Extraction Framework from Academic Articles","cited_arxiv_id":"1807.09009","evidence_quote":"Presents the metadata extraction approach the paper's web-scraping-plus-API method builds on."}],"review_version":1}