{"id":"d0b8d667-d6ac-48a1-a6b6-5a9ef7ee0557","arxiv_id":"2608.10774","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A reproducible OpenAlex graph pipeline shows that broad disciplinary scope, not citation insularity, is the robust structural signal of problematic venues, and that PageRank-based journal prestige is about ten times more resistant to injected citation cartels than count-based indicators.","lead":"This paper presents an open-source graph tool that turns public OpenAlex publication data into a seven-type network and screens it for unusual co-authorship, citation, and venue patterns. A controlled test on delisted journals shows that broad disciplinary scope, not citation insularity, is the robust warning sign, and that PageRank-based journal prestige is far harder to inflate with fake citations than simple counts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Field-entropy AUC 0.70 may be a field-composition artifact; the reported robust signal is not yet shown to discriminate within matched fields.","rationale":"The reader identified the same weakest assumption: matching only by size and indexing status, not by field or age (Section 7.5). My analysis agrees this is the most load-bearing concern. The field-entropy detector is the only robust signal surviving the matched design; if its AUC is inflated by field composition, the paper loses its quantitative validation of a screening signal for problematic venues. The PageRank gaming-resistance result (7.4) is a secondary simulation with unreported details, and the insularity null result is robust, but neither carries the same weight as the AUC 0.70 field-entropy claim. The proposed test is feasible: the authors already have venue-level field data and can rematch without new collection. The paper is otherwise solid in execution, honest about limitations, and its negative result is useful. Therefore the verdict remains CONDITIONAL: the central claim should be stated with this caveat or strengthened by field-matched rematching; the abstract's wording already risks overclaiming, and this concern makes the condition concrete.","tokens_in":15224,"tokens_out":1231,"duration_ms":10781,"concrete_test":"Rematch the Study C controls by field (e.g., same OpenAlex main field or subfield) and by age range (e.g., same start year of indexing or same publication-year distribution), keeping the size matching, then recompute the field-entropy AUC on the integrity subset and the full positive set. If the AUC drops to ~0.5 (or below 0.6), the paper's strongest surviving empirical claim is undone; if it remains ~0.7, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim that \"after matching the only robust surviving signal is the breadth of disciplinary scope (AUC 0.70)\" depends on the controls being comparable to delisted venues except for integrity status. The authors' Section 7.5 states the matching is \"only by size and indexing status (not by field or age)\". Field entropy is computed over OpenAlex fields, so if delisted venues are drawn disproportionately from broad, multi-field disciplines (e.g., multidisciplinary mega-journals) and size-matched still-indexed controls from narrower single-field journals, the AUC 0.70 could reflect legitimate disciplinary breadth rather than an integrity-linked \"scope creep\". This is not a contradiction, but it is a direct threat to the load-bearing conclusion, and the paper's own limitation statement flags it. The claimed \"scope creep\" signature (median entropy 2.33 vs 1.82) is exactly what one would expect from a field-composition imbalance; without field-stratified or field- and age-matched controls, the surviving signal is ambiguous.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a heterogeneous graph model of the academic publishing network over OpenAlex, with node/edge types, projections (citation and co-authorship), structural metrics, community detection, and three screening detectors. It demonstrates the methodology on an institutional corpus (VSB-TUO 2020–2025) and a worldwide LLM corpus, then reports a controlled venue-centric study (Study C) using journals delisted by Scopus/DOAJ as ground truth. The central empirical claims are: after size-matching controls, the only robust surviving discriminator of delisted venues is the breadth of disciplinary scope (field entropy, AUC 0.70, 95% CI [0.64;0.75]); and an open graph-based prestige measure (PageRank over the journal citation network) tracks a JIF proxy (Spearman 0.49) while being roughly an order of magnitude more resistant to a synthetic citation-cartel injection than count-based indicators. The paper releases the method as the open-source apnet library and emphasizes screening-with-evidence rather than binary classification.","tokens_in":15340,"tokens_out":4627,"duration_ms":49412,"significance":"If the claims hold, the paper makes two useful contributions: it provides a sober, falsifiable validation showing that a naïve case-control design produces a spurious 'prominence' signal, and it indicates that field-entropy breadth is the one structural feature that survives size matching in the full delisted set. It also demonstrates reproducible, commodity-hardware analysis of an open dataset, which is a concrete strength; the code release and explicit limitations statements support reproducibility. The PageRank gaming experiment is a testable, transparent robustness check. However, the significance of the empirical conclusions is conditional: the field-entropy result is threatened by the lack of field/age matching, and the gaming-resistance claim is based on a single arbitrarily parameterized attack model.","major_comments":[{"comment":"The claim that 'the only robust surviving signal is the breadth of disciplinary scope (AUC 0.70)' is not yet established, because the controls are matched only by size and indexing status, not by field or age, as the manuscript itself acknowledges in §7.5. Field entropy is computed over OpenAlex fields, so if delisted venues are drawn disproportionately from broad multidisciplinary venues while size-matched still-indexed controls are field-specialized, the median entropy difference (2.33 vs. 1.82) could reflect legitimate differences in disciplinary breadth rather than an integrity-linked 'scope creep.' This is a direct threat to the load-bearing empirical claim. The authors should provide a field-stratified or field/age-matched comparison, or a regression that controls for OpenAlex field composition, and report the AUC within relevant field categories; without such an analysis, the surviving-signal conclusion remains ambiguous.","section":"§7.3, §7.5"},{"comment":"The statement that PageRank is 'an order of magnitude more resistant to citation gaming than count-based indicators' is based on a single synthetic injection scenario (1,250 fictitious citations into one mid-prestige venue), giving inflation factors 84x vs. 8.5x. No confidence intervals or sensitivity analysis are reported, and the immediately subsequent Sybil attack (moving the venue from rank 172 to 12) shows that the resistance is highly conditional on the attack model. The claim should be qualified to the specific injection scenario, and the comparison should be repeated over a range of cartel sizes, numbers of fictitious citing venues, and target venue ranks, with identical attack budgets for both metrics.","section":"§7.4"},{"comment":"The size-matching protocol for the controls is underspecified. The text states that controls are 'size-comparable' and that up to 200 works per venue are sampled, but it does not state the matching variable(s), the matching algorithm (nearest-neighbor, caliper, exact), the matching ratio (one control per positive or several), or whether matching was performed with replacement. Without this information, the AUC values in Table 7 cannot be fully assessed or reproduced. The exact matching procedure and its code should be provided in the repository, along with diagnostics showing that the matching balances work count and other intended variables.","section":"§7.1"}],"minor_comments":[{"comment":"The caption of Figure 7 says candidates for closed citation loops 'lie near the diagonal,' but the detector score is the product of the within-corpus citation share and within-corpus citation count; the relationship between the plotted axes and the score component should be stated more explicitly.","section":"§6.5"},{"comment":"The phrase 'the only robust signal' in the abstract and in the main text is used for the full positive set, but the following paragraph reports additional surviving features (FWCI, median citations, cites out) in the integrity subset; the wording should clarify that 'only' applies to the full delisted set, not to the integrity subset.","section":"§7.3"},{"comment":"The labeling of the case studies is inconsistent: the LLM corpus is called 'case study B' in an appendix while Study C is described as 'the third study'; the numbering is confusing and should be harmonized.","section":"Appendix"},{"comment":"Several author names in Table 3 contain broken or escaped diacritics (e.g., 'R\\'obert'), and the table cells would benefit from rendering the names with proper Unicode characters.","section":"Table 3"},{"comment":"The statement that the evaluation is 'in-sample' is unclear because the study appears to compute descriptive AUCs without fitting any model; if no parameters are estimated, the term should be removed or replaced by a precise description of what is assessed on the same venues.","section":"§7.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for the journal and the open-source artifact is a genuine strength. The field-matching concern in Study C is the most important issue: without an analysis that controls for field composition, the headline AUC 0.70 does not uniquely identify an integrity link. The gaming-resistance claim also needs to be presented as conditional on the attack model. Neither issue seems fatal, but both require additional analysis or careful rewriting, so I recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing before you read it. First, the genuinely useful material is the confound warning and the negative result: a naive case-control design against delisted journals detects prominence, not integrity — median citations, FWCI, and output volume show AUC 0.77–0.86 before size-matching and fall to chance after. Citation insularity, the popular suspect, is dead as a distinguishing signal (AUC 0.38, with delisted venues self-citing less). Those findings hold up. Second, the surviving positive result — field entropy, AUC 0.70 — is measured honestly but the abstract states it more firmly than the design allows. Matching is by size and indexing status only; the paper's Section 7.5 says field and age are not matched. The 2.33 vs 1.82 entropy gap could be compositional — broad multi-field venues among the positives, narrower controls — rather than integrity-linked scope creep. Your stress-test concern lands. The paper does not hide this; it is the stated limitation. But the abstract should say 'after size-matching,' not 'the only robust surviving signal,' and 'scope creep' outruns the evidence: a cross-sectional entropy measure shows breadth, not creep.\n\nThe rest of the design is careful. Study C uses matched controls and bootstrap CIs, separates the DOAJ integrity subset (n=40, wide CIs), and finds a plausible secondary signature there — high citation impact plus broad scope. The PageRank gaming test discloses its own Sybil limitation, and the 84x vs 8.5x inflation is real for the injected scenario; calling it a proven general property would overreach, but the paper mostly stays inside the scenario. The third detector (thematically isolated venues) sits in tension with Study C, which finds insularity dead and breadth, not narrowness, associated with delisting. The paper acknowledges this in Section 6.5, so it is an unresolved tension in their own toolkit, not a hidden contradiction — still worth pressing at revision. Reproducibility is a stated third contribution, yet the arXiv text carries no repository link or commit hash, only CLI commands. For a paper whose pitch is auditable open tooling, that is a concrete gap.\n\nCircularity is genuinely low: validation uses external ground truth, and the companion-review self-citation is not a fitted-parameter loop. The institutional case study is a competent demonstration, not a strong result, and the paper mostly does not oversell it.\n\nThis is for people working on publishing-integrity screening and open bibliometrics: a reusable template, a real confound lesson, and one modest honest signal. It deserves a serious referee. The path to acceptance is clear — extend matching to field and age, link and version the code, soften the abstract. I would send it to review.","headline":"A careful, openly honest screening pipeline whose real contributions are the prominence confound and the death of insularity; the field-entropy signal still needs field-matched controls, and the abstract overstates it — but it deserves a serious referee.","tokens_in":15923,"tokens_out":9504,"would_cite":true,"duration_ms":86787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A heterogeneous graph model of scholarly publishing finds that after control matching, broad disciplinary scope is the only surviving structural signal, and graph-based prestige resists citation cartels about ten times better than…","keywords":["heterogeneous graph model","scholarly communication networks","publishing integrity screening","field entropy","PageRank prestige","citation cartels","community detection","case-control validation"],"falsifier":"A concrete test is to repeat the venue discrimination with controls matched additionally by disciplinary field and by age: if field entropy's AUC falls to about 0.5, the paper's main validation claim fails.","tokens_in":14927,"feed_emoji":"🔗","tokens_out":7993,"duration_ms":76734,"temperature":0.7,"pith_summary":"The paper tries to establish that a heterogeneous graph model over open bibliometric data, analysed through homogeneous projections, can turn screening for questionable publishing practices from black-box classification into ranked, evidence-backed signals for human review. Its controlled validation claims that after removing a prominence confound with size-matched controls, the only robust structural feature distinguishing delisted venues is broad disciplinary scope (field entropy, AUC 0.70). It further claims that a graph-based prestige measure over the journal citation network tracks a standard count-based impact proxy while resisting a simulated citation-cartel attack about ten times better. A reader would care because this offers an open, reproducible, structurally interpretable alternative to proprietary impact metrics and a defensible screening signal for integrity.","feed_headline":"Matched test: broad scope alone flags delisted venues","feed_subtitle":"A PageRank-style journal metric tracks impact while withstanding fake-citation attacks about ten times better than raw counts.","key_machinery":"The load-bearing machinery is the projection principle: instead of running algorithms directly on a mixed typed graph, the model derives homogeneous graphs—a directed Work-to-Work citation projection, an undirected weighted Author-to-Author co-authorship projection, and a Work-to-Field transitive mapping—and computes structural metrics there. The robust screening signal is field entropy, the Shannon entropy of a venue's field distribution, which operationalises 'disciplinary scope.' The prestige machinery is PageRank on the venue citation network with self-citations excluded, a network-prestige score that discounts citations from low-prestige sources and thereby resists bulk cartel citations. The matched case-control design, comparing delisted venues to size-matched still-indexed venues, is what separates the true signal from the prominence confound, and it is this combination of projection-based structural metrics and matched-control validation that carries the paper's argument.","core_discovery":"On the paper's own terms, the central discovery is that in a controlled, size-matched comparison of venues delisted by major indexing services against still-indexed venues, most intuitive structural signals—citation counts, self-citation share, output volume—are confounded by prominence and lose their discrimination, leaving the Shannon entropy of the venue's disciplinary field distribution as the one robust structural marker (AUC 0.70). A second discovery is that a PageRank-style prestige score computed on the venue citation graph tracks the standard count-based impact proxy (Spearman 0.49) and is roughly ten times more resistant to an injected citation cartel (an 84-fold inflation of the count metric versus 8.5-fold for graph prestige). These results support the paper's larger claim that graph-based, projection-centred analysis of open bibliometric data can be a practical screening layer for publishing integrity.","pith_inferences":["A natural extension, not tested in the paper, is to monitor a venue's field entropy over time; if scope creep precedes delisting, an entropy-trajectory early-warning tool could be built.","The demonstrated gaming resistance applies to one attack model (bulk fake citations from low-prestige sources); a determined attacker could first build prestige for fake venues, so practical use should pair the metric with cartel detection, as the paper itself notes.","The prominence confound identified here likely affects several published 'predatory journal detector' results, and reanalysis with size-matched controls could be a low-cost test of that literature.","Because matching was only by size and indexing status, the field-entropy result needs field- and age-matched replication before it is used operationally."],"forward_implications":["Institutions can obtain research-intelligence outputs—research groups, cross-disciplinary bridges, and anomaly candidates—from open data on commodity hardware, complementing expert review rather than replacing it.","The three screening detectors should be used to rank candidates with accompanying structural evidence, not to make binary integrity judgments.","In publishing-integrity research, matched control designs are mandatory; naive case-control comparisons can mistake 'was prominent' for 'is problematic.'","Venues with broadened disciplinary scope (scope creep) are the most defensible structural target for manual integrity review.","Open graph-based prestige can serve as a drop-in replacement for count-based impact metrics with substantially greater resistance to simple citation-cartel attacks."],"supporting_citations":[{"why":"Supplies the fully open scholarly metadata index that the entire graph model is built on.","marker":"[Priem et al., 2022]"},{"why":"Validates the reference coverage of the open index against proprietary databases, justifying its use for citation graphs.","marker":"[Culbert et al., 2025]"},{"why":"Provides the ground-truth rationale for treating delisted venues as signals of problematic publishing.","marker":"[Richtig et al., 2023]"},{"why":"Supplies the network-based prestige precedent that the venue PageRank metric extends.","marker":"[West et al., 2010]"},{"why":"Supplies the citation-cartel detection concept and the anomaly model used in the gaming-resistance test.","marker":"[Kojaku et al., 2021]"},{"why":"Supplies the community-detection algorithm used to reconstruct research groups in the institutional case study.","marker":"[Blondel et al., 2008]"},{"why":"Supplies the improved community-detection algorithm used to verify the robustness of the partitions.","marker":"[Traag et al., 2019]"},{"why":"Documents fabricated co-authorship networks, motivating the dense-co-authorship-clique screening detector.","marker":"[Porter and McIntosh, 2024]"}],"fun_headline_variants":["Only disciplinary breadth survives delisted-venue test","PageRank-style metric resists citation gaming 10x","Graph prestige withstands fake citations better than counts","Breadth of scope: sole robust signal for delisted venues"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The size-matched, still-indexed control venues are an adequate comparison, so the surviving field-entropy signal (AUC 0.70) reflects the problematic nature of delisted venues rather than their different field mix or age.","fun_headline_variants_meta":{"raw":{"variants":["Only disciplinary breadth survives delisted-venue test","PageRank-style metric resists citation gaming 10x","Graph prestige withstands fake citations better than counts","Breadth of scope: sole robust signal for delisted venues"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000486,"raw_usage":{"total_tokens":2445,"prompt_tokens":1040,"completion_tokens":1405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":656,"completion_tokens_details":{"reasoning_tokens":1352}},"tokens_in":656,"tokens_out":1405,"duration_ms":9466,"temperature":1.0,"reasoning_tokens":1352,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:37:15.275340+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to repeat the venue discrimination with controls matched additionally by disciplinary field and by age: if field entropy's AUC falls to about 0.5, the paper's main validation claim fails.","supporting_citations":[{"cited_title":"2022 , title =","cited_arxiv_id":null,"evidence_quote":"Supplies the fully open scholarly metadata index that the entire graph model is built on."},{"cited_title":"Predatory","cited_arxiv_id":null,"evidence_quote":"Provides the ground-truth rationale for treating delisted venues as signals of problematic publishing."},{"cited_title":"and Bergstrom, Theodore C","cited_arxiv_id":null,"evidence_quote":"Supplies the network-based prestige precedent that the venue PageRank metric extends."}],"review_version":1}