{"id":"e90144cc-e470-4dc3-b194-da666db4fef2","arxiv_id":"2503.13448","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review finds that deep learning AND methods, especially hybrid supervised-unsupervised approaches, improve performance but rely heavily on the AMiner dataset.","lead":"This paper is a systematic review of deep learning methods for author name disambiguation (AND) published between 2016 and 2024. It organizes 28 selected studies into supervised, unsupervised, and hybrid categories, and compares their reported F1 scores on common benchmarks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 ranks F1 scores across incompatible AMiner versions and evaluation protocols, so the 'hybrid methods are state-of-the-art' conclusion is not supported by the reported comparison.","rationale":"The paper's contribution is a systematic review, so its central claim is descriptive and comparative: deep learning has significantly advanced AND, and hybrid supervised/unsupervised methods are state of the art. The only quantitative evidence for the hybrid-SOTA subclaim is Section 4.4 and Table 1. My concern is not that the individual F1 scores are false; it is that they are not commensurable. The authors themselves state the ideal comparison criteria at the top of Section 4.4, then present a single ranked table whose 'AMiner' column mixes AMiner-WhoIsWho v3, AMiner-534K, AMiner plus CiteSeerX, AMiner plus Semantic Scholar, AMiner plus OAG, AMiner plus SNSF, and AMiner plus PubMed. Whether Xie et al. (89.7) outperforms Cheng et al. (87.72) depends on benchmark difficulty, not just method quality. This is load-bearing because the title, abstract, and conclusion specifically highlight hybrid methods as the key advance; without a valid comparison the claim is unsupported. I also note traceability issues: reference numbering in Table 1 is internally inconsistent, and the F1 values are not accompanied by confidence intervals or protocol details. If the concrete test passes—scores survive re-evaluation on one common benchmark—the qualitative conclusions about DL integration and AMiner dominance remain intact; if it fails, the SOTA ranking and the hybrid claim should be tempered. This agrees with the reader's weakest_assumption, so no verdict change is needed.","tokens_in":13843,"tokens_out":3297,"duration_ms":32111,"concrete_test":"For every row in Table 1, retrieve the cited paper and record the exact dataset version, train/test split, and F1 definition (pairwise vs. cluster-level, micro vs. macro). Then select the top three methods and re-evaluate them on a single held-out subset of AMiner-WhoIsWho v3 under an identical protocol. If the ranking changes or the top F1 scores are within noise, the hybrid state-of-the-art claim in Section 6 must be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central comparative claim—that hybrid methods achieve state-of-the-art AND results—rests on the ranked F1 column in Table 1 and the prose in Section 4.4. That column violates the authors' own stated comparability criteria (Section 4.4: same or comparable datasets; same metrics). Rows labeled 'AMiner' actually use different resources: Xie et al. (2022) is AMiner + CiteSeerX; Cheng et al. (2024) is AMiner-WhoIsWho v3; Gong et al. (2024) is AMiner-WhoIsWho + OAG; Sun et al. (2020) is AMiner + Semantic Scholar; Rettig et al. (2022) is AMiner + SNSF; Kim et al. (2019) is AMiner + PubMed. These dataset versions differ in size, labeling, and paper distribution, and the papers use different train/test splits and F1 definitions (pairwise vs. cluster-level, micro vs. macro). Absent a controlled re-evaluation, the headline 'hybrid methods... demonstrated state-of-the-art results' is an artifact of table construction rather than a measured finding. Traceability is further weakened by reference errors in the same table: [15] is cited for both Zhang, Yu, Liu, & Wang (2020) and Zhang, Y., Zhang, F., Yao, P., & Tang (2018), and [37] is listed as Yan et al. (2020) while the reference is Yan et al. (2019). These inconsistencies make verification of the reported scores difficult.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a literature review of deep learning-based author name disambiguation (AND) methods published between 2016 and 2024. It categorizes 28 selected studies into supervised, unsupervised, and mixed (hybrid) learning strategies, and provides a ranked F1-score comparison across methods evaluated on datasets derived from AMiner. The paper argues that deep learning has significantly advanced AND, that hybrid methods achieve state-of-the-art performance, and that the field's heavy reliance on AMiner limits generalizability.","tokens_in":14101,"tokens_out":6298,"duration_ms":51250,"significance":"If properly supported, the survey would fill a gap in recent AND reviews and provide a useful taxonomy of deep learning methods. The authors correctly identify a real and widely acknowledged challenge: the lack of diverse benchmarks and standardized evaluation protocols in AND. The compilation of recent methods is informative, and the discussion of AMiner's limitations is timely. However, the central comparative claim—that hybrid methods are state-of-the-art—rests on a ranked F1 comparison that pools results from incompatible dataset versions and evaluation protocols, and the 'systematic review' methodology is underdocumented. The descriptive synthesis is a plausible starting point, but the ranking and the hybrid-superiority conclusion should not be taken as evidence until the comparison is redone on compatible benchmarks with identical metrics, or the ranking is removed.","major_comments":[{"comment":"The ranked F1 comparison in Table 1 and Section 4.4 does not satisfy the two comparability criteria the authors themselves state at the beginning of Section 4.4. Rows labeled 'AMiner' use different dataset versions and extensions: AMiner + CiteSeerX, AMiner-WhoIsWho v3, AMiner-WhoIsWho + OAG, AMiner + Semantic Scholar, AMiner + SNSF, AMiner + PubMed, and AMiner + OC. These resources differ in scale, labeling protocol, and paper coverage, and the underlying papers report F1 computed on different train/test splits and possibly different granularities (pairwise vs. cluster-level, micro vs. macro). Consequently, the ranking and the Section 6 conclusion that 'hybrid methods ... demonstrated state-of-the-art results' are not supported by the evidence as presented. The authors should either remove the cross-method ranking and report per-benchmark results without ordering, or re-evaluate the methods on a common benchmark and a common metric.","section":"§4.4, Table 1"},{"comment":"The search strategy is described too loosely to support the label 'systematic review' used in the abstract and Section 1. The Google Scholar component takes only the first 50 results for one keyword pair, and the PURE Suggest snowballing is described as 'repeated twice' without specifying inclusion/exclusion criteria at the full-text screening stage, screening decisions, deduplication, or a flow diagram. No search dates are given. The authors should either document a reproducible protocol (including the full query, search dates, and the number of records excluded at each step) or revise the manuscript to present the work as a scoping or narrative review.","section":"§3.2"},{"comment":"There are multiple reference inconsistencies that prevent traceability of the reported F1 scores. In Table 1, '[15]' is attached to both Zhang, Yu, Liu, & Wang (2020) and Zhang, Y., Zhang, F., Yao, P., & Tang (2018), while the latter is reference [34] in the bibliography. The 'Yan et al.' entry is labeled (2020)[37] in the Section 4.4 text and (2019)[37] in Table 1, whereas reference [37] is a 2019 GCN paper and reference [39] is a different 2024 paper. Additionally, the DBLP sentence in Section 4.4 includes 'B[18]' with a spurious 'B'. These inconsistencies make verification of the reported scores difficult and must be corrected for the survey to be usable.","section":"§4.4, Table 1, References"},{"comment":"The central notion of 'hybrid' is used with two different meanings. In Section 4.1, Kim et al. (2019) is presented as a 'hybrid' method that combines structural and global features in a fully supervised pairwise classifier; in Sections 4.3 and 5, the term is reserved for pipelines that combine supervised and unsupervised learning. Moreover, Xie et al. (2022), classified as S+U in Table 1, is described in Section 4.3 without a clear supervised component. Because the headline finding depends on the hybrid category, the authors need to fix one operational definition of 'hybrid' and re-classify each method consistently before the state-of-the-art assertion can be assessed.","section":"§4.1, §4.3, §5"}],"minor_comments":[{"comment":"Two in-text citations are malformed and do not appear in the reference list: '(Zhang, Li & Lu, Wei & Yang, Jinqing. (2021). Biases in datasets...' and '(Sanyal, D. K., Bhowmick, P. K., & Das, P. P. (2021).' Replace these with proper numbered references or add full bibliographic entries.","section":"§2"},{"comment":"The sentence 'Among those based on the DBLP dataset, we have [16] with an F1 score of 0.98 and B[18] with a score of 0.975...' contains a stray 'B' before '[18]', and the F1 values are given in different numeric conventions (0.98 vs. 0.975) without stating the scale; normalize the notation.","section":"§4.4"},{"comment":"The year assigned to the CONNA entry is inconsistent: the prose in Section 4.1 cites 'Zhao et al. (2022) [20]' while Table 1 and Section 4.4 refer to 'Zhao et al. (2020)'. Correct the year and make the citation style uniform.","section":"§4.4, Table 1"},{"comment":"The phrase 'Looking at 1' appears to be missing 'Table'; it should read 'Looking at Table 1'.","section":"§5"},{"comment":"Reference [4] has an incomplete title '(????)' for Baglioni et al.; provide the full title and bibliographic details. Also verify that all DOI strings are complete and correctly formatted.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript falls within the scope of the venue and the topic is timely. The main concerns are fixable: the F1-ranking should be replaced with per-benchmark reporting or a controlled re-evaluation, the methodology section needs a reproducible protocol or a revised 'scoping review' label, and the reference errors must be corrected. If the authors cannot supply a controlled comparison, the claim of hybrid superiority should be removed or heavily qualified. The self-citation overlap (Santini et al., which includes an author) does not by itself affect the central claim, but the authors should ensure all cited scores are attributed to their original sources accurately."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a competent, useful survey of deep learning author name disambiguation (2016–2024), but its headline comparative claim—that hybrid methods have achieved state-of-the-art results—rests on a Table 1 that mixes F1 scores from different AMiner versions, additional datasets, and evaluation protocols. The conclusion is not as strong as the table suggests.\n\nWhat is actually new: existing surveys (Ferreira, Sanyal, De Bonis) cover earlier periods or graph methods. This one collects DL-specific work from 2016–2024, organizes it into supervised/unsupervised/hybrid, and gives a decent qualitative read of trends. The observation that almost everything is evaluated on AMiner and that this limits generalizability is correct and worth saying. The taxonomy is standard, but the synthesis is clear and the limitations section is honest.\n\nWhere I think the paper is soft: Section 4.4 and Table 1. The authors state the right criteria—same or comparable datasets, same metrics—and then violate them in the very table that supports the ranking. Look at the rows: Xie et al. use AMiner + CiteSeerX, Gong et al. use AMiner-WhoIsWho + OAG, Rettig et al. use AMiner + SNSF, Kim et al. use AMiner + PubMed. Those are not comparable evaluations. F1 definitions also differ (pairwise vs cluster-level, micro vs macro). So saying hybrid methods are state-of-the-art because the top row is a hybrid is an artifact of table construction. That doesn't sink the paper, because the qualitative point—hybrid methods are a promising direction—is supportable from the narrative and from the general pattern. But the specific ranking and the \"demonstrated state-of-the-art results\" sentence should be tempered or replaced with a call for controlled re-evaluation.\n\nThere are also minor traceability problems: [15] is used for both the 2020 Zhang et al. and the 2018 Zhang, Zhang, Yao, Tang paper (which is actually [34] in the references), and Yan et al. is listed as 2020 in Table 1 but the reference is 2019. These are easy fixes but they make it harder to verify the reported numbers.\n\nMethodology note: \"first 50 Google Scholar results\" is a thin basis for a systematic review. It doesn't invalidate the survey, but the authors should be transparent that coverage is partial and that the search strategy isn't reproducible in the usual sense.\n\nWho is this for: researchers working in AND or digital libraries who want a quick orientation to DL methods in the last eight years. It's a useful entry point, not a definitive benchmark comparison.\n\nMy take for review: I'd send it to peer review, but with a request to fix Table 1, either by restricting the comparison to truly comparable setups or by clearly labeling the rows as indicative rather than ranked, and to correct the reference inconsistencies. The paper deserves serious referee time because it fills a real gap and its qualitative conclusions are largely sound.","headline":"A useful DL-AND survey whose ranked F1 comparison overreaches because Table 1 mixes incompatible benchmarks; the qualitative conclusions are fine.","tokens_in":14633,"tokens_out":3968,"would_cite":true,"duration_ms":32007,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep learning has advanced author name disambiguation, but a survey of 28 studies finds hybrid methods lead while heavy reliance on the AMiner dataset limits how much the rankings can be trusted.","keywords":["author name disambiguation","deep learning","systematic review","hybrid methods","AMiner benchmark","supervised learning","unsupervised learning","bibliographic metadata"],"falsifier":"Re-run the top three methods (Xie et al., BOND, CONNA) on a single fixed benchmark with identical train/test splits and the same metric definitions; if the hybrid method no longer has the highest F1, the paper's central claim that hybrids are state-of-the-art would be undercut.","tokens_in":13619,"feed_emoji":"📚","tokens_out":5181,"duration_ms":44624,"temperature":0.7,"pith_summary":"This paper systematically reviews deep learning approaches to author name disambiguation (AND) published between 2016 and 2024, covering 28 studies. It argues that deep learning has substantially advanced AND by letting models combine structured metadata (co-authors, affiliations) with unstructured text (titles, abstracts), and that hybrid methods—those mixing supervised and unsupervised learning—currently achieve the highest reported scores. At the same time, it finds that most methods are evaluated on the AMiner dataset or its variants, which overrepresents Chinese names, and that inconsistent dataset versions and metrics make reported F1 scores hard to compare. The paper concludes that the field's next bottleneck is not model design but the scarcity of diverse, standardized benchmarks.","feed_headline":"Hybrid deep learning tops author name disambiguation","feed_subtitle":"Survey of 28 studies: hybrids beat pure supervised or unsupervised, yet AMiner-only tests weaken rankings.","key_machinery":"The review's organizing tool is a three-way taxonomy of deep learning approaches: supervised methods (trained on labeled author identities), unsupervised methods (clustering or embedding without labels), and hybrid/mixed methods that combine both. The load-bearing comparison device is a table that ranks studies by F1 score on AMiner-derived benchmarks, supplemented by results on DBLP, CiteSeerX, PubMed, and Scopus. This combination of taxonomy and benchmark table is what carries the conclusion that hybrid methods are state-of-the-art and that dataset diversity is the main open problem.","core_discovery":"The central claim is that deep learning has significantly improved author name disambiguation, and within this landscape hybrid approaches that combine supervised and unsupervised learning outperform purely supervised or purely unsupervised ones. The review's comparative table of F1 scores on AMiner-derived datasets places a hybrid method (Xie et al., 2022) at the top with 89.7, followed by an unsupervised method (BOND, 87.72) and a supervised method (CONNA, 86.22), showing no single paradigm dominates. The authors also find that the availability of labeled data in AMiner does not guarantee higher performance, and that the heavy reliance on AMiner—plus the lack of a standardized evaluation framework—limits confidence in the generalizability of these results.","pith_inferences":["If the AMiner-centric ranking is an artifact of the benchmark rather than a true property of methods, re-evaluating the top hybrid and supervised models on a deliberately non-AMiner dataset (e.g., a random sample of DBLP or Scopus) would likely reshuffle the order.","The review's suggestion that data diversity is the critical challenge implies that building labeled multi-script, multi-domain benchmarks could improve AND more than any single architectural innovation.","The authors' timeframe ends in early 2024, so large language model-based disambiguation pipelines—which may change the cost/benefit balance of supervised versus unsupervised approaches—are only partially represented in the surveyed evidence.","One testable extension: a meta-analysis regressing reported F1 on dataset version and metric definition could quantify how much of the performance gap between methods is actually explained by evaluation protocol differences."],"forward_implications":["Future AND systems should treat hybrid designs—supervised representation learning with unsupervised clustering—as the default starting point, since they currently lead the reported ranking.","A community benchmark with fixed splits and unified metrics across AMiner-WhoIsWho, DBLP, and PubMed would be needed to verify whether hybrid leadership holds outside AMiner.","Because AMiner overrepresents Chinese names, improvements measured there may not transfer to bibliographic databases with more Western or multi-script names.","Supervised labels alone are not a performance guarantee, so publication venues and data characteristics deserve as much attention as architecture choice.","The absence of standardized evaluation is itself a barrier to progress, since it prevents apples-to-apples comparison of new methods."],"supporting_citations":[{"why":"Provides the AMiner-WhoIsWho benchmark, the primary labeled dataset used by most reviewed methods and the basis for the comparative F1 table.","marker":"[11]"},{"why":"The top-ranked hybrid method in the comparison, used to support the conclusion that hybrid approaches achieve the best results.","marker":"[36]"},{"why":"BOND, the top-ranked unsupervised method, shows that unsupervised approaches can compete without labeled data and anchors the hybrid-versus-unsupervised comparison.","marker":"[30]"},{"why":"CONNA, the top-ranked supervised method, provides the supervised baseline and highlights the real-time assignment setting.","marker":"[20]"},{"why":"MORE, a recent hybrid framework, is cited as achieving state-of-the-art performance, reinforcing the hybrid trend.","marker":"[14]"},{"why":"The 2018 AMiner clustering and human-in-the-loop system is an early hybrid baseline that later methods are compared against.","marker":"[34]"},{"why":"The knowledge-graph embedding approach evaluated on AMiner and OpenCitations illustrates the dataset diversity challenges the review highlights.","marker":"[12]"},{"why":"The graph-based AND survey that this review positions itself against, establishing the gap of deep-learning-focused reviews that this paper fills.","marker":"[7]"}],"fun_headline_variants":["Hybrid deep learning dominates author disambiguation","Deep learning survey: hybrids beat pure AND methods","Author name disambiguation: hybrid DL leads the pack","Hybrid models top deep learning for author names","Review: hybrid DL best for author name disambiguation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's ranking and its conclusion that hybrid methods are best assume that F1 scores reported by different papers are comparable despite the papers using different versions of the AMiner dataset and different evaluation protocols.","fun_headline_variants_meta":{"raw":{"variants":["Hybrid deep learning dominates author disambiguation","Deep learning survey: hybrids beat pure AND methods","Author name disambiguation: hybrid DL leads the pack","Hybrid models top deep learning for author names","Review: hybrid DL best for author name disambiguation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1412,"prompt_tokens":834,"completion_tokens":578,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":450,"completion_tokens_details":{"reasoning_tokens":506}},"tokens_in":450,"tokens_out":578,"duration_ms":6457,"temperature":1.0,"reasoning_tokens":506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:07:04.734842+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the top three methods (Xie et al., BOND, CONNA) on a single fixed benchmark with identical train/test splits and the same metric definitions; if the hybrid method no longer has the highest F1, the paper's central claim that hybrids are state-of-the-art would be undercut.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the AMiner-WhoIsWho benchmark, the primary labeled dataset used by most reviewed methods and the basis for the comparative F1 table."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The top-ranked hybrid method in the comparison, used to support the conclusion that hybrid approaches achieve the best results."},{"cited_title":"Cheng, B","cited_arxiv_id":null,"evidence_quote":"BOND, the top-ranked unsupervised method, shows that unsupervised approaches can compete without labeled data and anchors the hybrid-versus-unsupervised comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CONNA, the top-ranked supervised method, provides the supervised baseline and highlights the real-time assignment setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MORE, a recent hybrid framework, is cited as achieving state-of-the-art performance, reinforcing the hybrid trend."},{"cited_title":"Zhang, F","cited_arxiv_id":null,"evidence_quote":"The 2018 AMiner clustering and human-in-the-loop system is an early hybrid baseline that later methods are compared against."},{"cited_title":"A Knowledge Graph Embeddings based Approach for Author Name Disambiguation using Literals","cited_arxiv_id":"2201.09555","evidence_quote":"The knowledge-graph embedding approach evaluated on AMiner and OpenCitations illustrates the dataset diversity challenges the review highlights."},{"cited_title":"De Bonis, F","cited_arxiv_id":null,"evidence_quote":"The graph-based AND survey that this review positions itself against, establishing the gap of deep-learning-focused reviews that this paper fills."}],"review_version":1}