REVIEW 4 major objections 5 minor 41 references
Recent Developments in Deep Learning-based Author Name Disambiguation
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Deep learning has advanced author name disambiguation, but a survey of 28 studies finds hybrid methods lead while heavy reliance on the AMiner dataset limits how much the rankings can be trusted.
desk verdict A useful DL-AND survey whose ranked F1 comparison overreaches because Table 1 mixes incompatible benchmarks; the qualitative conclusions are fine. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The review's organizing tool is a three-way taxonomy of deep learning approaches: supervised methods (trained on labeled author identities), unsupervised methods (clustering or embedding without labels), and hybrid/mixed methods that combine both. The load-bearing comparison device is a table that ranks studies by F1 score on AMiner-derived benchmarks, supplemented by results on DBLP, CiteSeerX, PubMed, and Scopus. This combination of taxonomy and benchmark table is what carries the conclusion that hybrid methods are state-of-the-art and that dataset diversity is the main open problem.
What would settle it
Re-run the top three methods (Xie et al., BOND, CONNA) on a single fixed benchmark with identical train/test splits and the same metric definitions; if the hybrid method no longer has the highest F1, the paper's central claim that hybrids are state-of-the-art would be undercut.
Extended reading notes
Core claim
The central claim is that deep learning has significantly improved author name disambiguation, and within this landscape hybrid approaches that combine supervised and unsupervised learning outperform purely supervised or purely unsupervised ones. The review's comparative table of F1 scores on AMiner-derived datasets places a hybrid method (Xie et al., 2022) at the top with 89.7, followed by an unsupervised method (BOND, 87.72) and a supervised method (CONNA, 86.22), showing no single paradigm dominates. The authors also find that the availability of labeled data in AMiner does not guarantee higher performance, and that the heavy reliance on AMiner—plus the lack of a standardized evaluation framework—limits confidence in the generalizability of these results.
Load-bearing premise
The review's ranking and its conclusion that hybrid methods are best assume that F1 scores reported by different papers are comparable despite the papers using different versions of the AMiner dataset and different evaluation protocols.
Editorial extensions
If this is right
- Future AND systems should treat hybrid designs—supervised representation learning with unsupervised clustering—as the default starting point, since they currently lead the reported ranking.
- A community benchmark with fixed splits and unified metrics across AMiner-WhoIsWho, DBLP, and PubMed would be needed to verify whether hybrid leadership holds outside AMiner.
- Because AMiner overrepresents Chinese names, improvements measured there may not transfer to bibliographic databases with more Western or multi-script names.
- Supervised labels alone are not a performance guarantee, so publication venues and data characteristics deserve as much attention as architecture choice.
- The absence of standardized evaluation is itself a barrier to progress, since it prevents apples-to-apples comparison of new methods.
Reading between the lines
- If the AMiner-centric ranking is an artifact of the benchmark rather than a true property of methods, re-evaluating the top hybrid and supervised models on a deliberately non-AMiner dataset (e.g., a random sample of DBLP or Scopus) would likely reshuffle the order.
- The review's suggestion that data diversity is the critical challenge implies that building labeled multi-script, multi-domain benchmarks could improve AND more than any single architectural innovation.
- The authors' timeframe ends in early 2024, so large language model-based disambiguation pipelines—which may change the cost/benefit balance of supervised versus unsupervised approaches—are only partially represented in the surveyed evidence.
- One testable extension: a meta-analysis regressing reported F1 on dataset version and metric definition could quantify how much of the performance gap between methods is actually explained by evaluation protocol differences.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents a literature review of deep learning-based author name disambiguation (AND) methods published between 2016 and 2024. It categorizes 28 selected studies into supervised, unsupervised, and mixed (hybrid) learning strategies, and provides a ranked F1-score comparison across methods evaluated on datasets derived from AMiner. The paper argues that deep learning has significantly advanced AND, that hybrid methods achieve state-of-the-art performance, and that the field's heavy reliance on AMiner limits generalizability.
Significance. If properly supported, the survey would fill a gap in recent AND reviews and provide a useful taxonomy of deep learning methods. The authors correctly identify a real and widely acknowledged challenge: the lack of diverse benchmarks and standardized evaluation protocols in AND. The compilation of recent methods is informative, and the discussion of AMiner's limitations is timely. However, the central comparative claim—that hybrid methods are state-of-the-art—rests on a ranked F1 comparison that pools results from incompatible dataset versions and evaluation protocols, and the 'systematic review' methodology is underdocumented. The descriptive synthesis is a plausible starting point, but the ranking and the hybrid-superiority conclusion should not be taken as evidence until the comparison is redone on compatible benchmarks with identical metrics, or the ranking is removed.
major comments (4)
- [§4.4, Table 1] The ranked F1 comparison in Table 1 and Section 4.4 does not satisfy the two comparability criteria the authors themselves state at the beginning of Section 4.4. Rows labeled 'AMiner' use different dataset versions and extensions: AMiner + CiteSeerX, AMiner-WhoIsWho v3, AMiner-WhoIsWho + OAG, AMiner + Semantic Scholar, AMiner + SNSF, AMiner + PubMed, and AMiner + OC. These resources differ in scale, labeling protocol, and paper coverage, and the underlying papers report F1 computed on different train/test splits and possibly different granularities (pairwise vs. cluster-level, micro vs. macro). Consequently, the ranking and the Section 6 conclusion that 'hybrid methods ... demonstrated state-of-the-art results' are not supported by the evidence as presented. The authors should either remove the cross-method ranking and report per-benchmark results without ordering, or re-evaluate the methods on a common benchmark and a common metric.
- [§3.2] The search strategy is described too loosely to support the label 'systematic review' used in the abstract and Section 1. The Google Scholar component takes only the first 50 results for one keyword pair, and the PURE Suggest snowballing is described as 'repeated twice' without specifying inclusion/exclusion criteria at the full-text screening stage, screening decisions, deduplication, or a flow diagram. No search dates are given. The authors should either document a reproducible protocol (including the full query, search dates, and the number of records excluded at each step) or revise the manuscript to present the work as a scoping or narrative review.
- [§4.4, Table 1, References] There are multiple reference inconsistencies that prevent traceability of the reported F1 scores. In Table 1, '[15]' is attached to both Zhang, Yu, Liu, & Wang (2020) and Zhang, Y., Zhang, F., Yao, P., & Tang (2018), while the latter is reference [34] in the bibliography. The 'Yan et al.' entry is labeled (2020)[37] in the Section 4.4 text and (2019)[37] in Table 1, whereas reference [37] is a 2019 GCN paper and reference [39] is a different 2024 paper. Additionally, the DBLP sentence in Section 4.4 includes 'B[18]' with a spurious 'B'. These inconsistencies make verification of the reported scores difficult and must be corrected for the survey to be usable.
- [§4.1, §4.3, §5] The central notion of 'hybrid' is used with two different meanings. In Section 4.1, Kim et al. (2019) is presented as a 'hybrid' method that combines structural and global features in a fully supervised pairwise classifier; in Sections 4.3 and 5, the term is reserved for pipelines that combine supervised and unsupervised learning. Moreover, Xie et al. (2022), classified as S+U in Table 1, is described in Section 4.3 without a clear supervised component. Because the headline finding depends on the hybrid category, the authors need to fix one operational definition of 'hybrid' and re-classify each method consistently before the state-of-the-art assertion can be assessed.
minor comments (5)
- [§2] Two in-text citations are malformed and do not appear in the reference list: '(Zhang, Li & Lu, Wei & Yang, Jinqing. (2021). Biases in datasets...' and '(Sanyal, D. K., Bhowmick, P. K., & Das, P. P. (2021).' Replace these with proper numbered references or add full bibliographic entries.
- [§4.4] The sentence 'Among those based on the DBLP dataset, we have [16] with an F1 score of 0.98 and B[18] with a score of 0.975...' contains a stray 'B' before '[18]', and the F1 values are given in different numeric conventions (0.98 vs. 0.975) without stating the scale; normalize the notation.
- [§4.4, Table 1] The year assigned to the CONNA entry is inconsistent: the prose in Section 4.1 cites 'Zhao et al. (2022) [20]' while Table 1 and Section 4.4 refer to 'Zhao et al. (2020)'. Correct the year and make the citation style uniform.
- [§5] The phrase 'Looking at 1' appears to be missing 'Table'; it should read 'Looking at Table 1'.
- [References] Reference [4] has an incomplete title '(????)' for Baglioni et al.; provide the full title and bibliographic details. Also verify that all DOI strings are complete and correctly formatted.
Circularity Check
No significant circularity: survey reports external results; the few self-citations are not load-bearing.
full rationale
This paper is a systematic literature survey, not a derivation or prediction paper. It reports F1 scores taken from original papers, classifies the underlying methods by learning strategy, and synthesizes qualitative findings. There is no fitted parameter later renamed as a prediction, and no claimed first-principles result that reduces to an input by construction. The only overlapping self-citations are [3] (Peroni and Shotton on OpenCitations, used as background) and [12] (Santini et al., which includes co-author Peroni, cited as the source of the AMiner-534K dataset description and as a low-ranked row in Table 1). Neither citation carries the paper's central claim that hybrid methods achieve state-of-the-art AND results; that claim is a synthesis of independently reported external numbers from many non-overlapping research groups. The skeptic's concern about Table 1 ranking F1 scores across incompatible AMiner versions and evaluation protocols is a substantive validity and comparability threat to the survey's ranking, but it is not circularity: the ranking is not an input to itself, and no argument in the paper is equivalent to its own conclusion by definition. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption The selected 28 papers are representative of the deep learning AND literature from 2016-2024.
- domain assumption Reported F1 scores from different papers are comparable despite using different versions of AMiner and different evaluation setups.
Cite this review
Pith. "Pith review of Recent Developments in Deep Learning-based Author Name Disambiguation." pith.science (2026). https://pith.science/paper/WPI7J4X3
@misc{pith2026250313448,
author = {Pith},
title = {Pith review of: Recent Developments in Deep Learning-based Author Name Disambiguation},
year = {2026},
howpublished = {\url{https://pith.science/paper/WPI7J4X3}},
note = {Machine review of arXiv:2503.13448}
}
read the original abstract
Author Name Disambiguation (AND) is a critical task for digital libraries aiming to link existing authors with their respective publications. Due to the lack of persistent identifiers used by researchers and the presence of intrinsic linguistic challenges, such as homonymy, the development of Deep Learning algorithms to address this issue has become widespread. Many AND deep learning methods have been developed, and surveys exist comparing the approaches in terms of techniques, complexity, performance. However, none explicitly addresses AND methods in the context of deep learning in the latest years (i.e. timeframe 2016-2024). In this paper, we provide a systematic review of state-of-the-art AND techniques based on deep learning, highlighting recent improvements, challenges, and open issues in the field. We find that DL methods have significantly impacted AND by enabling the integration of structured and unstructured data, and hybrid approaches effectively balance supervised and unsupervised learning.
Reference graph
Works this paper leans on
- [15]
-
[34]
Y. Zhang, F. Zhang, P. Yao, J. Tang, Name Disambiguation in AMiner: Clustering, Maintenance, and Human in the Loop., in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, ACM, London United Kingdom, 2018, pp. 1002–1011. URL: https://dl.acm.org/doi/10.1145/3219819.3219859. doi:10.1145/3219819.3219859
-
[37]
H. Yan, H. Peng, C. Li, J. Li, L. Wang, Bibliographic Name Disambiguation with Graph Convolutional Network, in: R. Cheng, N. Mamoulis, Y. Sun, X. Huang (Eds.), Web Information Systems Engineer- ing – WISE 2019, volume 11881, Springer International Publishing, Cham, 2019, pp. 538–551. URL: https://link.springer.com/10.1007/978-3-030-34223-4_34. doi: 10.100...
-
[39]
Q. Yan, AsirAsir, Synergizing Large Language Models and Tree-based Algorithms for Author Name Disambiguation, 2024. URL: https://openreview.net/forum?id=VEz1sq66pi
work page 2024
-
[18]
Bib2Auth: Deep Learning Approach for Author Disambiguation using Bibliographic Data
Z. Boukhers, N. Bahubali, A. T. Chandrasekaran, A. Anand, S. M. G. Prasadand, S. Aralappa, Bib2Auth: Deep Learning Approach for Author Disambiguation using Bibliographic Data, 2021. URL: http://arxiv.org/abs/2107.04382. doi:10.48550/arXiv.2107.04382, arXiv:2107.04382 [cs]
work page Pith review arXiv doi:10.48550/arxiv.2107.04382 2021
-
[1]
A. Ferreira, M. Gonçalves, A. Laender, A Brief Survey of Automatic Methods for Author Name Disambiguation, ACM SIGMOD Record 41 (2012) 15–26. doi:10.1145/2350036.2350040
-
[2]
A. Strotmann, D. Zhao, Author name disambiguation: What difference does it make in author-based citation analysis?, Journal of the American Society for Information Science and Technology 63 (2012) 1820–1833. URL: https://onlinelibrary.wiley.com/doi/abs/10.1002/asi.22695. doi:10.1002/ asi.22695, _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/asi.22695
-
[3]
S. Peroni, D. Shotton, Opencitations, an infrastructure organization for open scholarship, Quanti- tative Science Studies 1 (2020) 428–444. doi:10.1162/qss_a_00023
Show all 41 references
-
[4]
Baglioni, A
M. Baglioni, A. Mannocci, P. Manghi, C. Atzori, A. Bardi, S. L. Bruzzo, Reflections on the Misuses of ORCID iDs (????)
-
[5]
J. Kim, J. Kim, Effect of forename string on author name disambiguation, J. Assoc. Inf. Sci. Technol. 71 (2020) 839–855. URL: https://doi.org/10.1002/asi.24298. doi:10.1002/asi.24298
2020 doi
-
[6]
D. K. Sanyal, P. K. Bhowmick, P. P. Das, A review of author name disambiguation techniques for the PubMed bibliographic database, Journal of Information Science 47 (2021) 227–254. URL: https://doi.org/10.1177/0165551519888605. doi:10.1177/0165551519888605, publisher: SAGE Publ...
2021 doi
-
[7]
De Bonis, F
M. De Bonis, F. Falchi, P. Manghi, Graph-based methods for Author Name Disambiguation: a survey, PeerJ Computer Science 9 (2023) e1536. URL: https://peerj.com/articles/cs-1536. doi: 10. 7717/peerj-cs.1536
2023
-
[8]
Milojevic, Accuracy of simple, initials-based methods for author name disambiguation, ArXiv abs/1308.0749 (2013)
S. Milojevic, Accuracy of simple, initials-based methods for author name disambiguation, ArXiv abs/1308.0749 (2013). URL: https://api.semanticscholar.org/CorpusID:16417347
2013 arXiv
-
[9]
Manzoor, S
A. Manzoor, S. Asghar, T. Amjad, Toward a New Paradigm for Author Name Disambiguation, IEEE Access 10 (2022) 76055–76068. URL: https://ieeexplore.ieee.org/document/9826729/. doi:10. 1109/ACCESS.2022.3190088
2022
-
[10]
J. Kim, J. Kim, J. Kim, Effect of chinese characters on machine learning for chinese author name disambiguation: A counterfactual evaluation, Journal of Information Science 49 (2023) 711–725. doi:10.1177/01655515211018171
2023 doi
-
[11]
B. Chen, J. Zhang, F. Zhang, T. Han, Y. Cheng, X. Li, Y. Dong, J. Tang, Web-scale academic name disambiguation: the whoiswho benchmark, leaderboard, and toolkit, in: KDD 23: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data, ACM, 2023, pp. 3817–3828...
2023
- [12]
-
[13]
J. F. Burnham, Scopus database: a review, Biomedical Digital Libraries 3 (2006) 1. URL: https: //doi.org/10.1186/1742-5581-3-1. doi: 10.1186/1742-5581-3-1
2006 doi
-
[14]
J. Gong, X. Fang, J. Peng, Y. Zhao, J. Zhao, C. Wang, Y. Li, J. Zhang, S. Drew, MORE: Toward Improv- ing Author Name Disambiguation in Academic Knowledge Graphs, International Journal of Ma- chine Learning and Cybernetics 15 (2024) 37–50. URL: https://doi.org/10.1007/s13042-02...
2024 doi
-
[16]
Alqarni, S
Firdaus, W. Alqarni, S. Nurmaini, A. Darmawahyuni, A. I. Sapitri, M. N. Rachmatullah, S. D. Lestari, Author classification on bibliographic data using capsule networks architecture, in: 2022 9th International Conference on Electrical Engineering, Computer Science and Informati...
2022
-
[17]
Boukhers, N
Z. Boukhers, N. B. Asundi, Whois? Deep Author Name Disambiguation Using Bibliographic Data, in: G. Silvello, O. Corcho, P. Manghi, G. M. Di Nunzio, K. Golub, N. Ferro, A. Poggi (Eds.), Linking Theory and Practice of Digital Libraries, Springer International Publishing, Cham, 2...
2022 doi
-
[19]
K. Kim, S. Rohatgi, C. L. Giles, Hybrid Deep Pairwise Classification for Author Name Disambigua- tion, in: Proceedings of the 28th ACM International Conference on Information and Knowledge Management, ACM, Beijing China, 2019, pp. 2369–2372. URL: https://dl.acm.org/doi/10.1145...
2019
-
[20]
B. Chen, J. Zhang, J. Tang, L. Cai, Z. Wang, S. Zhao, H. Chen, C. Li, CONNA: Addressing Name Disambiguation on the Fly, IEEE Transactions on Knowledge and Data Engineering 34 (2022) 3139–
2022
-
[21]
S. Wang, Q. Li, R. Koopman, Co-attention-Based Pairwise Learning for Author Name Disambigua- tion, in: D. H. Goh, S.-J. Chen, S. Tuarob (Eds.), Leveraging Generative Intelligence in Digital Libraries: Towards Human-Machine Collaboration, Springer Nature, Singapore, 2023, pp. 2...
2023 doi
-
[22]
Firdaus, M
F. Firdaus, M. Anshori, S. P. Raflesia, A. Zarkasi, M. Afrina, S. Nurmaini, Deep Neural Network Struc- ture to Improve Individual Performance based Author Classification, Computer Engineering and Applications Journal 8 (2019) 77–83. URL: https://comengapp.unsri.ac.id/index.php...
2019 doi
-
[23]
Firdaus, I
F. Firdaus, I. Fahreza, S. Nurmaini, A. Darmawahyuni, A. I. Sapitri, M. N. Rachmatullah, S. D. Lestari, M. Fachrurrozi, M. Afrina, B. W. Putra, Identification of Indonesian Authors Using Deep Neural Net- works, Computer Engineering and Applications Journal 11 (2022) 15–24. URL...
2022 doi
-
[24]
Q. Zhou, W. Chen, W. Wang, J. Xu, L. Zhao, Multiple features driven author name disambiguation, in: 2021 IEEE International Conference on Web Services (ICWS), 2021, pp. 506–515. doi:10.1109/ ICWS53863.2021.00071
2021
-
[25]
Q. Zhou, W. Chen, P.-P. Zhao, A. Liu, J.-J. Xu, J.-F. Qu, L. Zhao, Towards Effective Author Name Disambiguation by Hybrid Attention, Journal of Computer Science and Technology 39 (2024) 929–950. URL: https://doi.org/10.1007/s11390-023-2070-z. doi: 10.1007/s11390-023-2070-z
2024 doi
-
[26]
Q. Sun, H. Peng, J. Li, S. Wang, X. Dong, L. Zhao, P. S. Yu, L. He, Pairwise Learning for Name Disambiguation in Large-Scale Heterogeneous Academic Networks, in: 2020 IEEE In- ternational Conference on Data Mining (ICDM), IEEE, Sorrento, Italy, 2020, pp. 511–520. URL: https://...
2020
-
[27]
D. Choi, J. Jang, S. Song, H. Lee, J. Lim, K. Bok, J. Yoo, Name Disambiguation Scheme Based on Heterogeneous Academic Sites, Applied Sciences 14 (2024) 192. URL: https://www.mdpi.com/ 2076-3417/14/1/192. doi:10.3390/app14010192, number: 1 Publisher: Multidisciplinary Digital P...
2024 doi
-
[28]
Zhang, C
Z. Zhang, C. Wu, Z. Li, J. Peng, H. Wu, H. Song, S. Deng, B. Wang, Author Name Disambiguation Using Multiple Graph Attention Networks, in: 2021 International Joint Conference on Neural Net- works (IJCNN), IEEE, Shenzhen, China, 2021, pp. 1–8. URL: https://ieeexplore.ieee.org/d...
2021
-
[29]
Müller, F
M.-C. Müller, F. Reitz, N. Roy, Data sets for author name disambiguation: an empirical anal- ysis and a new resource, Scientometrics 111 (2017) 1467–1500. URL: https://doi.org/10.1007/ s11192-017-2363-5. doi: 10.1007/s11192-017-2363-5
2017 doi
-
[30]
Cheng, B
Y. Cheng, B. Chen, F. Zhang, J. Tang, BOND: Bootstrapping From-Scratch Name Disambiguation with Multi-task Promoting, in: Proceedings of the ACM Web Conference 2024, WWW ’24, Association for Computing Machinery, New York, NY, USA, 2024, pp. 4216–4226. URL: https: //dl.acm.org/...
2024
-
[31]
Xiong, P
B. Xiong, P. Bao, Y. Wu, Learning semantic and relationship joint embedding for author name disambiguation, Neural Computing and Applications 33 (2021) 1987–1998. URL: https://doi.org/10. 1007/s00521-020-05088-y. doi: 10.1007/s00521-020-05088-y
2021 doi
-
[32]
Z. Qiao, Y. Du, Y. Fu, P. Wang, Y. Zhou, Unsupervised Author Disambiguation using Heterogeneous Graph Convolutional Network Embedding, 2019, pp. 910–919. doi: 10.1109/BigData47090. 2019.9005458
2019
-
[33]
Pooja, S
K. Pooja, S. Mondal, J. Chandra, Exploiting Higher Order Multi-dimensional Relationships with Self-attention for Author Name Disambiguation, ACM Transactions on Knowledge Discovery from Data 16 (2022) 1–23. URL: https://dl.acm.org/doi/10.1145/3502730. doi:10.1145/3502730
2022 doi
-
[35]
Y. Ma, Y. Wu, C. Lu, A Graph-Based Author Name Disambiguation Method and Analysis via Information Theory, Entropy 22 (2020) 416. URL: https://www.mdpi.com/1099-4300/22/4/416. doi:10.3390/e22040416
2020 doi
-
[36]
W. Xie, S. Liu, X. Wang, T. Jia, Author Name Disambiguation via Heterogeneous Network Embedding from Structural and Semantic Perspectives, in: 2022 IEEE 34th International Conference on Tools with Artificial Intelligence (ICTAI), IEEE, Macao, China, 2022, pp. 245–250. URL: htt...
2022
-
[38]
Rettig, K
L. Rettig, K. Baumann, S. Sigloch, P. Cudre-Mauroux, Leveraging Knowledge Graph Embeddings to Disambiguate Author Names in Scientific Data, in: 2022 IEEE International Conference on Big Data (Big Data), IEEE, Osaka, Japan, 2022, pp. 5549–5557. URL: https://ieeexplore.ieee.org/...
2022
-
[381]
URL: https://api.semanticscholar.org/CorpusID:218586707
-
[3152]
doi:10.1109/TKDE.2020.3021256
URL: https://ieeexplore.ieee.org/document/9184992/. doi:10.1109/TKDE.2020.3021256
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.