Completing features for author name disambiguation (AND): an empirical analysis

Data sets for author name disambiguation: an empirical analysis and a new resource

Scientometrics ◽

10.1007/s11192-017-2363-5 ◽

2017 ◽

Vol 111 (3) ◽

pp. 1467-1500 ◽

Cited By ~ 19

Author(s):

Mark-Christoph Müller ◽

Florian Reitz ◽

Nicolas Roy

Keyword(s):

Empirical Analysis ◽

Data Sets ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Ethnicity‐based name partitioning for author name disambiguation using supervised machine learning

Journal of the Association for Information Science and Technology ◽

10.1002/asi.24459 ◽

2021 ◽

Author(s):

Jinseok Kim ◽

Jenna Kim ◽

Jason Owen‐Smith

Keyword(s):

Machine Learning ◽

Supervised Machine Learning ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Multilayer heuristics based clustering framework (MHCF) for author name disambiguation

Scientometrics ◽

10.1007/s11192-021-04087-7 ◽

2021 ◽

Author(s):

Humaira Waqas ◽

Muhammad Abdul Qadir

Keyword(s):

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

AutoSense Model for Word Sense Induction

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v33i01.33016212 ◽

2019 ◽

Vol 33 ◽

pp. 6212-6219 ◽

Cited By ~ 1

Author(s):

Reinald Kim Amplayo ◽

Seung-won Hwang ◽

Min Song

Keyword(s):

Latent Variable ◽

Word Sense ◽

Name Disambiguation ◽

Variable Model ◽

Fine Grained ◽

Word Sense Induction ◽

Author Name Disambiguation ◽

Competing Models ◽

Word Senses ◽

Better Than

Word sense induction (WSI), or the task of automatically discovering multiple senses or meanings of a word, has three main challenges: domain adaptability, novel sense detection, and sense granularity flexibility. While current latent variable models are known to solve the first two challenges, they are not flexible to different word sense granularities, which differ very much among words, from aardvark with one sense, to play with over 50 senses. Current models either require hyperparameter tuning or nonparametric induction of the number of senses, which we find both to be ineffective. Thus, we aim to eliminate these requirements and solve the sense granularity problem by proposing AutoSense, a latent variable model based on two observations: (1) senses are represented as a distribution over topics, and (2) senses generate pairings between the target word and its neighboring word. These observations alleviate the problem by (a) throwing garbage senses and (b) additionally inducing fine-grained word senses. Results show great improvements over the stateof-the-art models on popular WSI datasets. We also show that AutoSense is able to learn the appropriate sense granularity of a word. Finally, we apply AutoSense to the unsupervised author name disambiguation task where the sense granularity problem is more evident and show that AutoSense is evidently better than competing models. We share our data and code here: https://github.com/rktamplayo/AutoSense.

Download Full-text

Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation

Journal of Information Science ◽

10.1177/01655515211018171 ◽

2021 ◽

pp. 016555152110181

Author(s):

Jinseok Kim ◽

Jenna Kim ◽

Jinmo Kim

Keyword(s):

Machine Learning ◽

Real World ◽

Digital Libraries ◽

Chinese Characters ◽

Name Disambiguation ◽

Authority Control ◽

Author Name Disambiguation ◽

Bibliographic Data ◽

Chinese Author

Chinese author names are known to be more difficult to disambiguate than other ethnic names because they tend to share surnames and forenames, thus creating many homonyms. In this study, we demonstrate how using Chinese characters can affect machine learning for author name disambiguation. For analysis, 15K author names recorded in Chinese are transliterated into English and simplified by initialising their forenames to create counterfactual scenarios, reflecting real-world indexing practices in which Chinese characters are usually unavailable. The results show that Chinese author names that are highly ambiguous in English or with initialised forenames tend to become less confusing if their Chinese characters are included in the processing. Our findings indicate that recording Chinese author names in native script can help researchers and digital libraries enhance authority control of Chinese author names that continue to increase in size in bibliographic data.

Download Full-text

LUCID: Author name disambiguation using graph Structural Clustering

2017 Intelligent Systems Conference (IntelliSys) ◽

10.1109/intellisys.2017.8324326 ◽

2017 ◽

Author(s):

Ijaz Hussain ◽

Sohail Asghar

Keyword(s):

Name Disambiguation ◽

Author Name Disambiguation ◽

Structural Clustering

Download Full-text

Correction to: Evaluating author name disambiguation for digital libraries: a case of DBLP

Scientometrics ◽

10.1007/s11192-018-2960-y ◽

2018 ◽

Vol 118 (1) ◽

pp. 383-383

Author(s):

Jinseok Kim

Keyword(s):

Digital Libraries ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Author name disambiguation for collaboration network analysis and visualization

Proceedings of the American Society for Information Science and Technology ◽

10.1002/meet.2009.1450460218 ◽

2009 ◽

Vol 46 (1) ◽

pp. 1-20 ◽

Cited By ~ 14

Author(s):

Andreas Strotmann ◽

Dangzhi Zhao ◽

Tania Bubela

Keyword(s):

Network Analysis ◽

Collaboration Network ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Off-the-shelf Semantic Author Name Disambiguation for Bibliographic Data Bases

Digital Libraries for Open Knowledge - Lecture Notes in Computer Science ◽

10.1007/978-3-030-30760-8_42 ◽

2019 ◽

pp. 397-400

Author(s):

Mark-Christoph Müller ◽

Adam Bannister ◽

Florian Reitz

Keyword(s):

Name Disambiguation ◽

Data Bases ◽

Author Name Disambiguation ◽

Bibliographic Data

Download Full-text

Author name disambiguation of bibliometric data: A comparison of several unsupervised approaches

Quantitative Science Studies ◽

10.1162/qss_a_00081 ◽

2020 ◽

Vol 1 (4) ◽

pp. 1510-1528

Author(s):

Alexander Tekles ◽

Lutz Bornmann

Keyword(s):

A Priori ◽

Bibliometric Data ◽

Name Disambiguation ◽

Controlled Conditions ◽

Author Name Disambiguation ◽

Overall Performance ◽

Bibliometric Databases ◽

Bibliometric Studies ◽

Unsupervised Approaches

Adequately disambiguating author names in bibliometric databases is a precondition for conducting reliable analyses at the author level. In the case of bibliometric studies that include many researchers, it is not possible to disambiguate each single researcher manually. Several approaches have been proposed for author name disambiguation, but there has not yet been a comparison of them under controlled conditions. In this study, we compare a set of unsupervised disambiguation approaches. Unsupervised approaches specify a model to assess the similarity of author mentions a priori instead of training a model with labeled data. To evaluate the approaches, we applied them to a set of author mentions annotated with a ResearcherID, this being an author identifier maintained by the researchers themselves. Apart from comparing the overall performance, we take a more detailed look at the role of the parametrization of the approaches and analyze the dependence of the results on the complexity of the disambiguation task. Furthermore, we examine which effects the differences in the set of metadata considered by the different approaches have on the disambiguation results. In the context of this study, the approach proposed by Caron and van Eck (2014) produced the best results.

Download Full-text