refsplitr: Author name disambiguation, author georeferencing, and mapping of coauthorship networks with Web of Science data

Word sense induction (WSI), or the task of automatically discovering multiple senses or meanings of a word, has three main challenges: domain adaptability, novel sense detection, and sense granularity flexibility. While current latent variable models are known to solve the first two challenges, they are not flexible to different word sense granularities, which differ very much among words, from aardvark with one sense, to play with over 50 senses. Current models either require hyperparameter tuning or nonparametric induction of the number of senses, which we find both to be ineffective. Thus, we aim to eliminate these requirements and solve the sense granularity problem by proposing AutoSense, a latent variable model based on two observations: (1) senses are represented as a distribution over topics, and (2) senses generate pairings between the target word and its neighboring word. These observations alleviate the problem by (a) throwing garbage senses and (b) additionally inducing fine-grained word senses. Results show great improvements over the stateof-the-art models on popular WSI datasets. We also show that AutoSense is able to learn the appropriate sense granularity of a word. Finally, we apply AutoSense to the unsupervised author name disambiguation task where the sense granularity problem is more evident and show that AutoSense is evidently better than competing models. We share our data and code here: https://github.com/rktamplayo/AutoSense.

Download Full-text

Effect of Chinese characters on machine learning for Chinese author name disambiguation: A counterfactual evaluation

Journal of Information Science ◽

10.1177/01655515211018171 ◽

2021 ◽

pp. 016555152110181

Author(s):

Jinseok Kim ◽

Jenna Kim ◽

Jinmo Kim

Keyword(s):

Machine Learning ◽

Real World ◽

Digital Libraries ◽

Chinese Characters ◽

Name Disambiguation ◽

Authority Control ◽

Author Name Disambiguation ◽

Bibliographic Data ◽

Chinese Author

Chinese author names are known to be more difficult to disambiguate than other ethnic names because they tend to share surnames and forenames, thus creating many homonyms. In this study, we demonstrate how using Chinese characters can affect machine learning for author name disambiguation. For analysis, 15K author names recorded in Chinese are transliterated into English and simplified by initialising their forenames to create counterfactual scenarios, reflecting real-world indexing practices in which Chinese characters are usually unavailable. The results show that Chinese author names that are highly ambiguous in English or with initialised forenames tend to become less confusing if their Chinese characters are included in the processing. Our findings indicate that recording Chinese author names in native script can help researchers and digital libraries enhance authority control of Chinese author names that continue to increase in size in bibliographic data.

Download Full-text

LUCID: Author name disambiguation using graph Structural Clustering

2017 Intelligent Systems Conference (IntelliSys) ◽

10.1109/intellisys.2017.8324326 ◽

2017 ◽

Author(s):

Ijaz Hussain ◽

Sohail Asghar

Keyword(s):

Name Disambiguation ◽

Author Name Disambiguation ◽

Structural Clustering

Download Full-text

Correction to: Evaluating author name disambiguation for digital libraries: a case of DBLP

Scientometrics ◽

10.1007/s11192-018-2960-y ◽

2018 ◽

Vol 118 (1) ◽

pp. 383-383

Author(s):

Jinseok Kim

Keyword(s):

Digital Libraries ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Author name disambiguation for collaboration network analysis and visualization

Proceedings of the American Society for Information Science and Technology ◽

10.1002/meet.2009.1450460218 ◽

2009 ◽

Vol 46 (1) ◽

pp. 1-20 ◽

Cited By ~ 14

Author(s):

Andreas Strotmann ◽

Dangzhi Zhao ◽

Tania Bubela

Keyword(s):

Network Analysis ◽

Collaboration Network ◽

Name Disambiguation ◽

Author Name Disambiguation

Download Full-text

Off-the-shelf Semantic Author Name Disambiguation for Bibliographic Data Bases

Digital Libraries for Open Knowledge - Lecture Notes in Computer Science ◽

10.1007/978-3-030-30760-8_42 ◽

2019 ◽

pp. 397-400

Author(s):

Mark-Christoph Müller ◽

Adam Bannister ◽

Florian Reitz

Keyword(s):

Name Disambiguation ◽

Data Bases ◽

Author Name Disambiguation ◽

Bibliographic Data

Download Full-text

Key performance indicators to improve the competitive dimensions of the construction company

Research Society and Development ◽

10.33448/rsd-v9i5.3130 ◽

2020 ◽

Vol 9 (5) ◽

pp. e54953130

Author(s):

Aparecida Massako Tomioka ◽

José Manoel Souza das Neves

Keyword(s):

Organizational Performance ◽

Construction Industry ◽

Performance Indicators ◽

Research Method ◽

Web Of Science ◽

Qualitative Approach ◽

Review Of The Literature ◽

Science Data ◽

Data Bases ◽

Construction Company

The construction industry is a significant economic and productive sector of a country. Due to the importance of the sector, this study is justified not only for the academia, but also for the productive and business circles. Identifying competitive dimensions and comprehend the organizational performance through performance indicators, allows managers to make decisions through these tools, according to the model in which the organization operates, as close as possible to their reality. The present work aims to analyze the application of performance indicators through the competitive dimensions of the construction company. The used research method was a qualitative approach, being of an applied nature, classified according to the objectives of the research in descriptive and explanatory. The procedure used was the review of the literature through scientific articles in the Web of Science data bases, for the last ten years.

Download Full-text