A novel machine-learning approach to measuring scientific knowledge flows using citation context analysis

Abstract Scientific discoveries do not occur in vacuum but rather by connecting existing pieces of knowledge in new and creative ways. Mapping the relation and structure of scientific knowledge is therefore central to our understanding of the dynamics of scientific production. Here we introduce a new approach to generate scientific knowledge maps based on a machine learning approach that, starting from the observed publication patterns of authors, generates an N-dimensional space where it is possible to measure the similarity or distance between different research topics and knowledge domains. We provide an implementation of the proposed approach that considers the American Physical Society publications database and generates a map of the research space in Physics that characterizes the relation among research topics over time. We use this map to measure two indicators, the research capacity fingerprint and the knowledge density, to profile the research activity in physical sciences of more than 400 urban areas across the world. We show that these indicators can be used to analyze and predict the evolution over time of the research capacity and specialization of specific geographical areas. Furthermore we provide an extensive analysis of the relation between socio-economic development indicators and the ability to produce new knowledge for 67 countries, as measured by our approach, highlighting some key correlates of scientific production capacity. The proposed approach is scalable to very large datasets and can be extended to study other disciplines and research areas without having to rely on ad-hoc science classification schemes.

Download Full-text

CONTEXT ANALYSIS FOR SEMANTIC MAPPING OF DATA SOURCES USING A MULTI-STRATEGY MACHINE LEARNING APPROACH

Proceedings of the Seventh International Conference on Enterprise Information Systems ◽

10.5220/0002539804450448 ◽

2005 ◽

Keyword(s):

Machine Learning ◽

Data Sources ◽

Semantic Mapping ◽

Learning Approach ◽

Context Analysis ◽

Machine Learning Approach

Download Full-text

Constructing and Validating Geographically Refined HAZUS-MH4 Hurricane Wind Risk Models: A Machine Learning Approach

Advances in Hurricane Engineering ◽

10.1061/9780784412626.092 ◽

2012 ◽

Cited By ~ 2

Author(s):

D. Subramanian ◽

J. Salazar ◽

L. Duenas-Osorio ◽

R. Stein

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Risk Models ◽

Hurricane Wind ◽

Machine Learning Approach

Download Full-text

The impact of economic plans on the Chinese education system: a machine learning approach

CADMO ◽

10.3280/cad2018-001005 ◽

2018 ◽

pp. 37-49

Author(s):

Wenjun Lin ◽

Xuefu Xu ◽

Francesco Dell’Anna

Keyword(s):

Machine Learning ◽

Education System ◽

Learning Approach ◽

Chinese Education ◽

System A ◽

Machine Learning Approach ◽

The Impact

Download Full-text

Improving Bandwidth Utilization and Fairness between TCP Flows based on a Machine-learning Approach

IEEJ Transactions on Electronics Information and Systems ◽

10.1541/ieejeiss.133.1259 ◽

2013 ◽

Vol 133 (6) ◽

pp. 1259-1268

Author(s):

Akihiro Shiozu ◽

Syunji Yazaki ◽

K^|^ocirc;ki Abe

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Bandwidth Utilization ◽

Machine Learning Approach

Download Full-text

A Machine Learning Approach to Anaphora Resolution in Arabic

International Review on Computers and Software (IRECOS) ◽

10.15866/irecos.v9i12.4786 ◽

2014 ◽

Vol 9 (12) ◽

pp. 1956

Author(s):

Abdullatif Abolohom ◽

Nazlia Omar

Keyword(s):

Machine Learning ◽

Learning Approach ◽

Anaphora Resolution ◽

Machine Learning Approach

Download Full-text

1552-P: Machine Learning Approach to Decision-Making for Initial Insulin Use in Japanese Patients with Type 2 Diabetes

Diabetes ◽

10.2337/db20-1552-p ◽

2020 ◽

Vol 69 (Supplement 1) ◽

pp. 1552-P

Author(s):

KAZUYA FUJIHARA ◽

MAYUKO H. YAMADA ◽

YASUHIRO MATSUBAYASHI ◽

MASAHIKO YAMAMOTO ◽

TOSHIHIRO IIZUKA ◽

...

Keyword(s):

Machine Learning ◽

Type 2 Diabetes ◽

Decision Making ◽

Japanese Patients ◽

Learning Approach ◽

Machine Learning Approach ◽

Insulin Use

Download Full-text

A Machine Learning Approach to Jet-Surface Interaction Noise Modeling

AIAA Scitech 2020 Forum ◽

10.2514/6.2020-1728 ◽

2020 ◽

Author(s):

Clifford A. Brown ◽

Jonny Dowdall ◽

Brian Whiteaker ◽

Lauren McIntyre

Keyword(s):

Machine Learning ◽

Surface Interaction ◽

Learning Approach ◽

Noise Modeling ◽

Machine Learning Approach

Download Full-text

A machine learning approach to evaluate IgE and IgG4 responses as patient-specific measure of exposure to sublingual allergen immunotherapy

10.26226/morressier.5afda3c8d64f25002cfc4098 ◽

2018 ◽

Author(s):

Thomas Stranzl

Keyword(s):

Machine Learning ◽

Allergen Immunotherapy ◽

Patient Specific ◽

Learning Approach ◽

Specific Measure ◽

Machine Learning Approach

Download Full-text

Mol2vec: Unsupervised Machine Learning Approach with Chemical Intuition

10.26434/chemrxiv.5513581.v1 ◽

2017 ◽

Author(s):

Sabrina Jaeger ◽

Simone Fulle ◽

Samo Turk

Keyword(s):

Machine Learning ◽

Language Processing ◽

Supervised Machine Learning ◽

Learning Approach ◽

Learning Approaches ◽

Unsupervised Machine Learning ◽

Feature Representations ◽

Machine Learning Approach ◽

The Individual ◽

Vector Representations

Inspired by natural language processing techniques we here introduce Mol2vec which is an unsupervised machine learning approach to learn vector representations of molecular substructures. Similarly, to the Word2vec models where vectors of closely related words are in close proximity in the vector space, Mol2vec learns vector representations of molecular substructures that are pointing in similar directions for chemically related substructures. Compounds can finally be encoded as vectors by summing up vectors of the individual substructures and, for instance, feed into supervised machine learning approaches to predict compound properties. The underlying substructure vector embeddings are obtained by training an unsupervised machine learning approach on a so-called corpus of compounds that consists of all available chemical matter. The resulting Mol2vec model is pre-trained once, yields dense vector representations and overcomes drawbacks of common compound feature representations such as sparseness and bit collisions. The prediction capabilities are demonstrated on several compound property and bioactivity data sets and compared with results obtained for Morgan fingerprints as reference compound representation. Mol2vec can be easily combined with ProtVec, which employs the same Word2vec concept on protein sequences, resulting in a proteochemometric approach that is alignment independent and can be thus also easily used for proteins with low sequence similarities.

Download Full-text