Named Entity Recognition and Relation Extraction for COVID-19: Explainable Active Learning with Word2vec Embeddings and Transformer-Based BERT Models

AHIAP: An Agile Medical Named Entity Recognition and Relation Extraction Framework Based on Active Learning

Health Information Science - Lecture Notes in Computer Science ◽

10.1007/978-3-030-61951-0_7 ◽

2020 ◽

pp. 68-75

Author(s):

Ming Sheng ◽

Jing Dong ◽

Yong Zhang ◽

Yuelin Bu ◽

Anqi Li ◽

...

Keyword(s):

Active Learning ◽

Named Entity Recognition ◽

Relation Extraction ◽

Entity Recognition ◽

Named Entity

Download Full-text

Text Mining of Hazard and Operability Analysis Reports Based on Active Learning

Processes ◽

10.3390/pr9071178 ◽

2021 ◽

Vol 9 (7) ◽

pp. 1178

Author(s):

Zhenhua Wang ◽

Beike Zhang ◽

Dong Gao

Keyword(s):

Active Learning ◽

Named Entity Recognition ◽

Entity Recognition ◽

Chemical System ◽

Chemical Safety ◽

High Quality ◽

Data Set ◽

Final Model ◽

Named Entity ◽

Sampling Algorithms

In the field of chemical safety, a named entity recognition (NER) model based on deep learning can mine valuable information from hazard and operability analysis (HAZOP) text, which can guide experts to carry out a new round of HAZOP analysis, help practitioners optimize the hidden dangers in the system, and be of great significance to improve the safety of the whole chemical system. However, due to the standardization and professionalism of chemical safety analysis text, it is difficult to improve the performance of traditional models. To solve this problem, in this study, an improved method based on active learning is proposed, and three novel sampling algorithms are designed, Variation of Token Entropy (VTE), HAZOP Confusion Entropy (HCE) and Amplification of Least Confidence (ALC), which improve the ability of the model to understand HAZOP text. In this method, a part of data is used to establish the initial model. The sampling algorithm is then used to select high-quality samples from the data set. Finally, these high-quality samples are used to retrain the whole model to obtain the final model. The experimental results show that the performance of the VTE, HCE, and ALC algorithms are better than that of random sampling algorithms. In addition, compared with other methods, the performance of the traditional model is improved effectively by the method proposed in this paper, which proves that the method is reliable and advanced.

Download Full-text

Named Entity Recognition and Relation Extraction with Graph Neural Networks in Semi Structured Documents

2020 25th International Conference on Pattern Recognition (ICPR) ◽

10.1109/icpr48806.2021.9412669 ◽

2021 ◽

Author(s):

Manuel Carbonell ◽

Pau Riba ◽

Mauricio Villegas ◽

Alicia Fornes ◽

Josep Llados

Keyword(s):

Neural Networks ◽

Named Entity Recognition ◽

Relation Extraction ◽

Entity Recognition ◽

Named Entity ◽

Structured Documents ◽

Graph Neural Networks

Download Full-text

Named Entity Recognition and Relation Extraction

ACM Computing Surveys ◽

10.1145/3445965 ◽

2021 ◽

Vol 54 (1) ◽

pp. 1-39

Author(s):

Zara Nasar ◽

Syed Waqar Jaffry ◽

Muhammad Kamran Malik

Keyword(s):

Deep Learning ◽

State Of The Art ◽

Named Entity Recognition ◽

Relation Extraction ◽

The State ◽

Entity Recognition ◽

Joint Models ◽

Named Entity ◽

Textual Data ◽

Benchmark Datasets

With the advent of Web 2.0, there exist many online platforms that result in massive textual-data production. With ever-increasing textual data at hand, it is of immense importance to extract information nuggets from this data. One approach towards effective harnessing of this unstructured textual data could be its transformation into structured text. Hence, this study aims to present an overview of approaches that can be applied to extract key insights from textual data in a structured way. For this, Named Entity Recognition and Relation Extraction are being majorly addressed in this review study. The former deals with identification of named entities, and the latter deals with problem of extracting relation between set of entities. This study covers early approaches as well as the developments made up till now using machine learning models. Survey findings conclude that deep-learning-based hybrid and joint models are currently governing the state-of-the-art. It is also observed that annotated benchmark datasets for various textual-data generators such as Twitter and other social forums are not available. This scarcity of dataset has resulted into relatively less progress in these domains. Additionally, the majority of the state-of-the-art techniques are offline and computationally expensive. Last, with increasing focus on deep-learning frameworks, there is need to understand and explain the under-going processes in deep architectures.

Download Full-text

An Attention-Based Model Using Character Composition of Entities in Chinese Relation Extraction

Information ◽

10.3390/info11020079 ◽

2020 ◽

Vol 11 (2) ◽

pp. 79 ◽

Cited By ~ 2

Author(s):

Xiaoyu Han ◽

Yue Zhang ◽

Wenkai Zhang ◽

Tinglei Huang

Keyword(s):

Language Processing ◽

Large Scale ◽

Named Entity Recognition ◽

Relation Extraction ◽

Entity Recognition ◽

Additional Information ◽

Named Entity ◽

Proposed Model ◽

The Relationship ◽

Crucial Part

Relation extraction is a vital task in natural language processing. It aims to identify the relationship between two specified entities in a sentence. Besides information contained in the sentence, additional information about the entities is verified to be helpful in relation extraction. Additional information such as entity type getting by NER (Named Entity Recognition) and description provided by knowledge base both have their limitations. Nevertheless, there exists another way to provide additional information which can overcome these limitations in Chinese relation extraction. As Chinese characters usually have explicit meanings and can carry more information than English letters. We suggest that characters that constitute the entities can provide additional information which is helpful for the relation extraction task, especially in large scale datasets. This assumption has never been verified before. The main obstacle is the lack of large-scale Chinese relation datasets. In this paper, first, we generate a large scale Chinese relation extraction dataset based on a Chinese encyclopedia. Second, we propose an attention-based model using the characters that compose the entities. The result on the generated dataset shows that these characters can provide useful information for the Chinese relation extraction task. By using this information, the attention mechanism we used can recognize the crucial part of the sentence that can express the relation. The proposed model outperforms other baseline models on our Chinese relation extraction dataset.

Download Full-text

Active learning for ontological event extraction incorporating named entity recognition and unknown word handling

Journal of Biomedical Semantics ◽

10.1186/s13326-016-0059-z ◽

2016 ◽

Vol 7 (1) ◽

Cited By ~ 2

Author(s):

Xu Han ◽

Jung-jae Kim ◽

Chee Keong Kwoh

Keyword(s):

Active Learning ◽

Named Entity Recognition ◽

Event Extraction ◽

Entity Recognition ◽

Unknown Word ◽

Named Entity

Download Full-text

LTP: A New Active Learning Strategy for CRF-Based Named Entity Recognition

Neural Processing Letters ◽

10.1007/s11063-021-10737-x ◽

2022 ◽

Author(s):

Mingyi Liu ◽

Zhiying Tu ◽

Tong Zhang ◽

Tonghua Su ◽

Xiaofei Xu ◽

...

Keyword(s):

Active Learning ◽

Learning Strategy ◽

Named Entity Recognition ◽

Entity Recognition ◽

Named Entity ◽

Active Learning Strategy

Download Full-text

Fine-Grained Mechanical Chinese Named Entity Recognition Based on ALBERT-AttBiLSTM-CRF and Transfer Learning

Symmetry ◽

10.3390/sym12121986 ◽

2020 ◽

Vol 12 (12) ◽

pp. 1986

Author(s):

Liguo Yao ◽

Haisong Huang ◽

Kuan-Wei Wang ◽

Shih-Huan Chen ◽

Qiaoqiao Xiong

Keyword(s):

Active Learning ◽

Manufacturing Industry ◽

Learning Strategy ◽

Named Entity Recognition ◽

Entity Recognition ◽

Utilization Rate ◽

Data Types ◽

Fine Grained ◽

Named Entity ◽

Model Transfer

Manufacturing text often exists as unlabeled data; the entity is fine-grained and the extraction is difficult. The above problems mean that the manufacturing industry knowledge utilization rate is low. This paper proposes a novel Chinese fine-grained NER (named entity recognition) method based on symmetry lightweight deep multinetwork collaboration (ALBERT-AttBiLSTM-CRF) and model transfer considering active learning (MTAL) to research fine-grained named entity recognition of a few labeled Chinese textual data types. The method is divided into two stages. In the first stage, the ALBERT-AttBiLSTM-CRF was applied for verification in the CLUENER2020 dataset (Public dataset) to get a pretrained model; the experiments show that the model obtains an F1 score of 0.8962, which is better than the best baseline algorithm, an improvement of 9.2%. In the second stage, the pretrained model was transferred into the Manufacturing-NER dataset (our dataset), and we used the active learning strategy to optimize the model effect. The final F1 result of Manufacturing-NER was 0.8931 after the model transfer (it was higher than 0.8576 before the model transfer); so, this method represents an improvement of 3.55%. Our method effectively transfers the existing knowledge from public source data to scientific target data, solving the problem of named entity recognition with scarce labeled domain data, and proves its effectiveness.

Download Full-text

Using error decay prediction to overcome practical issues of deep active learning for named entity recognition

Machine Learning ◽

10.1007/s10994-020-05897-1 ◽

2020 ◽

Vol 109 (9-10) ◽

pp. 1749-1778

Author(s):

Haw-Shiuan Chang ◽

Shankar Vembu ◽

Sunil Mohan ◽

Rheeya Uppaal ◽

Andrew McCallum

Keyword(s):

Active Learning ◽

Named Entity Recognition ◽

Entity Recognition ◽

Named Entity

Download Full-text

A Low-Cost Named Entity Recognition Research Based on Active Learning

Scientific Programming ◽

10.1155/2018/1890683 ◽

2018 ◽

Vol 2018 ◽

pp. 1-10 ◽

Cited By ~ 1

Author(s):

Han Huang ◽

Hongyu Wang ◽

Dawei Jin

Keyword(s):

Active Learning ◽

Language Processing ◽

Selection Process ◽

Conditional Random Field ◽

Low Cost ◽

Named Entity Recognition ◽

Entity Recognition ◽

Training Set ◽

Processing Technologies ◽

Named Entity

Named entity recognition (NER) is an indispensable and very important part of many natural language processing technologies, such as information extraction, information retrieval, and intelligent Q & A. This paper describes the development of the AL-CRF model, which is a NER approach based on active learning (AL). The algorithmic sequence of the processes performed by the AL-CRF model is the following: first, the samples are clustered using the k-means approach. Then, stratified sampling is performed on the produced clusters in order to obtain initial samples, which are used to train the basic conditional random field (CRF) classifier. The next step includes the initiation of the selection process which uses the criterion of entropy. More specifically, samples having the highest entropy values are added to the training set. Afterwards, the learning process is repeated, and the CRF classifier is retrained based on the obtained training set. The learning and the selection process of the AL is running iteratively until the harmonic mean F stabilizes and the final NER model is obtained. Several NER experiments are performed on legislative and medical cases in order to validate the AL-CRF performance. The testing data include Chinese judicial documents and Chinese electronic medical records (EMRs). Testing indicates that our proposed algorithm has better recognition accuracy and recall rate compared to the conventional CRF model. Moreover, the main advantage of our approach is that it requires fewer manually labelled training samples, and at the same time, it is more effective. This can result in a more cost effective and more reliable process.

Download Full-text