An attention-based deep learning model for clinical named entity recognition of Chinese electronic medical records

Abstract Background Clinical named entity recognition (CNER) is important for medical information mining and establishment of high-quality knowledge map. Due to the different text features from natural language and a large number of professional and uncommon clinical terms in Chinese electronic medical records (EMRs), there are still many difficulties in clinical named entity recognition of Chinese EMRs. It is of great importance to eliminate semantic interference and improve the ability of autonomous learning of internal features of the model under the small training corpus. Methods From the perspective of deep learning, we integrated the attention mechanism into neural network, and proposed an improved clinical named entity recognition method for Chinese electronic medical records called BiLSTM-Att-CRF, which could capture more useful information of the context and avoid the problem of missing information caused by long-distance factors. In addition, medical dictionaries and part-of-speech (POS) features were also introduced to improve the performance of the model. Results Based on China Conference on Knowledge Graph and Semantic Computing (CCKS) 2017 and 2018 Chinese EMRs corpus, our BiLSTM-Att-CRF model finally achieved better performance than other widely-used models without additional features(F1-measure of 85.4% in CCKS 2018, F1-measure of 90.29% in CCKS 2017), and achieved the best performance with POS and dictionary features (F1-measure of 86.11% in CCKS 2018, F1-measure of 90.48% in CCKS 2017). In particular, the BiLSTM-Att-CRF model had significant effect on the improvement of Recall. Conclusions Our work preliminarily confirmed the validity of attention mechanism in discovering key information and mining text features, which might provide useful ideas for future research in clinical named entity recognition of Chinese electronic medical records. In the future, we will explore the deeper application of attention mechanism in neural network.

Download Full-text

Research on named entity recognition of chinese electronic medical records based on multi-head attention mechanism and character-word information fusion

Journal of Intelligent & Fuzzy Systems ◽

10.3233/jifs-212495 ◽

2021 ◽

pp. 1-12

Author(s):

Qinghui Zhang ◽

Meng Wu ◽

Pengtao Lv ◽

Mengya Zhang ◽

Hongwei Yang

Keyword(s):

Medical Record ◽

Electronic Medical Record ◽

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Attention Mechanism ◽

Entity Recognition ◽

Long Distance ◽

Named Entity ◽

The Impact

In the medical field, Named Entity Recognition (NER) plays a crucial role in the process of information extraction through electronic medical records and medical texts. To address the problems of long distance entity, entity confusion, and difficulty in boundary division in the Chinese electronic medical record NER task, we propose a Chinese electronic medical record NER method based on the multi-head attention mechanism and character-word fusion. This method uses a new character-word joint feature representation based on the pre-training model BERT and self-constructed domain dictionary, which can accurately divide the entity boundary and solve the impact of unregistered words. Subsequently, on the basis of the BiLSTM-CRF model, a multi-head attention mechanism is introduced to learn the dependency relationship between remote entities and entity information in different semantic spaces, which effectively improves the performance of the model. Experiments show that our models have better performance and achieves significant improvement compared to baselines. The specific performance is that the F1 value on the Chinese electronic medical record data set reaches 95.22%, which is 2.67%higher than the F1 value of the baseline model.

Download Full-text

A multiclass classification method based on deep learning for named entity recognition in electronic medical records

2016 New York Scientific Data Summit (NYSDS) ◽

10.1109/nysds.2016.7747810 ◽

2016 ◽

Cited By ~ 20

Author(s):

Xishuang Dong ◽

Lijun Qian ◽

Yi Guan ◽

Lei Huang ◽

Qiubin Yu ◽

...

Keyword(s):

Deep Learning ◽

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Multiclass Classification ◽

Entity Recognition ◽

Classification Method ◽

Named Entity

Download Full-text

Combined Attention Mechanism for Named Entity Recognition in Chinese Electronic Medical Records

2019 IEEE International Conference on Healthcare Informatics (ICHI) ◽

10.1109/ichi.2019.8904812 ◽

2019 ◽

Author(s):

Luqi Li ◽

Li Hou

Keyword(s):

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Attention Mechanism ◽

Entity Recognition ◽

Named Entity

Download Full-text

Clinical Named Entity Recognition from Chinese Electronic Medical Records Based on Deep Learning Pretraining

Journal of Healthcare Engineering ◽

10.1155/2020/8829219 ◽

2020 ◽

Vol 2020 ◽

pp. 1-8

Author(s):

Lejun Gong ◽

Zhifei Zhang ◽

Shiqi Chen

Keyword(s):

Deep Learning ◽

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Clinical Entity ◽

Fine Tuning ◽

Entity Recognition ◽

Recognition Model ◽

Named Entity ◽

Model Based

Background. Clinical named entity recognition is the basic task of mining electronic medical records text, which are with some challenges containing the language features of Chinese electronic medical records text with many compound entities, serious missing sentence components, and unclear entity boundary. Moreover, the corpus of Chinese electronic medical records is difficult to obtain. Methods. Aiming at these characteristics of Chinese electronic medical records, this study proposed a Chinese clinical entity recognition model based on deep learning pretraining. The model used word embedding from domain corpus and fine-tuning of entity recognition model pretrained by relevant corpus. Then BiLSTM and Transformer are, respectively, used as feature extractors to identify four types of clinical entities including diseases, symptoms, drugs, and operations from the text of Chinese electronic medical records. Results. 75.06% Macro-P, 76.40% Macro-R, and 75.72% Macro-F1 aiming at test dataset could be achieved. These experiments show that the Chinese clinical entity recognition model based on deep learning pretraining can effectively improve the recognition effect. Conclusions. These experiments show that the proposed Chinese clinical entity recognition model based on deep learning pretraining can effectively improve the recognition performance.

Download Full-text

Deep learning for named entity recognition on Chinese electronic medical records: Combining deep transfer learning with multitask bi-directional LSTM RNN

PLoS ONE ◽

10.1371/journal.pone.0216046 ◽

2019 ◽

Vol 14 (5) ◽

pp. e0216046 ◽

Cited By ~ 5

Author(s):

Xishuang Dong ◽

Shanta Chowdhury ◽

Lijun Qian ◽

Xiangfang Li ◽

Yi Guan ◽

...

Keyword(s):

Deep Learning ◽

Electronic Medical Records ◽

Transfer Learning ◽

Medical Records ◽

Named Entity Recognition ◽

Entity Recognition ◽

Named Entity

Download Full-text

A Hybrid Model for Named Entity Recognition on Chinese Electronic Medical Records

ACM Transactions on Asian and Low-Resource Language Information Processing ◽

10.1145/3436819 ◽

2021 ◽

Vol 20 (2) ◽

pp. 1-12

Author(s):

Yu Wang ◽

Yining Sun ◽

Zuchang Ma ◽

Lisheng Gao ◽

Yang Xu

Keyword(s):

Neural Network ◽

Electronic Medical Records ◽

Hybrid Model ◽

Medical Records ◽

Medical Information ◽

Short Term Memory ◽

Clinical Symptoms ◽

Named Entity Recognition ◽

Entity Recognition ◽

Named Entity

Electronic medical records (EMRs) contain valuable information about the patients, such as clinical symptoms, diagnostic results, and medications. Named entity recognition (NER) aims to recognize entities from unstructured text, which is the initial step toward the semantic understanding of the EMRs. Extracting medical information from Chinese EMRs could be a more complicated task because of the difference between English and Chinese. Some researchers have noticed the importance of Chinese NER and used the recurrent neural network or convolutional neural network (CNN) to deal with this task. However, it is interesting to know whether the performance could be improved if the advantages of the RNN and CNN can be both utilized. Moreover, RoBERTa-WWM, as a pre-training model, can generate the embeddings with word-level features, which is more suitable for Chinese NER compared with Word2Vec. In this article, we propose a hybrid model. This model first obtains the entities identified by bidirectional long short-term memory and CNN, respectively, and then uses two hybrid strategies to output the final results relying on these entities. We also conduct experiments on raw medical records from real hospitals. This dataset is provided by the China Conference on Knowledge Graph and Semantic Computing in 2019 (CCKS 2019). Results demonstrate that the hybrid model can improve performance significantly.

Download Full-text

A deep learning model incorporating part of speech and self-matching attention for named entity recognition of Chinese electronic medical records

BMC Medical Informatics and Decision Making ◽

10.1186/s12911-019-0762-7 ◽

2019 ◽

Vol 19 (S2) ◽

Cited By ~ 6

Author(s):

Xiaoling Cai ◽

Shoubin Dong ◽

Jinlong Hu

Keyword(s):

Deep Learning ◽

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Learning Model ◽

Entity Recognition ◽

Named Entity ◽

Part Of Speech ◽

Deep Learning Model

Download Full-text

Overview of CCKS 2020 Task 3: Named Entity Recognition and Event Extraction in Chinese Electronic Medical Records

Data Intelligence ◽

10.1162/dint_a_00093 ◽

2021 ◽

pp. 1-13

Author(s):

Xia Li ◽

Qinghua Wen ◽

Zengtao Jiao ◽

Jiangtao Zhang

Keyword(s):

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Event Extraction ◽

Entity Recognition ◽

Language Models ◽

Data Sets ◽

External Resources ◽

Named Entity ◽

Evaluation Task

Abstract The China Conference on Knowledge Graph and Semantic Computing (CCKS) 2020 Evaluation Task 3 presented clinical named entity recognition and event extraction for the Chinese electronic medical records. Two annotated data sets and some other additional resources for these two subtasks were provided for participators. This evaluation competition attracted 354 teams and 46 of them successfully submitted the valid results. The pre-trained language models are widely applied in this evaluation task. Data argumentation and external resources are also helpful.

Download Full-text

Information Extraction from Electronic Medical Records Using Multitask Recurrent Neural Network with Contextual Word Embedding

Applied Sciences ◽

10.3390/app9183658 ◽

2019 ◽

Vol 9 (18) ◽

pp. 3658 ◽

Cited By ~ 6

Author(s):

Jianliang Yang ◽

Yuenan Liu ◽

Minghui Qian ◽

Chenghua Guan ◽

Xiangfei Yuan

Keyword(s):

Electronic Medical Records ◽

Medical Records ◽

Large Scale ◽

Short Term Memory ◽

Conditional Random Field ◽

Named Entity Recognition ◽

Recognition Task ◽

Entity Recognition ◽

Language Models ◽

Named Entity

Clinical named entity recognition is an essential task for humans to analyze large-scale electronic medical records efficiently. Traditional rule-based solutions need considerable human effort to build rules and dictionaries; machine learning-based solutions need laborious feature engineering. For the moment, deep learning solutions like Long Short-term Memory with Conditional Random Field (LSTM–CRF) achieved considerable performance in many datasets. In this paper, we developed a multitask attention-based bidirectional LSTM–CRF (Att-biLSTM–CRF) model with pretrained Embeddings from Language Models (ELMo) in order to achieve better performance. In the multitask system, an additional task named entity discovery was designed to enhance the model’s perception of unknown entities. Experiments were conducted on the 2010 Informatics for Integrating Biology & the Bedside/Veterans Affairs (I2B2/VA) dataset. Experimental results show that our model outperforms the state-of-the-art solution both on the single model and ensemble model. Our work proposes an approach to improve the recall in the clinical named entity recognition task based on the multitask mechanism.

Download Full-text

Overview of CCKS 2018 Task 1: Named Entity Recognition in Chinese Electronic Medical Records

Communications in Computer and Information Science - Knowledge Graph and Semantic Computing: Knowledge Computing and Language Understanding ◽

10.1007/978-981-15-1956-7_14 ◽

2019 ◽

pp. 158-164

Author(s):

Jiangtao Zhang ◽

Juanzi Li ◽

Zengtao Jiao ◽

Jun Yan

Keyword(s):

Electronic Medical Records ◽

Medical Records ◽

Named Entity Recognition ◽

Entity Recognition ◽

Named Entity

Download Full-text