Scalable Attentive Sentence Pair Modeling via Distilled Sentence Embedding

Recent state-of-the-art natural language understanding models, such as BERT and XLNet, score a pair of sentences (A and B) using multiple cross-attention operations – a process in which each word in sentence A attends to all words in sentence B and vice versa. As a result, computing the similarity between a query sentence and a set of candidate sentences, requires the propagation of all query-candidate sentence-pairs throughout a stack of cross-attention layers. This exhaustive process becomes computationally prohibitive when the number of candidate sentences is large. In contrast, sentence embedding techniques learn a sentence-to-vector mapping and compute the similarity between the sentence vectors via simple elementary operations. In this paper, we introduce Distilled Sentence Embedding (DSE) – a model that is based on knowledge distillation from cross-attentive models, focusing on sentence-pair tasks. The outline of DSE is as follows: Given a cross-attentive teacher model (e.g. a fine-tuned BERT), we train a sentence embedding based student model to reconstruct the sentence-pair scores obtained by the teacher model. We empirically demonstrate the effectiveness of DSE on five GLUE sentence-pair tasks. DSE significantly outperforms several ELMO variants and other sentence embedding methods, while accelerating computation of the query-candidate sentence-pairs similarities by several orders of magnitude, with an average relative degradation of 4.6% compared to BERT. Furthermore, we show that DSE produces sentence embeddings that reach state-of-the-art performance on universal sentence representation benchmarks. Our code is made publicly available at https://github.com/microsoft/Distilled-Sentence-Embedding.

Download Full-text

Deep Cascade Multi-Task Learning for Slot Filling in Online Shopping Assistant

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v33i01.33016465 ◽

2019 ◽

Vol 33 ◽

pp. 6465-6472 ◽

Cited By ~ 3

Author(s):

Yu Gong ◽

Xusheng Luo ◽

Yu Zhu ◽

Wenwu Ou ◽

Zhao Li ◽

...

Keyword(s):

Natural Language ◽

Knowledge Base ◽

Online Shopping ◽

State Of The Art ◽

Language Understanding ◽

Dialog Systems ◽

Named Entity ◽

Online Test ◽

Benchmark Datasets ◽

Slot Filling

Slot filling is a critical task in natural language understanding (NLU) for dialog systems. State-of-the-art approaches treat it as a sequence labeling problem and adopt such models as BiLSTM-CRF. While these models work relatively well on standard benchmark datasets, they face challenges in the context of E-commerce where the slot labels are more informative and carry richer expressions. In this work, inspired by the unique structure of E-commerce knowledge base, we propose a novel multi-task model with cascade and residual connections, which jointly learns segment tagging, named entity tagging and slot filling. Experiments show the effectiveness of the proposed cascade and residual structures. Our model has a 14.6% advantage in F1 score over the strong baseline methods on a new Chinese E-commerce shopping assistant dataset, while achieving competitive accuracies on a standard dataset. Furthermore, online test deployed on such dominant E-commerce platform shows 130% improvement on accuracy of understanding user utterances. Our model has already gone into production in the E-commerce platform.

Download Full-text

How to Select One Among All ? An Empirical Study Towards the Robustness of Knowledge Distillation in Natural Language Understanding

10.18653/v1/2021.findings-emnlp.65 ◽

2021 ◽

Author(s):

Tianda Li ◽

Ahmad Rashid ◽

Aref Jafari ◽

Pranav Sharma ◽

Ali Ghodsi ◽

...

Keyword(s):

Empirical Study ◽

Natural Language ◽

Natural Language Understanding ◽

Language Understanding ◽

Knowledge Distillation

Download Full-text

Probing Natural Language Inference Models through Semantic Fragments

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v34i05.6397 ◽

2020 ◽

Vol 34 (05) ◽

pp. 8713-8721

Author(s):

Kyle Richardson ◽

Hai Hu ◽

Lawrence Moss ◽

Ashish Sabharwal

Keyword(s):

Natural Language ◽

State Of The Art ◽

Fine Tuning ◽

Language Understanding ◽

Inference Models ◽

Linguistic Understanding ◽

Benchmark Datasets ◽

Linguistic Behavior ◽

Linguistic Models

Do state-of-the-art models for language understanding already have, or can they easily learn, abilities such as boolean coordination, quantification, conditionals, comparatives, and monotonicity reasoning (i.e., reasoning about word substitutions in sentential contexts)? While such phenomena are involved in natural language inference (NLI) and go beyond basic linguistic understanding, it is unclear the extent to which they are captured in existing NLI benchmarks and effectively learned by models. To investigate this, we propose the use of semantic fragments—systematically generated datasets that each target a different semantic phenomenon—for probing, and efficiently improving, such capabilities of linguistic models. This approach to creating challenge datasets allows direct control over the semantic diversity and complexity of the targeted linguistic phenomena, and results in a more precise characterization of a model's linguistic behavior. Our experiments, using a library of 8 such semantic fragments, reveal two remarkable findings: (a) State-of-the-art models, including BERT, that are pre-trained on existing NLI benchmark datasets perform poorly on these new fragments, even though the phenomena probed here are central to the NLI task; (b) On the other hand, with only a few minutes of additional fine-tuning—with a carefully selected learning rate and a novel variation of “inoculation”—a BERT-based model can master all of these logic and monotonicity fragments while retaining its performance on established NLI benchmarks.

Download Full-text

Leveraging Pre-trained Checkpoints for Sequence Generation Tasks

Transactions of the Association for Computational Linguistics ◽

10.1162/tacl_a_00313 ◽

2020 ◽

Vol 8 ◽

pp. 264-280

Author(s):

Sascha Rothe ◽

Shashi Narayan ◽

Aliaksei Severyn

Keyword(s):

Natural Language Processing ◽

Empirical Study ◽

Natural Language ◽

Language Processing ◽

State Of The Art ◽

Text Summarization ◽

Neural Models ◽

Language Understanding ◽

Sequence Generation ◽

Compute Time

Unsupervised pre-training of large neural models has recently revolutionized Natural Language Processing. By warm-starting from the publicly released checkpoints, NLP practitioners have pushed the state-of-the-art on multiple benchmarks while saving significant amounts of compute time. So far the focus has been mainly on the Natural Language Understanding tasks. In this paper, we demonstrate the efficacy of pre-trained checkpoints for Sequence Generation. We developed a Transformer-based sequence-to-sequence model that is compatible with publicly available pre-trained BERT, GPT-2, and RoBERTa checkpoints and conducted an extensive empirical study on the utility of initializing our model, both encoder and decoder, with these checkpoints. Our models result in new state-of-the-art results on Machine Translation, Text Summarization, Sentence Splitting, and Sentence Fusion.

Download Full-text

Where do Clinical Language Models Break Down? A Critical Behavioural Exploration of the ClinicalBERT Deep Transformer Model

Journal of Computational Vision and Imaging Systems ◽

10.15353/jcvis.v6i1.3548 ◽

2021 ◽

Vol 6 (1) ◽

pp. 1-4

Author(s):

Alexander MacLean ◽

Alexander Wong

Keyword(s):

Natural Language ◽

Language Processing ◽

State Of The Art ◽

Language Model ◽

Language Models ◽

Clinical Knowledge ◽

Language Understanding ◽

Improved Performance ◽

Transformer Model ◽

Clinical Domain

The introduction of Bidirectional Encoder Representations from Transformers (BERT) was a major breakthrough for transfer learning in natural language processing, enabling state-of-the-art performance across a large variety of complex language understanding tasks. In the realm of clinical language modeling, the advent of BERT led to the creation of ClinicalBERT, a state-of-the-art deep transformer model pretrained on a wealth of patient clinical notes to facilitate for downstream predictive tasks in the clinical domain. While ClinicalBERT has been widely leveraged by the research community as the foundation for building clinical domain-specific predictive models given its overall improved performance in the Medical Natural Language inference (MedNLI) challenge compared to the seminal BERT model, the fine-grained behaviour and intricacies of this popular clinical language model has not been well-studied. Without this deeper understanding, it is very challenging to understand where ClinicalBERT does well given its additional exposure to clinical knowledge, where it doesn't, and where it can be improved in a meaningful manner. Motivated to garner a deeper understanding, this study presents a critical behaviour exploration of the ClinicalBERT deep transformer model using MedNLI challenge dataset to better understanding the following intricacies: 1) decision-making similarities between ClinicalBERT and BERT (leverage a new metric we introduce called Model Alignment), 2) where ClinicalBERT holds advantages over BERT given its clinical knowledge exposure, and 3) where ClinicalBERT struggles when compared to BERT. The insights gained about the behaviour of ClinicalBERT will help guide towards new directions for designing and training clinical language models in a way that not only addresses the remaining gaps and facilitates for further improvements in clinical language understanding performance, but also highlights the limitation and boundaries of use for such models.

Download Full-text

Knowledge Distillation with Noisy Labels for Natural Language Understanding

10.18653/v1/2021.wnut-1.33 ◽

2021 ◽

Author(s):

Shivendra Bhardwaj ◽

Abbas Ghaddar ◽

Ahmad Rashid ◽

Khalil Bibi ◽

Chengyang Li ◽

...

Keyword(s):

Natural Language ◽

Natural Language Understanding ◽

Language Understanding ◽

Knowledge Distillation ◽

Noisy Labels

Download Full-text

Learning Directional Sentence-Pair Embedding for Natural Language Reasoning (Student Abstract)

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v34i10.7184 ◽

2020 ◽

Vol 34 (10) ◽

pp. 13825-13826

Author(s):

Yuchen Jiang ◽

Zhenxin Xiao ◽

Kai-Wei Chang

Keyword(s):

Natural Language ◽

Abductive Reasoning ◽

Sentence Pair ◽

Reasoning Task ◽

Language Understanding ◽

Cause And Effect ◽

Evaluation Task ◽

Effect Relation ◽

Relation Prediction ◽

Directional Relations

Enabling the models with the ability of reasoning and inference over text is one of the core missions of natural language understanding. Despite deep learning models have shown strong performance on various cross-sentence inference benchmarks, recent work has shown that they are leveraging spurious statistical cues rather than capturing deeper implied relations between pairs of sentences. In this paper, we show that the state-of-the-art language encoding models are especially bad at modeling directional relations between sentences by proposing a new evaluation task: Cause-and-Effect relation prediction task. Back by our curated Cause-and-Effect Relation dataset (Cℰℛ), we also demonstrate that a mutual attention mechanism can guide the model to focus on capturing directional relations between sentences when added to existing transformer-based models. Experiment results show that the proposed approach improves the performance on downstream applications, such as the abductive reasoning task.

Download Full-text

KDAS-ReID: Architecture Search for Person Re-Identification via Distilled Knowledge with Dynamic Temperature

Algorithms ◽

10.3390/a14050137 ◽

2021 ◽

Vol 14 (5) ◽

pp. 137

Author(s):

Zhou Lei ◽

Kangkang Yang ◽

Kai Jiang ◽

Shengbo Chen

Keyword(s):

State Of The Art ◽

Identification Algorithm ◽

Student Model ◽

Deep Convolutional Neural Networks ◽

Fast Speed ◽

Training Stage ◽

Knowledge Distillation ◽

And Training ◽

Better Than ◽

Teacher Model

Person re-Identification(Re-ID) based on deep convolutional neural networks (CNNs) achieves remarkable success with its fast speed. However, prevailing Re-ID models are usually built upon backbones that manually design for classification. In order to automatically design an effective Re-ID architecture, we propose a pedestrian re-identification algorithm based on knowledge distillation, called KDAS-ReID. When the knowledge of the teacher model is transferred to the student model, the importance of knowledge in the teacher model will gradually decrease with the improvement of the performance of the student model. Therefore, instead of applying the distillation loss function directly, we consider using dynamic temperatures during the search stage and training stage. Specifically, we start searching and training at a high temperature and gradually reduce the temperature to 1 so that the student model can better learn from the teacher model through soft targets. Extensive experiments demonstrate that KDAS-ReID performs not only better than other state-of-the-art Re-ID models on three benchmarks, but also better than the teacher model based on the ResNet-50 backbone.

Download Full-text

English-Vietnamese Cross-Lingual Paraphrase Identification Using MT-DNN

Engineering, Technology & Applied Science Research ◽

10.48084/etasr.4300 ◽

2021 ◽

Vol 11 (5) ◽

pp. 7598-7604

Author(s):

H. V. T. Chi ◽

D. L. Anh ◽

N. L. Thanh ◽

D. Dinh

Keyword(s):

Neural Network ◽

Information Retrieval ◽

Natural Language ◽

Deep Neural Network ◽

State Of The Art ◽

Language Models ◽

Language Understanding ◽

Cross Language Information Retrieval ◽

Cross Lingual ◽

Cross Language

Paraphrase identification is a crucial task in natural language understanding, especially in cross-language information retrieval. Nowadays, Multi-Task Deep Neural Network (MT-DNN) has become a state-of-the-art method that brings outstanding results in paraphrase identification [1]. In this paper, our proposed method based on MT-DNN [2] to detect similarities between English and Vietnamese sentences, is proposed. We changed the shared layers of the original MT-DNN from original the BERT [3] to other pre-trained multi-language models such as M-BERT [3] or XLM-R [4] so that our model could work on cross-language (in our case, English and Vietnamese) information retrieval. We also added some tasks as improvements to gain better results. As a result, we gained 2.3% and 2.5% increase in evaluated accuracy and F1. The proposed method was also implemented on other language pairs such as English – German and English – French. With those implementations, we got a 1.0%/0.7% improvement for English – German and a 0.7%/0.5% increase for English – French.

Download Full-text

Syntactic Structure Distillation Pretraining for Bidirectional Encoders

Transactions of the Association for Computational Linguistics ◽

10.1162/tacl_a_00345 ◽

2020 ◽

Vol 8 ◽

pp. 776-794

Author(s):

Adhiguna Kuncoro ◽

Lingpeng Kong ◽

Daniel Fried ◽

Dani Yogatama ◽

Laura Rimell ◽

...

Keyword(s):

Natural Language ◽

Relative Error ◽

Marginal Distribution ◽

Syntactic Structure ◽

Language Model ◽

Structured Prediction ◽

Language Understanding ◽

Knowledge Distillation ◽

Textual Representation ◽

Open Question

Textual representation learners trained on large amounts of data have achieved notable success on downstream tasks; intriguingly, they have also performed well on challenging tests of syntactic competence. Hence, it remains an open question whether scalable learners like BERT can become fully proficient in the syntax of natural language by virtue of data scale alone, or whether they still benefit from more explicit syntactic biases. To answer this question, we introduce a knowledge distillation strategy for injecting syntactic biases into BERT pretraining, by distilling the syntactically informative predictions of a hierarchical—albeit harder to scale—syntactic language model. Since BERT models masked words in bidirectional context, we propose to distill the approximate marginal distribution over words in context from the syntactic LM. Our approach reduces relative error by 2–21% on a diverse set of structured prediction tasks, although we obtain mixed results on the GLUE benchmark. Our findings demonstrate the benefits of syntactic biases, even for representation learners that exploit large amounts of data, and contribute to a better understanding of where syntactic biases are helpful in benchmarks of natural language understanding.

Download Full-text