Unsupervised domain adaptation for lip reading based on cross-modal knowledge distillation

AbstractWe present an unsupervised domain adaptation (UDA) method for a lip-reading model that is an image-based speech recognition model. Most of conventional UDA methods cannot be applied when the adaptation data consists of an unknown class, such as out-of-vocabulary words. In this paper, we propose a cross-modal knowledge distillation (KD)-based domain adaptation method, where we use the intermediate layer output in the audio-based speech recognition model as a teacher for the unlabeled adaptation data. Because the audio signal contains more information for recognizing speech than lip images, the knowledge of the audio-based model can be used as a powerful teacher in cases where the unlabeled adaptation data consists of audio-visual parallel data. In addition, because the proposed intermediate-layer-based KD can express the teacher as the sub-class (sub-word)-level representation, this method allows us to use the data of unknown classes for the adaptation. Through experiments on an image-based word recognition task, we demonstrate that the proposed approach can not only improve the UDA performance but can also use the unknown-class adaptation data.

Download Full-text

Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation

2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) ◽

10.1109/asru.2017.8268911 ◽

2017 ◽

Cited By ~ 17

Author(s):

Wei-Ning Hsu ◽

Yu Zhang ◽

James Glass

Keyword(s):

Speech Recognition ◽

Data Augmentation ◽

Domain Adaptation ◽

Robust Speech Recognition ◽

Unsupervised Domain Adaptation ◽

Variational Autoencoder

Download Full-text

Hearing Lips: Improving Lip Reading by Distilling Speech Recognizers

Proceedings of the AAAI Conference on Artificial Intelligence ◽

10.1609/aaai.v34i04.6174 ◽

2020 ◽

Vol 34 (04) ◽

pp. 6917-6924 ◽

Cited By ~ 1

Author(s):

Ya Zhao ◽

Rui Xu ◽

Xinchao Wang ◽

Peng Hou ◽

Haihong Tang ◽

...

Keyword(s):

Deep Learning ◽

Speech Recognition ◽

Error Rate ◽

Large Scale ◽

State Of The Art ◽

Lip Reading ◽

Speech Recognizers ◽

Lip Movement ◽

Knowledge Distillation ◽

The One

Lip reading has witnessed unparalleled development in recent years thanks to deep learning and the availability of large-scale datasets. Despite the encouraging results achieved, the performance of lip reading, unfortunately, remains inferior to the one of its counterpart speech recognition, due to the ambiguous nature of its actuations that makes it challenging to extract discriminant features from the lip movement videos. In this paper, we propose a new method, termed as Lip by Speech (LIBS), of which the goal is to strengthen lip reading by learning from speech recognizers. The rationale behind our approach is that the features extracted from speech recognizers may provide complementary and discriminant clues, which are formidable to be obtained from the subtle movements of the lips, and consequently facilitate the training of lip readers. This is achieved, specifically, by distilling multi-granularity knowledge from speech recognizers to lip readers. To conduct this cross-modal knowledge distillation, we utilize an efficacious alignment scheme to handle the inconsistent lengths of the audios and videos, as well as an innovative filtering strategy to refine the speech recognizer's prediction. The proposed method achieves the new state-of-the-art performance on the CMLR and LRS2 datasets, outperforming the baseline by a margin of 7.66% and 2.75% in character error rate, respectively.

Download Full-text

Joint Progressive Knowledge Distillation and Unsupervised Domain Adaptation

2020 International Joint Conference on Neural Networks (IJCNN) ◽

10.1109/ijcnn48605.2020.9206989 ◽

2020 ◽

Cited By ~ 1

Author(s):

Le Thanh Nguyen-Meidine ◽

Eric Granger ◽

Madhu Kiran ◽

Jose Dolz ◽

Louis-Antoine Blais-Morin

Keyword(s):

Domain Adaptation ◽

Unsupervised Domain Adaptation ◽

Knowledge Distillation

Download Full-text

Development of Visual and Audio Speech Recognition Systems Using Deep Neural Networks

10.20948/graphicon-2021-3027-905-916 ◽

2021 ◽

Author(s):

Denis Ivanko ◽

Dmitry Ryumin

Keyword(s):

Speech Recognition ◽

State Of The Art ◽

Recognition Performance ◽

Recognition Task ◽

Reading Task ◽

Lip Reading ◽

Recognition Systems ◽

End To End ◽

Future Work ◽

Single Modality

In this paper we design end-to-end neural network for the low-resource lip-reading task and audio speech recognition task using 3D CNNs, pre-trained CNN weights of several state-of- the-art models (e.g. VGG19, InceptionV3, MobileNetV2, etc.) and LSTMs. We present two phrase-level speech recognition pipelines: for lip-reading and acoustic speech recognition. We evaluate different combinations of front-end and back-end modules on the RUSAVIC dataset. We compare our results with traditional 2D CNN approach and demonstrate the increase in recognition accuracy up to 14%. Moreover, we carefully studied existing state-of-the-art models to be use for augmentation. Based on the conducted analysis we have chosen 5 most promising model’s architectures and evaluated them on own data. We have tested our systems on a real-word data of two different scenarios: recorded in idling vehicle and during actual driving. Our independently trained systems demonstrated acoustic speech accuracy up to 90% and lip-reading accuracy up to 61%. Future work will focus on the fusion of visual and audio speech modalities and on speaker adaptation. We expect that fused multi-modal information will help to further improve recognition performance compared to a single modality. Another possible direction could be the research of different NN-based architectures to better tackle end-to-end lip-reading task.

Download Full-text

Unsupervised Domain Adaptation for Speech Recognition via Uncertainty Driven Self-Training

ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) ◽

10.1109/icassp39728.2021.9414299 ◽

2021 ◽

Author(s):

Sameer Khurana ◽

Niko Moritz ◽

Takaaki Hori ◽

Jonathan Le Roux

Keyword(s):

Speech Recognition ◽

Domain Adaptation ◽

Unsupervised Domain Adaptation

Download Full-text

Unsupervised Domain Adaptation for Facial Expression Recognition Using Generative Adversarial Networks

Computational Intelligence and Neuroscience ◽

10.1155/2018/7208794 ◽

2018 ◽

Vol 2018 ◽

pp. 1-10 ◽

Cited By ~ 8

Author(s):

Xiaoqing Wang ◽

Xiangjun Wang ◽

Yubo Ni

Keyword(s):

Facial Expression ◽

Facial Expression Recognition ◽

Domain Adaptation ◽

Recognition Task ◽

Fine Tuning ◽

Generative Adversarial Networks ◽

Expression Recognition ◽

Generative Adversarial Network ◽

Unsupervised Domain Adaptation ◽

Adversarial Network

In the facial expression recognition task, a good-performing convolutional neural network (CNN) model trained on one dataset (source dataset) usually performs poorly on another dataset (target dataset). This is because the feature distribution of the same emotion varies in different datasets. To improve the cross-dataset accuracy of the CNN model, we introduce an unsupervised domain adaptation method, which is especially suitable for unlabelled small target dataset. In order to solve the problem of lack of samples from the target dataset, we train a generative adversarial network (GAN) on the target dataset and use the GAN generated samples to fine-tune the model pretrained on the source dataset. In the process of fine-tuning, we give the unlabelled GAN generated samples distributed pseudolabels dynamically according to the current prediction probabilities. Our method can be easily applied to any existing convolutional neural networks (CNN). We demonstrate the effectiveness of our method on four facial expression recognition datasets with two CNN structures and obtain inspiring results.

Download Full-text