Improving the Performance of ASR System by Building Acoustic Models using Spectro-Temporal and Phase-Based Features

Circuits Systems and Signal Processing ◽

10.1007/s00034-021-01848-w ◽

2021 ◽

Author(s):

Anirban Dutta ◽

G. Ashishkumar ◽

Ch. V. Rama Rao

Keyword(s):

Acoustic Models ◽

Download Full-text

Development and analysis of Punjabi ASR system for mobile phones under different acoustic models

International Journal of Speech Technology ◽

10.1007/s10772-019-09593-x ◽

2019 ◽

Vol 22 (1) ◽

pp. 219-230 ◽

Author(s):

Puneet Mittal ◽

Navdeep Singh

Keyword(s):

Mobile Phones ◽

Acoustic Models ◽

Download Full-text

An Investigation of Multilingual TDNN-BLSTM Acoustic Modeling for Hindi Speech Recognition

International Journal of Sensors Wireless Communications and Control ◽

10.2174/2210327911666210118143758 ◽

2021 ◽

Vol 11 ◽

Author(s):

Ankit Kumar ◽

Rajesh Kumar Aggarwal

Keyword(s):

Neural Network ◽

Speech Recognition ◽

High Accuracy ◽

Training Data ◽

Acoustic Modeling ◽

Training Dataset ◽

Acoustic Model ◽

Indian Languages ◽

Acoustic Models ◽

Background: In India, thousands of languages or dialects are in use. Most Indian dialects are low asset dialects. A well-performing Automatic Speech Recognition (ASR) system for Indian languages is unavailable due to a lack of resources. Hindi is one of them as large vocabulary Hindi speech datasets are not freely available. We have only a few hours of transcribed Hindi speech dataset. There is a lot of time and money involved in creating a well-transcribed speech dataset. Thus, developing a real-time ASR system with a few hours of the training dataset is the most challenging task. The different techniques like data augmentation, semi-supervised training, multilingual architecture, and transfer learning, have been reported in the past to tackle the fewer speech data issues. In this paper, we examine the effect of multilingual acoustic modeling in ASR systems for the Hindi language. Objective: This article’s objective is to develop a high accuracy Hindi ASR system with a reasonable computational load and high accuracy using a few hours of training data. Method: To achieve this goal we used Multilingual training with Time Delay Neural Network- Bidirectional Long Short Term Memory (TDNN-BLSTM) acoustic modeling. Multilingual acoustic modeling has significantly improved the ASR system's performance for low and limited resource languages. The common practice is to train the acoustic model by merging data from similar languages. In this work, we use three Indian languages, namely Hindi, Marathi, and Bengali. Hindi with 2.5 hours of training data and Marathi with 5.5 hours of training data and Bengali with 28.5 hours of transcribed data, was used in this work to train the proposed model. Results: The Kaldi toolkit was used to perform all the experiments. The paper is investigated over three main points. First, we present the monolingual ASR system using various Neural Network (NN) based acoustic models. Second, we show that Recurrent Neural Network (RNN) language modeling helps to improve the ASR performance further. Finally, we show that a multilingual ASR system significantly reduces the Word Error Rate (WER) (absolute 2% WER reduction for Hindi and 3% for the Marathi language). In all the three languages, the proposed TDNN-BLSTM-A multilingual acoustic models help to get the lowest WER. Conclusion: The multilingual hybrid TDNN-BLSTM-A architecture shows a 13.67% relative improvement over the monolingual Hindi ASR system. The best WER of 8.65% was recorded for Hindi ASR. For Marathi and Bengali, the proposed TDNN-BLSTM-A acoustic model reports the best WER of 30.40% and 10.85%.

Download Full-text

How Neural Network Depth Compensates for HMM Conditional Independence Assumptions in DNN-HMM Acoustic Models

10.21437/interspeech.2016-283 ◽

2016 ◽

Author(s):

Suman Ravuri ◽

Steven Wegmann

Keyword(s):

Neural Network ◽

Conditional Independence ◽

Acoustic Models ◽

Independence Assumptions

Download Full-text

Impact of Aliasing on Deep CNN-Based End-to-End Acoustic Models

10.21437/interspeech.2018-1371 ◽

2018 ◽

Author(s):

Yuan Gong ◽

Christian Poellabauer

Keyword(s):

Acoustic Models ◽

Deep Cnn ◽

Download Full-text

Supervised Learning of Acoustic Models in a Zero Resource Setting to Improve DPGMM Clustering

10.21437/interspeech.2016-988 ◽

2016 ◽

Author(s):

Michael Heck ◽

Sakriani Sakti ◽

Satoshi Nakamura

Keyword(s):

Supervised Learning ◽

Acoustic Models ◽

Resource Setting

Download Full-text

Domain Adaptation of CNN Based Acoustic Models Under Limited Resource Settings

10.21437/interspeech.2016-1161 ◽

2016 ◽

Author(s):

Masayuki Suzuki ◽

Ryuki Tachibana ◽

Samuel Thomas ◽

Bhuvana Ramabhadran ◽

George Saon

Keyword(s):

Domain Adaptation ◽

Limited Resource ◽

Acoustic Models

Download Full-text

Ensembles of Multi-Scale VGG Acoustic Models

10.21437/interspeech.2017-920 ◽

2017 ◽

Author(s):

Michael Heck ◽

Masayuki Suzuki ◽

Takashi Fukuda ◽

Gakuto Kurata ◽

Satoshi Nakamura

Keyword(s):

Acoustic Models ◽

Download Full-text

Improving DNN Bluetooth Narrowband Acoustic Models by Cross-Bandwidth and Cross-Lingual Initialization

10.21437/interspeech.2017-1129 ◽

2017 ◽

Author(s):

Xiaodan Zhuang ◽

Arnab Ghoshal ◽

Antti-Veikko Rosti ◽

Matthias Paulik ◽

Daben Liu

Keyword(s):

Acoustic Models ◽

Download Full-text

Improved Multilingual Training of Stacked Neural Network Acoustic Models for Low Resource Languages

10.21437/interspeech.2016-1426 ◽

2016 ◽

Author(s):

Tanel Alumäe ◽

Stavros Tsakalidis ◽

Richard Schwartz

Keyword(s):

Neural Network ◽

Acoustic Models ◽

Download Full-text

CTC Training of Multi-Phone Acoustic Models for Speech Recognition

10.21437/interspeech.2017-505 ◽

2017 ◽

Author(s):

Olivier Siohan

Keyword(s):

Speech Recognition ◽

Acoustic Models

Download Full-text